PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Computational identification of residues that modulate voltage sensitivity of voltage-gated potassium channels.

BACKGROUND: Studies of the structure-function relationship in proteins for which no 3D structure is available are often based on inspection of multiple sequence alignments. Many functionally important residues of proteins can be identified because they are conserved during evolution. However, residues that vary can also be critically important if their variation is responsible for diversity of protein function and improved phenotypes. If too few sequences are studied, the support for hypotheses on the role of a given residue will be weak, but analysis of large multiple alignments is too complex for simple inspection. When a large body of sequence and functional data are available for a protein family, mature data mining tools, such as machine learning, can be applied to extract information more easily, sensitively and reliably. We have undertaken such an analysis of voltage-gated potassium channels, a transmembrane protein family whose members play indispensable roles in electrically excitable cells. RESULTS: We applied different learning algorithms, combined in various implementations, to obtain a model that predicts the half activation voltage of a voltage-gated potassium channel based on its amino acid sequence. The best result was obtained with a k-nearest neighbor classifier combined with a wrapper algorithm for feature selection, producing a mean absolute error of prediction of 7.0 mV. The predictor was validated by permutation test and evaluation of independent experimental data. Feature selection identified a number of residues that are predicted to be involved in the voltage sensitive conformation changes; these residues are good target candidates for mutagenesis analysis. CONCLUSION: Machine learning analysis can identify new testable hypotheses about the structure/function relationship in the voltage-gated potassium channel family. This approach should be applicable to any protein family if the number of training examples and the sequence diversity of the training set that are necessary for robust prediction are empirically validated. The predictor and datasets can be found at the VKCDB web site.

Algorithms↗

Combinatorial use of mRNA and two-dimensional electrophoresis expression data to choose relevant features for mass spectrometric identification.

It is only recently that quantitative studies of differential proteome analysis (DPA) have become possible. In this paper the issues involved in quantitative DPA are discussed and novel tools to select features for identification by mass spectrometry (MS) are described. The problem of comparing two sets of gels on a global level is explored as well as how to find specific protein features that differentiate two sets of two-dimensional electrophoresis gels. The concept of a 'virtual' gel, derived from gene expression data, is introduced. The virtual gel enables the co-analysis of data from gene and protein expression. We discuss the value of such an approach, and consider what new information can be gained by using gene and protein expression together. These tools are illustrated by analysis of data from tandem gene and protein expression experiments. Features that are highlighted by the above methods are putative candidates for MS identification. Tools are described that integrate the process of feature selection, cutting, and MS analysis.

Breast↗

Machine learning based pattern recognition applied to microarray data.

MOTIVATION: Microarrays have allowed the expression level of thousands of genes or proteins to be measured simultaneously. Data sets generated by these arrays consist of a small number of observations (e.g., 20-100 samples) on a very large number of variables (e.g., 10,000 genes or proteins). The observations in these data sets often have other attributes associated with them such as a class label denoting the pathology of the subject. Finding the genes or proteins that are correlated to these attributes is often a difficult task since most of the variables do not contain information about the pathology and as such can mask the identity of the relevant features. We describe a genetic algorithm (GA) that employs both supervised and unsupervised learning to mine gene expression and proteomic data. The pattern recognition GA selects features that increase clustering, while simultaneously searching for features that optimize the separation of the classes in a plot of the two or three largest principal components of the data. Because the largest principal components capture the bulk of the variance in the data, the features chosen by the GA contain information primarily about differences between classes in the data set. The principal component analysis routine embedded in the fitness function of the GA acts as an information filter, significantly reducing the size of the search space since it restricts the search to feature sets whose principal component plots show clustering on the basis of class. The algorithm integrates aspects of artificial intelligence and evolutionary computations to yield a smart one pass procedure for feature selection, clustering, classification, and prediction.

Algorithms↗

Assessing swine thermal comfort by image analysis of postural behaviors.

Postural behavior is an integral response of animals to complex environmental factors. Huddling, nearly contacting one another on the side, and spreading are common postural behaviors of group-housed animals undergoing cold, comfortable, and warm/hot sensations, respectively. These postural patterns have been routinely used by animal caretakers to assess thermal comfort of the animals and to make according adjustment on the environmental settings or management schemes. This manual adjustment approach, however, has the inherent limitations of daily discontinuity and inconsistency between caretakers in interpretation of the animal comfort behavior. The goal of this project was to explore a novel, automated image analysis system that would assess the thermal comfort of swine and make proper environmental adjustments to enhance animal wellbeing and production efficiency. This paper describes the progress and on-going work toward the achievement of our proposed goal. The feasibility of classifying the thermal comfort state of young pigs by neural network (NN) analysis of their postural images was first examined. It included exploration of using certain feature selections of the postural behavioral images as the input to a three-layer NN that was trained to classify the corresponding thermal comfort state as being cold, comfortable, or warm. The image feature selections, a critical step for the classification, examined in this study included Fourier coefficient (FC), moment (M), perimeter and area (P&A), and combination of M and P&A of the processed binary postural images. The result was positive, with the combination of M and P&A as the input feature to the NN yielding the highest correct classification rate. Subsequent work included the development of hardware and computational algorithms that enable automatic image segmentation, motion detection, and the selection of the behavioral images suitable for use in the classification. Work is in progress to quantify the relationships of postural behavior and physiological responses of pigs using thermographs. The results are expected to facilitate objective training of NN, hence improving the accuracy of the postural image-based assessment of the thermal comfort state. Work is also in progress to implement the analysis and assessment algorithms into computer codes for real-time application.

Animal Welfare↗

PACK: Profile Analysis using Clustering and Kurtosis to find molecular classifiers in cancer.

MOTIVATION: Elucidating the molecular taxonomy of cancers and finding biological and clinical markers from microarray experiments is problematic due to the large number of variables being measured. Feature selection methods that can identify relevant classifiers or that can remove likely false positives prior to supervised analysis are therefore desirable. RESULTS: We present a novel feature selection procedure based on a mixture model and a non-gaussianity measure of a gene's expression profile. The method can be used to find genes that define either small outlier subgroups or major subdivisions, depending on the sign of kurtosis. The method can also be used as a filtering step, prior to supervised analysis, in order to reduce the false discovery rate. We validate our methodology using six independent datasets by rediscovering major classifiers in ER negative and ER positive breast cancer and in prostate cancer. Furthermore, our method finds two novel subtypes within the basal subgroup of ER negative breast tumours, associated with apoptotic and immune response functions respectively, and with statistically different clinical outcome. AVAILABILITY: An R-function pack that implements the methods used here has been added to vabayelMix, available from (www.cran.r-project.org). CONTACT: aet21@cam.ac.uk SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.

Algorithms↗

Pituicytoma: diagnostic features on selective carotid angiography and MR imaging.

We report a case of pituicytoma, a rare primary tumor of the neurohypophysis. A 64-year-old man presented with progressive visual complaints, bitemporal hemianopsia, and headache. Imaging studies revealed distinctive features of a mass lesion that thickened the pituitary stalk with a bilobed protrusion extending into the hypothalamus. Angiography demonstrated tumor vascular supply from the superior hypophyseal arteries representing the diencephalic branches of the internal carotid arteries. We discuss the imaging and pathology of this unusual tumor.

Astrocytoma↗

Stabilometric parameters are affected by anthropometry and foot placement.

OBJECTIVE: To recognize and quantify the influence of biomechanical factors, namely anthropometry and foot placement, on the more common measures of stabilometric performance, including new-generation stochastic parameters. DESIGN: Fifty normal-bodied young adults were selected in order to cover a sufficiently wide range of anthropometric properties. They were allowed to choose their preferred side-by-side foot position and their quiet stance was recorded with eyes open and closed by a force platform. BACKGROUND: biomechanical factors are known to influence postural stability but their impact on stabilometric parameters has not been extensively explored yet. METHODS: Principal component analysis was used for feature selection among several biomechanical factors. A collection of 55 stabilometric parameters from the literature was estimated from the center-of-pressure time series. Linear relations between stabilometric parameters and selected biomechanical factors were investigated by robust regression techniques. RESULTS: The feature selection process returned height, weight, maximum foot width, base-of-support area, and foot opening angle as the relevant biomechanical variables. Only eleven out of the 55 stabilometric parameters were completely immune from a linear dependence on these variables. The remaining parameters showed a moderate to high dependence that was strengthened upon eye closure. For these parameters, a normalization procedure was proposed, to remove what can well be considered, in clinical investigations, a spurious source of between-subject variability. CONCLUSION: Care should be taken when quantifying postural sway through stabilometric parameters. It is suggested as a good practice to include some anthropometric measurements in the experimental protocol, and to standardize or trace foot position. RELEVANCE: Although the role of anthropometry and foot placement has been investigated in specific studies, there are no studies in the literature that systematically explore the relationship between such BF and stabilometric parameters. This knowledge may contribute to better defining the experimental protocol and improving the functional evaluation of postural sway for clinical purposes, e.g. by removing through normalization the spurious effects of body properties and foot position on postural performance.

Adult↗

Personalized diagnosis by cached solutions with hypertension as a study model.

Statistical modeling of links between genetic profiles with environmental and clinical data to aid in medical diagnosis is a challenge. Here, we present a computational approach for rapidly selecting important clinical data to assist in medical decisions based on personalized genetic profiles. What could take hours or days of computing is available on-the-fly, making this strategy feasible to implement as a routine without demanding great computing power. The key to rapidly obtaining an optimal/nearly optimal mathematical function that can evaluate the "disease stage" by combining information of genetic profiles with personal clinical data is done by querying a precomputed solution database. The database is previously generated by a new hybrid feature selection method that makes use of support vector machines, recursive feature elimination and random sub-space search. Here, to evaluate the method, data from polymorphisms in the renin-angiotensin-aldosterone system genes together with clinical data were obtained from patients with hypertension and control subjects. The disease "risk" was determined by classifying the patients' data with a support vector machine model based on the optimized feature; then measuring the Euclidean distance to the hyperplane decision function. Our results showed the association of renin-angiotensin-aldosterone system gene haplotypes with hypertension. The association of polymorphism patterns with different ethnic groups was also tracked by the feature selection process. A demonstration of this method is also available online on the project's web site.

Algorithms↗

The influence of sustained selective attention on stimulus selectivity in macaque visual area MT.

Remarkable alterations of perception during long-lasting attentional processes have been described in several recent studies. Although these findings have gained much interest, almost nothing is known about the modulation of neuronal responses during sustained attention. Therefore, we investigated the effect of prolonged selective attention on neuronal feature selectivity. Awake macaque monkeys were trained to perform a motion-tracking task that required attending one of two simultaneously presented moving bars for up to 15 sec. Extracellular recordings were obtained from neurons in macaque motion-sensitive middle temporal visual area (MT/V5). Under conditions of attention, we found high and constant direction selectivity over time. This was expressed by a strong and persistent response contrast between presentations of preferred and nonpreferred stimuli in successive motion cycles. With attention directed to another moving bar, neuronal responses to the behaviorally irrelevant stimulus became continuously less specific for the direction of motion. In particular, increasingly higher firing rates for motion in null direction caused a strong reduction of direction selectivity, which further increased with enhanced proximity between target and distracter bar. A passive condition experiment revealed that this reduction occurred only when motion remained the behaviorally relevant feature but disappeared when attention was withdrawn from this feature domain. Thus, sustained attention seems to stabilize direction selectivity of neurons in area MT against a time and competition-dependent degradation, whereas nonattended objects suffer from a reduced neuronal representation.

Animals↗

Bayesian neural networks for classification: how useful is the evidence framework?

This paper presents an empirical assessment of the Bayesian evidence framework for neural networks using four synthetic and four real-world classification problems. We focus on three issues; model selection, automatic relevance determination (ARD) and the use of committees. Model selection using the evidence criterion is only tenable if the number of training examples exceeds the number of network weights by a factor of five or ten. With this number of available examples, however, cross-validation is a viable alternative. The ARD feature selection scheme is only useful in networks with many hidden units and for data sets containing many irrelevant variables. ARD is also useful as a hard feature selection method. Results on applying the evidence framework to the real-world data sets showed that committees of Bayesian networks achieved classification accuracies similar to the best alternative methods. Importantly, this was achievable with a minimum of human intervention.

Journal Article↗

The use of pathologic features in selecting the extent of surgical resection necessary for breast cancer patients treated by primary radiation therapy.

The extent of the surgical resection necessary for breast cancer patients treated by primary radiation therapy is unknown. A simple gross excision of the tumor provides the best cosmetic result, but a wide local resection may be important to prevent local recurrence in some patients. In order to identify patients who are not adequately treated by gross excision of the tumor and radiation therapy, we performed a retrospective clinical-pathologic review of 221 treated women with infiltrating duct carcinoma. There were 53 cases in which the excision specimen showed a constellation of three pathologic features: prominent intraductal carcinoma in the tumor, intraductal carcinoma in the grossly-normal adjacent tissue, and poorly-differentiated nuclei. These cases had a 37% risk of a local recurrence at 6 years compared to eight per cent for all other cases (p less than 0.0001). In cases with all three features, the use of a supplemental dose of radiation to the primary site did not significantly reduce the risk of a local recurrence. Local recurrence at 6 years was 34% in cases with all three features, who received supplemental local radiation, compared to 49% in cases not receiving a supplemental dose (p = 0.28). Survival was also worse for patients with all three features compared to other cases (69% vs. 90% at 6 years, p = 0.002). These results indicate that patients with all three pathologic features have a high risk of local recurrence following gross excision of the tumor and radiation therapy. If primary radiation therapy is selected for these patients, they should first undergo a re-excision of the tumor site in order to be certain that areas of extensive intraductal carcinoma have been adequately resected. Patients whose tumors do not show all three features are adequately treated by gross excision of the tumor prior to radiation therapy.

Breast↗

Classification and knowledge discovery in protein databases.

We consider the problem of classification in noisy, high-dimensional, and class-imbalanced protein datasets. In order to design a complete classification system, we use a three-stage machine learning framework consisting of a feature selection stage, a method addressing noise and class-imbalance, and a method for combining biologically related tasks through a prior-knowledge based clustering. In the first stage, we employ Fisher's permutation test as a feature selection filter. Comparisons with the alternative criteria show that it may be favorable for typical protein datasets. In the second stage, noise and class imbalance are addressed by using minority class over-sampling, majority class under-sampling, and ensemble learning. The performance of logistic regression models, decision trees, and neural networks is systematically evaluated. The experimental results show that in many cases ensembles of logistic regression classifiers may outperform more expressive models due to their robustness to noise and low sample density in a high-dimensional feature space. However, ensembles of neural networks may be the best solution for large datasets. In the third stage, we use prior knowledge to partition unlabeled data such that the class distributions among non-overlapping clusters significantly differ. In our experiments, training classifiers specialized to the class distributions of each cluster resulted in a further decrease in classification error.

Algorithms↗

A novel approach to electrode signal analysis for glucose determination.

The methodology proposed in this presentation consists in considering the stationary Pt-electrode of an electrocatalytic sensor aimed at glucose measurement together with the reference electrode as a "black box" for which a mathematical model is assumed. The model correlates selected features of the output signal to the concentration of glucose and of interfering substances (urea, amino acids) and to their interactions. The model parameters are experimentally identified. During the measurement, the values of previously selected features of sensor output signal are determined; then they serve as the input data for computation of concentrations of glucose and of interfering substances.

Biosensing Techniques↗

Three-dimensional reconstruction of biological sections.

To observe internal detail in biological structures using light or electron microscopy, specimens need to be prepared from thin sections. Quantitative analysis of these section requires the transfer of complete section images, or selected features from them, into a computer. Where a complete structure is represented by a series of consecutive sections, images may be combined to represent the three dimensional structure. An interactive computer system is described which enables selected features of serial section images to be entered into the computer, edited, reconstructed in three-dimensions and displayed in any orientation. Sections can be analysed individually using appropriate, compatible programs, and an assessment of the complete structure obtained by interpolating between them. Categorisation of the different substructures within the sections allows them to be analysed and displayed separately.

Animals↗

Naming in young children: a dumb attentional mechanism?

Previous studies have shown that young children selectively attend to some object properties and ignore others when generalizing a newly learned object name. Moreover, the specific properties children attend to depend on the stimulus and task context. The present study tested an attentional account: that children's feature selection in name generalization is guided by non-strategic attentional processes that are minimally influenced by new conceptual information presented in the task. Four experiments presented 3-year-old children and adults with novel artifacts consisting of distinctive base objects with appended parts. In a Name condition, subjects were asked whether test objects had the same name as the exemplar. In a Similarity condition, subjects made similarity judgments for the same objects. Subjects in two experiments were shown a function for either the base object or the parts. Both adults' naming and similarity judgments were influenced by the functional information. Children's similarity judgments were also influenced by the functions. However, children's naming was immune to influence from information about function. Instead, children's feature selection in naming was shifted only by changes in the relative salience of base objects and parts. The results are consistent with the idea that dumb attentional processes are responsible for young children's smart generalizations of novel words to new instances. Potential mechanisms to explain these findings are discussed.

Adult↗

Semi-supervised learning via penalized mixture model with application to microarray sample classification.

MOTIVATION: It is biologically interesting to address whether human blood outgrowth endothelial cells (BOECs) belong to or are closer to large vessel endothelial cells (LVECs) or microvascular endothelial cells (MVECs) based on global expression profiling. An earlier analysis using a hierarchical clustering and a small set of genes suggested that BOECs seemed to be closer to MVECs. By taking advantage of the two known classes, LVEC and MVEC, while allowing BOEC samples to belong to either of the two classes or to form their own new class, we take a semi-supervised learning approach; for high-dimensional data as encountered here, we propose a penalized mixture model with a weighted L1 penalty to realize automatic feature selection while fitting the model. RESULTS: We applied our penalized mixture model to a combined dataset containing 27 BOEC, 28 LVEC and 25 MVEC samples. Analysis results indicated that the BOEC samples appeared to form their own new class. A simulation study confirmed that, compared with the standard mixture model with or without initial variable selection, the penalized mixture model performed much better in identifying relevant genes and forming corresponding clusters. The penalized mixture model seems to be promising for high-dimensional data with the capability of novel class discovery and automatic feature selection.

Algorithms↗

Sequence Data Analysis for Long Disordered Regions Prediction in the Calcineurin Family.

Our recently reported results (PSB 3:471-482, 1998; Proc. IEEE Intnl. Conf. Neural Networks 1:90-95, 1997; PSB 3:435-446, 1998) provide strong support for a hypothesis that some amino acid sequences code for disordered regions rather than structured ones and that such disordered regions are commonly involved in function. General and family-specific neural network predictors developed in those previous studies suggest that different classes of disordered regions exist. Here, family-specific data preprocessing for disorder prediction in the calcineurin (CaN) family is explored. The results show that prediction of order and disorder on CaN sequence data benefits significantly from the use of family-specific preprocessing, with feature extraction through principal components analysis (PCA) outperforming feature selection techniques, although all methods do a good job of discriminating CaN-specific disordered regions from CaN-specific ordered regions. On the other hand, for the discrimination of CaN-specific disordered regions from general (unrelated to CaN) ordered regions, feature selection approaches proved to be more appropriate than PCA. The results further support a hypothesis that different kinds of disordered regions exist, as all family-specific disorder predictors developed in this study significantly outperformed a previously reported general multi-family disorder predictor.

Journal Article↗

Mechanisms for precuing superiority in visual recognition.

Superior recognition, when similar alternatives are studied prior to brief target pictures of common objects, is usually attributed to selective feature analysis, which is beneficial only when the target is not readily discriminable in the alternatives. The present study found equal precuing superiority, however, for similar and dissimilar alternatives with abstract geometric stimuli, and with pictures of common objects, precuing superiority occurred only when each alternative represented the same semantic category. Precuing effects did not appear when different categories were represented in the alternatives, and performance was similar to other conditions in which only a categorical match between target and alternatives was possible. These results suggest that selective feature analysis and precuing superiority only occur when the target cannot be identified by its semantic category. With dissimilar alternatives activation of the target's category can be used to select a response, and there is no benefit from precuing.

Attention↗