PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A genetic algorithm based nearest neighbor classification to breast cancer diagnosis.

This paper presents an application of a hybrid approach (the genetic algorithms and the k-nearest neighbour) proposed by Ishbuchi to Wisconsin breast cancer data. For the diagnosis of breast cancer, the determination of the presence of benign/malignant breast tumors represents a very complex problem (even for an experienced cytologist). Therefore the automatic classification of benign and malignant symptoms is highly desirable as a valuable aid to assist oncologists in the decision making of the diagnosis of breast cancer. In this paper, the genetic algorithm based k-nearest neighbour method for classification of benign and malignant breast tumors is presented. The genetic-algorithm (GA) is used for finding a compact reference set by selecting a small number of reference patterns from a large number of training patterns in nearest neighbor classification. The GA simultaneously performs feature selection and pattern selection and prunes unnecessary features. The goal is to maximize the classification performance of the reference set and minimize the number of selected patterns and features. Results are also compared with a fuzzy-genetic approach where each reference patten represents a fuzzy if-then rule with a circular-cone-type membership function.

Algorithms↗

Individualized markers optimize class prediction of microarray data.

BACKGROUND: Identification of molecular markers for the classification of microarray data is a challenging task. Despite the evident dissimilarity in various characteristics of biological samples belonging to the same category, most of the marker--selection and classification methods do not consider this variability. In general, feature selection methods aim at identifying a common set of genes whose combined expression profiles can accurately predict the category of all samples. Here, we argue that this simplified approach is often unable to capture the complexity of a disease phenotype and we propose an alternative method that takes into account the individuality of each patient-sample. RESULTS: Instead of using the same features for the classification of all samples, the proposed technique starts by creating a pool of informative gene-features. For each sample, the method selects a subset of these features whose expression profiles are most likely to accurately predict the sample's category. Different subsets are utilized for different samples and the outcomes are combined in a hierarchical framework for the classification of all samples. Moreover, this approach can innately identify subgroups of samples within a given class which share common feature sets thus highlighting the effect of individuality on gene expression. CONCLUSION: In addition to high classification accuracy, the proposed method offers a more individualized approach for the identification of biological markers, which may help in better understanding the molecular background of a disease and emphasize the need for more flexible medical interventions.

Biomarkers↗

BioViews: Java-based tools for genomic data visualization.

Visualization tools for bioinformatics ideally should provide universal access to the most current data in an interactive and intuitive graphical user interface. Since the introduction of Java, a language designed for distributed programming over the Web, the technology now exists to build a genomic data visualization tool that meets these requirements. Using Java we have developed a prototype genome browser applet (BioViews) that incorporates a three-level graphical view of genomic data: a physical map, an annotated sequence map, and a DNA sequence display. Annotated biological features are displayed on the physical and sequence-based maps, and the different views are interconnected. The applet is linked to several databases and can retrieve features and display hyperlinked textual data on selected features. In addition to browsing genomic data, different types of analyses can be performed interactively and the results of these analyses visualized alongside prior annotations. Our genome browser is built on top of extensible, reusable graphic components specifically designed for bioinformatics. Other groups can (and do) reuse this work in various ways. Genome centers can reuse large parts of the genome browser with minor modifications, bioinformatics groups working on sequence analysis can reuse components to build front ends for analysis programs, and biology laboratories can reuse components to publish results as dynamic Web documents.

Animals↗

Three-dimensional computer modelling system for the study of biological structures.

A three-dimensional computer modelling system has been developed for use in biology, and is currently running on a Sun3 computer. The data originate as a series of two-dimensional micrographs which are digitised via a TV camera. The two-dimensional images are used to select features of interest and to construct a three-dimensional model. This model can be viewed in vector or solid format, it can be rotated about three orthogonal axes and can be viewed in three dimensions as a stereo pair or an anaglyph. The system has been used in a large number of projects over the past 10-15 years, for example, to examine physiological and nerve structures. The time-consuming part of the process is the selection of features, which involves a high level of biological expertise. Present developments are concerned with reduction of the time spent in feature recognition and involve the introduction of expert systems together with human-computer interaction to deal with problems of identification.

Computer Simulation↗

Predicting radiotherapy-induced cardiac perfusion defects.

The purpose of this work is to compare the efficacy of mathematical models in predicting the occurrence of radiotherapy-induced left ventricular perfusion defects assessed using single-photon emission computed tomography (SPECT). The basis of this study is data from 73 left-sided breast/ chestwall patients treated with tangential photon fields. The mathematical models compared were three commonly used parametric models [Lyman normal tissue complication probability (LNTCP), relative serialty (RS), generalized equivalent uniform dose (gEUD)] and a nonparametric model (Linear discriminant analysis--LDA). Data used by the models were the left ventricular dose--volume histograms, or SPECT-based dose-function histograms, and the presence/absence of SPECT perfusion defects 6 months postradiation therapy (21 patients developed defects). For the parametric models, maximum likelihood estimation and F-tests were used to fit the model parameters. The nonparametric LDA model step-wise selected features (volumes/function above dose levels) using a method based on receiver operating characteristics (ROC) analysis to best separate the groups with and without defects. Optimistic (upper bound) and pessimistic (lower bound) estimates of each model's predictive capability were generated using ROC curves. A higher area under the ROC curve indicates a more accurate model (a model that is always accurate has area = 1). The areas under these curves for different models were used to statistically test for differences between them. Pessimistic estimates of areas under the ROC curve using dose-volume histogram/ dose-function histogram inputs, in order of increasing prediction accuracy, were LNTCP (0.79/0.75), RS (0.80/0.77), gEUD (0.81/0.78), and LDA (0.84/0.86). Only the LDA model benefited from SPECT-based regional functional information. In general, the LDA model was statistically superior to the parametric models. The LDA model selected as features the left ventricular volumes above approximately 23 Gy (V23), essentially volume in field, and 33 Gy (V33), as best separating the groups with and without defects. In conclusion, the nonparametric LDA model appears to be a more accurate predictor of radiotherapy-induced left ventricular perfusion defects than commonly used parametric models.

Dose-Response Relationship, Radiation↗

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross‐taxon transferability↗

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans↗

Linear and nonlinear QSAR study of N-hydroxy-2-[(phenylsulfonyl)amino]acetamide derivatives as matrix metalloproteinase inhibitors.

The inhibitory activity (IC50) toward matrix metalloproteinases (MMP-1, MMP-2, MMP-3, MMP-9, and MMP-13) of N-hydroxy-2-[(phenylsulfonyl)amino]acetamide derivatives (HPSAAs) has been successfully modeled using 2D autocorrelation descriptors. The relevant molecular descriptors were selected by linear and nonlinear genetic algorithm (GA) feature selection using multiple linear regression (MLR) and Bayesian-regularized neural network (BRANN) approaches, respectively. The quality of the models was evaluated by means of cross-validation experiments and the best results correspond to nonlinear ones (Q2>0.7 for all models). Despite the high correlation between the studied compound IC50 values, the 2D autocorrelation space brings different descriptors for each MMP inhibition. On the basis of these results, these models contain useful molecular information about the ligand specificity for MMP S'1, S1, and S'2 pockets.

Acetamides↗

Bayesian-regularized genetic neural networks applied to the modeling of non-peptide antagonists for the human luteinizing hormone-releasing hormone receptor.

Bayesian-regularized genetic neural networks (BRGNNs) were used to model the binding affinity (IC(50)) for 128 non-peptide antagonists for the human luteinizing hormone-releasing hormone (LHRH) receptor using 2D spatial autocorrelation vectors. As a preliminary step, a linear dependence was established by multiple linear regression (MLR) approach, selecting the relevant descriptors by genetic algorithm (GA) feature selection. The linear model showed to fit the training set (N=102) with R(2)=0.746, meanwhile BRGNN exhibited a higher value of R(2)=0.871. Beyond the improvement of training set fitting, the BRGNN model overcame the linear one by being able to describe 85% of test set (N=26) variance in comparison with 73% the MLR model. Our non-linear QSAR model illustrates the importance of an adequate distribution of atomic properties represented in topological frames and reveals the electronegativities, masses and polarizabilities as the most influencing atomic properties in the structures of the heterocycles under analysis for having an appropriate LHRH antagonistic activity. Furthermore, the ability of the non-linear selected variables for differentiating the data was evidenced when total data set was well distributed in a Kohonen self-organizing map (SOM).

Algorithms↗

Optimized between-group classification: a new jackknife-based gene selection procedure for genome-wide expression data.

BACKGROUND: A recent publication described a supervised classification method for microarray data: Between Group Analysis (BGA). This method which is based on performing multivariate ordination of groups proved to be very efficient for both classification of samples into pre-defined groups and disease class prediction of new unknown samples. Classification and prediction with BGA are classically performed using the whole set of genes and no variable selection is required. We hypothesize that an optimized selection of highly discriminating genes might improve the prediction power of BGA. RESULTS: We propose an optimized between-group classification (OBC) which uses a jackknife-based gene selection procedure. OBC emphasizes classification accuracy rather than feature selection. OBC is a backward optimization procedure that maximizes the percentage of between group inertia by removing the least influential genes one by one from the analysis. This selects a subset of highly discriminative genes which optimize disease class prediction. We apply OBC to four datasets and compared it to other classification methods. CONCLUSION: OBC considerably improved the classification and predictive accuracy of BGA, when assessed using independent data sets and leave-one-out cross-validation. AVAILABILITY: The R code is freely available [see Additional file 1] as well as supplementary information [see Additional file 2].

Algorithms↗

Using global optimization to improve classification for medical diagnosis and prognosis.

Global optimization-based techniques are studied in order to increase the accuracy of medical diagnosis and prognosis with data from various databases. First, we discuss feature selection, the problem of determining the most informative features for classification in the databases under consideration. Then, we apply a technique based on convex and global optimization for classification in these databases. The third application of this technique is a method that calculates centers of clusters to predict when breast cancer is likely to recur in patients for which cancer has been removed. The technique achieves high accuracy with these databases. Better classifiers will lead to improved assistance in making medical diagnostic and prognostic decisions.

Algorithms↗

Personal recognition using hand shape and texture.

This paper proposes a new bimodal biometric system using feature-level fusion of hand shape and palm texture. The proposed combination is of significance since both the palmprint and hand-shape images are proposed to be extracted from the single hand image acquired from a digital camera. Several new hand-shape features that can be used to represent the hand shape and improve the performance are investigated. The new approach for palmprint recognition using discrete cosine transform coefficients, which can be directly obtained from the camera hardware, is demonstrated. None of the prior work on hand-shape or palmprint recognition has given any attention on the critical issue of feature selection. Our experimental results demonstrate that while majority of palmprint or hand-shape features are useful in predicting the subjects identity, only a small subset of these features are necessary in practice for building an accurate model for identification. The comparison and combination of proposed features is evaluated on the diverse classification schemes; naive Bayes (normal, estimated, multinomial), decision trees (C4.5, LMT), k-NN, SVM, and FFN. Although more work remains to be done, our results to date indicate that the combination of selected hand-shape and palmprint features constitutes a promising addition to the biometrics-based personal recognition systems.

Algorithms↗

Artificial neural network prediction of retention factors of some benzene derivatives and heterocyclic compounds in micellar electrokinetic chromatography.

A 5-4-1 artificial neural network (ANN) was constructed and trained for prediction of the retention factors of some benzene derivatives and heterocyclic compounds in micellar electrokinetic chromatography (MEKC) based on quantitative structure-property relationship (QSPR). The inputs of this network are theoretically derived descriptors that were chosen by the stepwise variable selection techniques. These descriptors are: molecular surface area, maximum value of electron density on atom in molecule, path four connectivity index, average molecular weight, and sum of atomic polarizability which were selected by using stepwise multiple linear regression as a feature selection technique. The standard errors of training, test, and validation sets for the ANN model are 0.091, 0.119, and 0.114, respectively. Results obtained showed that nonlinear model can simulate the relationship between the structural descriptors and the retention factors of the molecules in data set accurately. Also the appearance of these descriptors in QSPR models reveals the role of electronic and steric interactions in solute retention in MEKC.

Benzene Derivatives↗

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

Computerized analysis of mammographic parenchymal patterns for assessing breast cancer risk: effect of ROI size and location.

The long-term goal of our research is to develop computerized radiographic markers for assessing breast density and parenchymal patterns that may be used together with clinical measures for determining the risk of breast cancer and assessing the response to preventive treatment. In our earlier studies, we found that women at high risk tended to have dense breasts with mammographic patterns that were coarse and low in contrast. With our method, computerized texture analysis is performed on a region of interest (ROI) within the mammographic image. In our current study, we investigate the effect of ROI size and ROI location on the computerized texture features obtained from 90 subjects (30 BRCA1/BRCA2 gene-mutation carriers and 60 age-matched women deemed to be at low risk for breast cancer). Mammograms were digitized at 0.1 mm pixel size and various ROI sizes were extracted from different breast regions in the craniocaudal (CC) view. Seventeen features, which characterize the density and texture of the parenchymal patterns, were extracted from the ROIs on these digitized mammograms. Stepwise feature selection and linear discriminant analysis were applied to identify features that differentiate between the low-risk women and the BRCA1/BRCA2 gene-mutation carriers. ROC analysis was used to assess the performance of the features in the task of distinguishing between these two groups. Our results show that there was a statistically significant decrease in the performance of the computerized texture features, as the ROI location was varied from the central region behind the nipple. However, we failed to show a statistically significant decrease in the performance of the computerized texture features with decreasing ROI size for the range studied.

Algorithms↗

Transformation of stimulus information from very short-term memory.

Observers were shown a simple stimulus pattern, a mask pattern, and a test pattern on each trial. Different types of test patterns were used to assess transformations of material from very short-term memory. The durations of the initial stimulus pattern and the mask pattern were also varied. Significant differences in average response data were found between types of test pattern over exposure durations; however, sensitivity measures showed minimal differences between types of test pattern. This suggested some distinctions between input processing and response selection in the paradigm. Selective feature analysis seems to be a characteristic of response selection.

Discrimination, Psychological↗

Comparison of feature extraction and selection methods in mammogram recognition.

This paper presents a comparison of feature extraction and selection methods in the design of mammogram recognition systems. Mammographic images were classified into two categories, normal and cancerous. The following methods of feature extraction were investigated: two-dimensional Haar wavelets, histograms, and singular value decomposition. The feature patterns were reduced and selected using principal component analysis (PCA) and rough sets. The rough sets methods were applied to the final selection of the pattern features. Classification of mammograms was realized using an error backpropagation neural network.

Breast↗

Discriminative analysis of lip motion features for speaker identification and speech-reading.

There have been several studies that jointly use audio, lip intensity, and lip geometry information for speaker identification and speech-reading applications. This paper proposes using explicit lip motion information, instead of or in addition to lip intensity and/or geometry information, for speaker identification and speech-reading within a unified feature selection and discrimination analysis framework, and addresses two important issues: 1) Is using explicit lip motion information useful, and, 2) if so, what are the best lip motion features for these two applications? The best lip motion features for speaker identification are considered to be those that result in the highest discrimination of individual speakers in a population, whereas for speech-reading, the best features are those providing the highest phoneme/word/phrase recognition rate. Several lip motion feature candidates have been considered including dense motion features within a bounding box about the lip, lip contour motion features, and combination of these with lip shape features. Furthermore, a novel two-stage, spatial, and temporal discrimination analysis is introduced to select the best lip motion features for speaker identification and speech-reading applications. Experimental results using an hidden-Markov-model-based recognition system indicate that using explicit lip motion information provides additional performance gains in both applications, and lip motion features prove more valuable in the case of speech-reading application.

Algorithms↗