PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

A review of caveats in statistical nuclear image analysis.

A large body of the published literature in nuclear image analysis do not evaluate their findings on an independent data set. Hence, if several features are evaluated on a limited data set over-optimistic results are easily achieved. In order to find features that separate different outcome classes of interest, statistical evaluation of the nuclear features must be performed. Furthermore, to classify an unknown sample using image analysis, a classification rule must be designed and evaluated. Unfortunately, statistical evaluation methods used in the literature of nuclear image analysis are often inappropriate. The present article discusses some of the difficulties in statistical evaluation of nuclear image analysis, and a study of cervical cancer is presented in order to illustrate the problems. In conclusion, some of the most severe errors in nuclear image analysis occur in analysis of a large feature set, including few patients, without confirming the results on an independent data set. To select features, Bonferroni correction for multiple test is recommended, together with a standard feature set selection method. Furthermore, we consider that the minimum requirement of performing statistical evaluation in nuclear image analysis is confirmation of the results on an independent data set. We suggest that a consensus of how to perform evaluation of diagnostic and prognostic features is necessary, in order to develop reliable tools for clinical use, based on nuclear image analysis.

Data Interpretation, Statistical↗

Multiple neural network classification scheme for detection of colonic polyps in CT colonography data sets.

RATIONALE AND OBJECTIVES: A new classification system for colonic polyp detection, designed to increase sensitivity and reduce the number of false-positive findings with computed tomographic colonography, was developed and tested in this study. MATERIALS AND METHODS: The system involves classification by a committee of neural networks (NNs), each using largely distinct subsets of features selected from a general set. Back-propagation NNs trained with the Levenberg-Marquardt algorithm were used as primary classifiers (committee members). The set of features included region density, Gaussian and mean curvature and sphericity, lesion size, colon wall thickness, and the means and standard deviations of all of these values. Subsets of variables were initially selected because of their effectiveness according to training and test sample misclassification rates. The final decision for each case is based on the majority vote across the networks and reflects the weighted votes of all networks. The authors also introduce a smoothed cross-validation method designed to improve estimation of the true misclassification rates by reducing bias and variance. RESULTS: This committee method reduced the false-positive rate by 36%, a clinically meaningful reduction, and improved sensitivity by an average of 6.9% compared with decisions made by any single NN. The overall sensitivity and specificity were 82.9% and 95.3%, respectively, when sensitivity was estimated by means of smoothed cross-validation. CONCLUSION: The proposed method of using multiple classifiers and majority voting is recommended for classification tasks with large sets of input features, particularly when selected feature subsets may not be equally effective and do not provide satisfactory true- and false-positive rates. This approach reduces variance in estimates of misclassification rates.

Algorithms↗

Is contrast just another feature for visual selective attention?

The biased-competition theory of attention [Annual Review of Neuroscience 18 (1995) 193] suggests that attention and stimulus contrast trade off, and implies that high-contrast stimuli should be easy to attend to and hard to ignore. To test this, observers searched displays for a target digit. Observers were well able to exclude high-contrast distractors when attempting to search only among low-contrast stimuli (Experiment 1). In Experiments 2 and 3, location determined which stimuli were relevant. When contrast of relevant and irrelevant stimuli was uncertain (due to contrast varying between trials, Experiment 2), increasing the contrast of distractors impaired performance. However, when contrast was certain (due to blocking of trials, Experiment 3) and targets were of low contrast, high contrast distractors produced less interference than low contrast distractors. The ability of subjects to attend selectively to low vs. high contrast items in Experiments 1 and 3 suggests that selectivity for stimulus contrast might be similar to other types of feature selectivity (e.g., color and location). Such findings are inconsistent with the biased competition theory regarding the interplay of contrast and attention. However, results from Experiment 2 suggest that, when target contrast varies, the default tendency is to attend to high-contrast items.

Adaptation, Physiological↗

Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction.

MOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. The aim of this paper is to systematically benchmark the role of non-linear versus linear techniques and dimensionality reduction methods. RESULTS: A systematic benchmarking study is performed by comparing linear versions of standard classification and dimensionality reduction techniques with their non-linear versions based on non-linear kernel functions with a radial basis function (RBF) kernel. A total of 9 binary cancer classification problems, derived from 7 publicly available microarray datasets, and 20 randomizations of each problem are examined. CONCLUSIONS: Three main conclusions can be formulated based on the performances on independent test sets. (1) When performing classification with least squares support vector machines (LS-SVMs) (without dimensionality reduction), RBF kernels can be used without risking too much overfitting. The results obtained with well-tuned RBF kernels are never worse and sometimes even statistically significantly better compared to results obtained with a linear kernel in terms of test set receiver operating characteristic and test set accuracy performances. (2) Even for classification with linear classifiers like LS-SVM with linear kernel, using regularization is very important. (3) When performing kernel principal component analysis (kernel PCA) before classification, using an RBF kernel for kernel PCA tends to result in overfitting, especially when using supervised feature selection. It has been observed that an optimal selection of a large number of features is often an indication for overfitting. Kernel PCA with linear kernel gives better results.

Algorithms↗

A genetic algorithm based nearest neighbor classification to breast cancer diagnosis.

This paper presents an application of a hybrid approach (the genetic algorithms and the k-nearest neighbour) proposed by Ishbuchi to Wisconsin breast cancer data. For the diagnosis of breast cancer, the determination of the presence of benign/malignant breast tumors represents a very complex problem (even for an experienced cytologist). Therefore the automatic classification of benign and malignant symptoms is highly desirable as a valuable aid to assist oncologists in the decision making of the diagnosis of breast cancer. In this paper, the genetic algorithm based k-nearest neighbour method for classification of benign and malignant breast tumors is presented. The genetic-algorithm (GA) is used for finding a compact reference set by selecting a small number of reference patterns from a large number of training patterns in nearest neighbor classification. The GA simultaneously performs feature selection and pattern selection and prunes unnecessary features. The goal is to maximize the classification performance of the reference set and minimize the number of selected patterns and features. Results are also compared with a fuzzy-genetic approach where each reference patten represents a fuzzy if-then rule with a circular-cone-type membership function.

Algorithms↗

BioViews: Java-based tools for genomic data visualization.

Visualization tools for bioinformatics ideally should provide universal access to the most current data in an interactive and intuitive graphical user interface. Since the introduction of Java, a language designed for distributed programming over the Web, the technology now exists to build a genomic data visualization tool that meets these requirements. Using Java we have developed a prototype genome browser applet (BioViews) that incorporates a three-level graphical view of genomic data: a physical map, an annotated sequence map, and a DNA sequence display. Annotated biological features are displayed on the physical and sequence-based maps, and the different views are interconnected. The applet is linked to several databases and can retrieve features and display hyperlinked textual data on selected features. In addition to browsing genomic data, different types of analyses can be performed interactively and the results of these analyses visualized alongside prior annotations. Our genome browser is built on top of extensible, reusable graphic components specifically designed for bioinformatics. Other groups can (and do) reuse this work in various ways. Genome centers can reuse large parts of the genome browser with minor modifications, bioinformatics groups working on sequence analysis can reuse components to build front ends for analysis programs, and biology laboratories can reuse components to publish results as dynamic Web documents.

Animals↗

Three-dimensional computer modelling system for the study of biological structures.

A three-dimensional computer modelling system has been developed for use in biology, and is currently running on a Sun3 computer. The data originate as a series of two-dimensional micrographs which are digitised via a TV camera. The two-dimensional images are used to select features of interest and to construct a three-dimensional model. This model can be viewed in vector or solid format, it can be rotated about three orthogonal axes and can be viewed in three dimensions as a stereo pair or an anaglyph. The system has been used in a large number of projects over the past 10-15 years, for example, to examine physiological and nerve structures. The time-consuming part of the process is the selection of features, which involves a high level of biological expertise. Present developments are concerned with reduction of the time spent in feature recognition and involve the introduction of expert systems together with human-computer interaction to deal with problems of identification.

Computer Simulation↗

Predicting radiotherapy-induced cardiac perfusion defects.

The purpose of this work is to compare the efficacy of mathematical models in predicting the occurrence of radiotherapy-induced left ventricular perfusion defects assessed using single-photon emission computed tomography (SPECT). The basis of this study is data from 73 left-sided breast/ chestwall patients treated with tangential photon fields. The mathematical models compared were three commonly used parametric models [Lyman normal tissue complication probability (LNTCP), relative serialty (RS), generalized equivalent uniform dose (gEUD)] and a nonparametric model (Linear discriminant analysis--LDA). Data used by the models were the left ventricular dose--volume histograms, or SPECT-based dose-function histograms, and the presence/absence of SPECT perfusion defects 6 months postradiation therapy (21 patients developed defects). For the parametric models, maximum likelihood estimation and F-tests were used to fit the model parameters. The nonparametric LDA model step-wise selected features (volumes/function above dose levels) using a method based on receiver operating characteristics (ROC) analysis to best separate the groups with and without defects. Optimistic (upper bound) and pessimistic (lower bound) estimates of each model's predictive capability were generated using ROC curves. A higher area under the ROC curve indicates a more accurate model (a model that is always accurate has area = 1). The areas under these curves for different models were used to statistically test for differences between them. Pessimistic estimates of areas under the ROC curve using dose-volume histogram/ dose-function histogram inputs, in order of increasing prediction accuracy, were LNTCP (0.79/0.75), RS (0.80/0.77), gEUD (0.81/0.78), and LDA (0.84/0.86). Only the LDA model benefited from SPECT-based regional functional information. In general, the LDA model was statistically superior to the parametric models. The LDA model selected as features the left ventricular volumes above approximately 23 Gy (V23), essentially volume in field, and 33 Gy (V33), as best separating the groups with and without defects. In conclusion, the nonparametric LDA model appears to be a more accurate predictor of radiotherapy-induced left ventricular perfusion defects than commonly used parametric models.

Dose-Response Relationship, Radiation↗

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross‐taxon transferability↗

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans↗

Using global optimization to improve classification for medical diagnosis and prognosis.

Global optimization-based techniques are studied in order to increase the accuracy of medical diagnosis and prognosis with data from various databases. First, we discuss feature selection, the problem of determining the most informative features for classification in the databases under consideration. Then, we apply a technique based on convex and global optimization for classification in these databases. The third application of this technique is a method that calculates centers of clusters to predict when breast cancer is likely to recur in patients for which cancer has been removed. The technique achieves high accuracy with these databases. Better classifiers will lead to improved assistance in making medical diagnostic and prognostic decisions.

Algorithms↗

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

Computerized analysis of mammographic parenchymal patterns for assessing breast cancer risk: effect of ROI size and location.

The long-term goal of our research is to develop computerized radiographic markers for assessing breast density and parenchymal patterns that may be used together with clinical measures for determining the risk of breast cancer and assessing the response to preventive treatment. In our earlier studies, we found that women at high risk tended to have dense breasts with mammographic patterns that were coarse and low in contrast. With our method, computerized texture analysis is performed on a region of interest (ROI) within the mammographic image. In our current study, we investigate the effect of ROI size and ROI location on the computerized texture features obtained from 90 subjects (30 BRCA1/BRCA2 gene-mutation carriers and 60 age-matched women deemed to be at low risk for breast cancer). Mammograms were digitized at 0.1 mm pixel size and various ROI sizes were extracted from different breast regions in the craniocaudal (CC) view. Seventeen features, which characterize the density and texture of the parenchymal patterns, were extracted from the ROIs on these digitized mammograms. Stepwise feature selection and linear discriminant analysis were applied to identify features that differentiate between the low-risk women and the BRCA1/BRCA2 gene-mutation carriers. ROC analysis was used to assess the performance of the features in the task of distinguishing between these two groups. Our results show that there was a statistically significant decrease in the performance of the computerized texture features, as the ROI location was varied from the central region behind the nipple. However, we failed to show a statistically significant decrease in the performance of the computerized texture features with decreasing ROI size for the range studied.

Algorithms↗

Transformation of stimulus information from very short-term memory.

Observers were shown a simple stimulus pattern, a mask pattern, and a test pattern on each trial. Different types of test patterns were used to assess transformations of material from very short-term memory. The durations of the initial stimulus pattern and the mask pattern were also varied. Significant differences in average response data were found between types of test pattern over exposure durations; however, sensitivity measures showed minimal differences between types of test pattern. This suggested some distinctions between input processing and response selection in the paradigm. Selective feature analysis seems to be a characteristic of response selection.

Discrimination, Psychological↗

Comparison of feature extraction and selection methods in mammogram recognition.

This paper presents a comparison of feature extraction and selection methods in the design of mammogram recognition systems. Mammographic images were classified into two categories, normal and cancerous. The following methods of feature extraction were investigated: two-dimensional Haar wavelets, histograms, and singular value decomposition. The feature patterns were reduced and selected using principal component analysis (PCA) and rough sets. The rough sets methods were applied to the final selection of the pattern features. Classification of mammograms was realized using an error backpropagation neural network.

Breast↗

Deciphering the swordtail's tale: a molecular and evolutionary quest.

The power of sexual selection to influence the evolution of morphological traits was first proposed more than 130 years ago by Darwin. Though long a controversial idea, it has been documented in recent decades for a host of animal species. Yet few of the established sexually selected features have been explored at the level of their genetic or molecular foundations. In a recent report, Zauner et al.1 describe some of the molecular features associated with one of the best characterized of sexually selected traits, the male-specific tail "sword" seen in certain species of the fish genus Xiphophorus. Zauner et al. find that the msxC gene, a gene previously implicated in fin development from work in zebrafish, is dramatically and specifically upregulated in the development of the ventral caudal fin rays, which give rise to the sword, in males. The results provide the first molecular insight into the development of this sexually selected trait while prompting new questions about the structure of the entire genetic network that underlies this trait. To fully understand the molecular-genetic and evolutionary history of this network, however, it will be essential to determine whether sword-development is a basal or derived trait in Xiphophorus.

Animals↗

Oscillatory responses in cat visual cortex exhibit inter-columnar synchronization which reflects global stimulus properties.

A fundamental step in visual pattern recognition is the establishment of relations between spatially separate features. Recently, we have shown that neurons in the cat visual cortex have oscillatory responses in the range 40-60 Hz (refs 1, 2) which occur in synchrony for cells in a functional column and are tightly correlated with a local oscillatory field potential. This led us to hypothesize that the synchronization of oscillatory responses of spatially distributed, feature selective cells might be a way to establish relations between features in different parts of the visual field. In support of this hypothesis, we demonstrate here that neurons in spatially separate columns can synchronize their oscillatory responses. The synchronization has, on average, no phase difference, depends on the spatial separation and the orientation preference of the cells and is influenced by global stimulus properties.

Animals↗

Epstein-Barr virus-negative post-transplant lymphoproliferative disorders: a distinct entity?

Post-transplant lymphoproliferative disorders (PTLDs) are usually but not invariably associated with Epstein-Barr virus (EBV). The reported incidence, however, of EBV-negative PTLDs varies widely, and it is uncertain whether they should be considered analogous to EBV-positive PTLDs and whether they have any distinctive features. Therefore, the EBV status of 133 PTLDs from 80 patients was determined using EBV-encoded small ribonucleic acid (EBER) in situ hybridization stains with or without Southern blot EBV terminal repeat analysis. The morphologic, immunophenotypic, genotypic, and clinical features of the EBV-negative PTLDs were reviewed, and selected features were compared with EBV-positive cases. Twenty-one percent of patients had at least one EBV-negative PTLD (14% of biopsies). The initial EBV-negative PTLDs occurred a median of 50 months post-transplantation compared with 10 months for EBV-positive cases. Although only 2% of PTLDs from before 1991 were EBV negative, 23% of subsequent PTLDs were EBV negative (p <0.001). Of the EBV-negative PTLDs, 67% were of monomorphic type (M-PTLD) compared with 42% of EBV-positive cases (p <0.05). The other EBV-negative PTLDs were of infectious mononucleosis-like, plasma cell-rich (n = 2), small B-cell lymphoid neoplasm, large granular lymphocyte disorder (n = 4) and polymorphic (P) types. B-cell clonality was established in 14 specimens and T-cell clonality was established in three (two patients). None of the remaining specimens were studied with Southern blot analysis and some had no ancillary studies. Rearrangement of c-MYC was identified in two M-PTLDs with small noncleaved-like features, and rearrangement of BCL-2 was found in one large noncleaved-like M-PTLD. Ten patients were alive at 3 to 63 months (only three patients received chemotherapy). Seven patients, all with M-PTLDs, are dead at 0.3 to 6 months. Therefore, EBV-negative PTLDs have distinct features, but some do respond to decreased immunosuppression, similar to EBV-positive cases, suggesting that EBV positivity should not be an absolute criterion for the diagnosis of a PTLD.

Adult↗