PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Fine needle aspiration biopsy diagnosis of mucoepidermoid carcinoma. Statistical analysis.

Fine needle aspiration (FNA) biopsy is an increasingly popular method for the evaluation of salivary gland tumors. Of the common salivary gland tumors, mucoepidermoid carcinoma is probably the most difficult to diagnose accurately by this means. A series of 96 FNA biopsy specimens of salivary gland masses, including 34 mucoepidermoid carcinomas, 51 other benign and malignant neoplasms, 7 nonneoplastic lesions and 4 normal salivary glands, were analyzed in order to identify the most useful criteria for diagnosing mucoepidermoid carcinoma. Thirteen cytologic criteria were evaluated in the FNA specimens, and a stepwise logistic regression analysis was performed. The three cytologic features selected as most predictive of mucoepidermoid carcinoma were intermediate cells, squamous cells and overlapping epithelial groups. Using these three features together, the sensitivity and specificity of accurately diagnosing mucoepidermoid carcinoma were 97% and 100%, respectively.

Biopsy, Needle↗

Use of logistic regression analysis to improve prediction of prognosis in acute myeloid leukaemia.

The prognostic usefulness of a range of factors has been examined for patients with acute myeloid leukaemia. Although there was a statistical association between some of these factors and remission rate, the association was only partial. To improve the usefulness of the data, multiple logistic regressional analysis was used. The features selected for use in the analysis were age, blood blast count, FAB classification and colony growth pattern. The last three features could be used as categorical variables, since blood blast counts of greater than 100 X 10(9)/1, FAB group 1 and a prolific pattern of colony growth were associated with a low remission rate. Age was used as a continuous variable. Using these features, eight regression groups were defined. Thus when this data for an individual patient is analysed, it is possible to obtain a value for the probability of that patient achieving remission.

Adult↗

[Reproducibility of radiologic diagnosis in gonarthrosis].

AIM OF STUDY: Ongoing efforts of the "German Society of Orthopaedic Surgery and Traumatology" (DGOT) to standardize diagnosis and therapy of osteoarthritis, necessitated this study, where the reproducibility of different radiographic features of knee-OA was assessed. METHODS: Three readers graded 100 antero-posterior and lateral knee radiographs for selected features (femorotibial osteophytes, joint space narrowing, sclerosis and chondrocalcinosis; patellofemoral osteophytes) and an overall-score (Kellgren and Lawrence, 1963) at two time points 3 months apart. Intra- and inter-observer-reliability were calculated by intra-class correlation-coefficient (ICC). RESULTS: Osteophytes in the femorotibial as well as in the patellofemoral joint could be assessed with a high intra- and inter-observer-reliability. While for joint space narrowing intra-observer-reliability is excellent, the inter-observer-reliability is less satisfactory, and subchondral sclerosis as well as chondrocalcinosis showed even less reproducibility. The Kellgren-Lawrence global score proofed to be highly reproducible. CONCLUSION: Based on the results of this examination, we can recommend a reliable radiographic classification of knee osteoarthritis by grading of relevant individual features (osteophytes and joint space narrowing) and overall assessment.

Follow-Up Studies↗

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem↗

Cell culture modeling of specialized tissue: identification of genes expressed specifically by follicle-associated epithelium of Peyer's patch by expression profiling of Caco-2/Raji co-cultures.

Peyer's patch follicle-associated epithelium (FAE) regulates intestinal antigen access to the immune system in part through the action of microfold (M) cells which mediate transcytosis of antigens and microorganisms. Studies on M cells have been limited by the difficulties in isolating purified cells, so we applied TOGA mRNA expression profiling to identify genes associated with the in vitro induction of M cell-like features in Caco-2 cells and tested them against normal Peyer's patch tissue for their expression in FAE. Among the genes identified by this method, laminin beta3, a matrix metalloproteinase and a tetraspan family member, showed enriched expression in FAE of mouse Peyer's patches. Moreover, the C. perfringens enterotoxin receptor (CPE-R) appeared to be expressed more strongly by UEA-1(+) M cells relative to neighboring FAE. Expression of the tetraspan TM4SF3 gene and CPE-R was also confirmed in human Peyer's patch FAE. Our results suggest that while the Caco-2 differentiation model is associated with some functional features of M cells, the genes induced may instead reflect the acquisition of a more general FAE phenotype, sharing only select features with the M cell subset.

Animals↗

Molecular conformational space analysis using computer graphics: going beyond FRODO.

The molecular graphics program FRODO has been modified to support analytical animation of molecular dynamics trajectories. The enhanced program, mdFRODO, supports all features available in FRODO and is interfaced to GROMOS. A variety of analytical animation modes is included. Extensive coloring and atom selection features are implemented to aid the user in distinguishing features of interest in a set of conformations. Molecular conformational space can be analyzed efficiently and comprehended. Animations may be viewed in stereo, and the animated object can be overlaid with any of the standard FRODO objects. The mdFRODO program is of wide use in molecular dynamics, X-ray crystallography and two-dimensional NMR work. Examples illustrating various aspects of collective motion in protein molecules are given and discussed.

Computer Graphics↗

The development of perceived structure and attention: evidence from divided and selective attention tasks.

Three experiments provide converging evidence for the view that both perceived structure and attention change during the elementary school years. Kindergarteners, second graders, and adults performed three speeded tasks: divided attention to conjunctions of features, selective attention to orthogonal dimensions and selective attention to correlated dimensions. The tasks were performed with sizes and shapes that were either spatially integrated or spatially separated. In the divided attention task, conjunctions were identified as quickly as single features with integrated stimuli at all ages, but conjunctions were identified more slowly than single features with separated stimuli by all age group. In the orthogonal dimensions task, interference was observed with integrated stimuli across ages, but the interference in adult performance was asymmetric. With separated stimuli, interference was gradually eliminated with increasing age. In correlated dimensions tasks, younger children showed a redundancy gain with integrated stimuli, but no gain was observed in the performances of the older subjects. With separated stimuli there was no redundancy gain at any age. These results were interpreted to mean that integrated stimuli are initially perceived as wholes by all subjects, but that features become more accessible with increasing age. Even so, attention remains constrained by stimulus structure. In contrast, separated stimuli are initially perceived as features at all ages, and the improvement in performance with increasing age is attributable to the increasing command of attentional resources that accompanies development. Our discussion of these findings focuses on three issues: multiple trends in perceptual development, the characteristics of an adequate theory of perceptual representation and processing, and a comparison of the separability hypothesis and other developmental accounts of perceptual development.

Adult↗

Methodological aspects of using decision trees to characterise leiomyomatous tumors.

The aim of the present work is to present the potential uses of a classification technique labeled the "decision tree" for tumor characterisation when faced with a large number of features. The decision tree technique enables multifeature logical classification rules to be produced by determining discriminatory values for each feature selected. In this report, we propose a methodology that used decision trees to compare and evaluate the information contributed by different types of features for tumor characterisation. This methodology is able to produce a set of hypotheses related to a diagnosis and or prognosis problem. For example, hypotheses can be producted (on the basis of a set of descriptive features) to explain why tumor cases belong to a given histopathological group. To illustrate our purpose, this methodology was applied to the difficult problem of leiomyomatous tumour diagnosis. The aim was to illustrate what kind of diagnostic information can be extracted from a sample data set including 23 smooth muscle tumors (14 benign leiomyomas and 9 malignant leiomyosarcomas) described by a large set of computer-assisted, microscope-generated features. Three groups of features were used relating to: (1) ploidy level determination (10 features), (2) quantitative chromatin pattern description (15 features), and (3) immunohistochemically related antigen specificities (6 features). All these features were quantified by digital cell image analysis. The results suggest that an objective distinction between leiomyomas and leiomyosarcomas can be established by means of simple logical rules depending on only a few features among which the immunohistochemically revealed antigen expression of desmin plays a preponderant part. One of the combinations of features proposed by the methodology is interesting for pathologists, because it includes two features describing the appearance of a nucleus in terms of chromatin distribution homogeneity and density, two features widely used by pathologists in tumor-grading systems.

Adolescent↗

Automated computerized classification of malignant and benign masses on digitized mammograms.

RATIONALE AND OBJECTIVES: To develop a method for differentiating malignant from benign masses in which a computer automatically extracts lesion features and merges them into an estimated likelihood of malignancy. MATERIALS AND METHODS: Ninety-five mammograms depicting masses in 65 patients were digitized. Various features related to the margin and density of each mass were extracted automatically from the neighborhoods of the computer-identified mass regions. Selected features were merged into an estimated likelihood of malignancy by, using three different automated classifiers. The performance of the three classifiers in distinguishing between benign and malignant masses was evaluated by receiver operating characteristic analysis and compared with the performance of an experienced mammographer and that of five less experienced mammographers. RESULTS: Our computer classification scheme yielded an area under the receiver operating characteristic curve (Az) value of 0.94, which was similar to that for an experienced mammographer (Az = 0.91) and was statistically significantly higher than the average performance of the radiologists with less mammographic experience (Az = 0.81) (P = .013). With the database used, the computer scheme achieved, at 100% sensitivity, a positive predictive value of 83%, which was 12% higher than that for the performance of the experienced mammographer and 21% higher than that for the average performance of the less experienced mammographers (P < .0001). CONCLUSION: Automated computerized classification schemes may be useful in helping radiologists distinguish between benign and malignant masses and thus reducing the number of unnecessary biopsies.

Breast Neoplasms↗

Quantitative structure-activity relationship studies of progesterone receptor binding steroids.

The selection of appropriate descriptors is an important step in the successful formulation of quantitative structure-activity relationships (QSARs). This paper compares a number of feature selection routines and mapping methods that are in current use. They include forward stepping regression (FSR), genetic function approximation (GFA), generalized simulated annealing (GSA), and genetic neural network (GNN). On the basis of a data set of steroids of known in vitro binding affinity to the progsterone receptor, a number of QSAR models are constructed. A comparison of the predictive qualities for both training and test compounds demonstrates that the GNN protocol achieves the best results among the 2D QSAR that are considered. Analysis of the choice of descriptors by the GNN method shows that the results are consistent with established SARs on this series of compounds.

Neural Networks, Computer↗

Plasma free fatty acid levels and the risk of ischemic heart disease in men: prospective results from the Québec Cardiovascular Study.

Insulin resistance, through numerous related disturbances in glucose and lipoprotein-lipid metabolism, is associated with an increased risk of ischemic heart disease (IHD). The purpose of the present study was to examine the relationship between increased plasma free fatty acid (FFA) concentrations, as a feature of the insulin resistance syndrome, and the risk of IHD in men. Analyses were carried out in a nested, case-control sample of men selected from a population of 2103 individuals without IHD at baseline among whom 114 developed IHD during a 5-year follow-up period. Incident IHD cases were matched with controls for age, body mass index, smoking habits and alcohol intake. Analyses were performed while excluding (88 cases and 98 controls) and including (103 cases and 99 controls) patients with type 2 diabetes. Among non-diabetic individuals, elevated plasma FFA concentrations (3rd tertile of the distribution) yielded a twofold increase in the risk of IHD (odds ratio [OR] 2.1, P=0.05) compared with lower plasma FFA levels (lowest tertile) after adjusting for non-lipid risk factors. Further adjustment for insulin, triglycerides, apolipoprotein B, HDL cholesterol and small dense LDL attenuated significantly the relationship between plasma FFA concentrations and the risk of IHD. High plasma FFA levels showed no synergism with selected features of the insulin resistance syndrome in determining the risk of IHD. Inclusion of diabetic subjects in the study did not improve FFA independent prognostic value to the risk of IHD. These results suggest that elevated plasma FFA concentrations are associated with an increased risk of IHD. However, a single fasting measurement of plasma FFA levels does not appear to improve our ability to predict IHD onset in men when information on other risk factors is considered.

Adult↗

Oriented principal component analysis for large margin classifiers.

Large margin classifiers (such as MLPs) are designed to assign training samples with high confidence (or margin) to one of the classes. Recent theoretical results of these systems show why the use of regularisation terms and feature extractor techniques can enhance their generalisation properties. Since the optimal subset of features selected depends on the classification problem, but also on the particular classifier with which they are used, global learning algorithms for large margin classifiers that use feature extractor techniques are desired. A direct approach is to optimise a cost function based on the margin error, which also incorporates regularisation terms for controlling capacity. These terms must penalise a classifier with the largest margin for the problem at hand. Our work shows that the inclusion of a PCA term can be employed for this purpose. Since PCA only achieves an optimal discriminatory projection for some particular distribution of data, the margin of the classifier can then be effectively controlled. We also propose a simple constrained search for the global algorithm in which the feature extractor and the classifier are trained separately. This allows a degree of flexibility for including heuristics that can enhance the search and the performance of the computed solution. Experimental results demonstrate the potential of the proposed method.

Algorithms↗

Selection of electrode positions for an EEG-based brain computer interface (BCI).

One major question in designing an EEG-based Brain Computer Interface to bypass the normal motor pathways is the selection of proper electrode positions. This study investigates electrode selection with a Distinction Sensitive Learning Vector Quantizer (DSLVQ). DSLVQ is an extended Learning Vector Quantizer (LVQ) which employs a weighted distance function for dynamical scaling and feature selection. The data analysed and classified were 56-channel EEG recordings over sensorimotor areas during preparation for discrete left or right index finger flexions. Data from 3 subjects are reported. It was found by DSLVQ that the most important electrode positions for differentiation between planning of left and right finger movement overlie cortical finger/hand areas over both hemispheres.

Adult↗

CLIFF: clustering of high-dimensional microarray data via iterative feature filtering using normalized cuts.

We present CLIFF, an algorithm for clustering biological samples using gene expression microarray data. This clustering problem is difficult for several reasons, in particular the sparsity of the data, the high dimensionality of the feature (gene) space, and the fact that many features are irrelevant or redundant. Our algorithm iterates between two computational processes, feature filtering and clustering. Given a reference partition that approximates the correct clustering of the samples, our feature filtering procedure ranks the features according to their intrinsic discriminability, relevance to the reference partition, and irredundancy to other relevant features, and uses this ranking to select the features to be used in the following round of clustering. Our clustering algorithm, which is based on the concept of a normalized cut, clusters the samples into a new reference partition on the basis of the selected features. On a well-studied problem involving 72 leukemia samples and 7130 genes, we demonstrate that CLIFF outperforms standard clustering approaches that do not consider the feature selection issue, and produces a result that is very close to the original expert labeling of the sample set.

Algorithms↗

An intelligent framework for the classification of the 12-lead ECG.

An intelligent framework has been proposed to classify an unknown 12-Lead electrocardiogram into one of a possible number of mutually exclusive and combined diagnostic classes. The framework segregates the classification problem into a number of bi-dimensional classification problems, requiring individual bi-group classifiers for each individual diagnostic class. The bi-group classifiers were generated employing Neural Networks (NN), combined with a combination framework containing an Evidential Reasoning framework to accommodate for any conflicting situations between the bi-group classifiers. A number of different feature selection techniques were investigated with the aim of generating the most appropriate input vector for the bi-group classifiers. It was found that by reducing the original input feature vector, the generalisation ability of the classifiers, when exposed to unseen data, was enhanced and subsequently this reduced the computational requirements of the network itself. The entire framework was compared with a conventional approach to NN classification and a rule based classification approach. The framework attained a significantly higher level of classification in comparison with the other methods; 80.0% compared with 66.7% for the rule based technique and 68.00% for the conventional neural approach.

Computer Simulation↗

Graph neural network-based risk stratification of prostate cancer using gene expression and SHAP interpretability.

Accurate risk stratification is essential for guiding treatment decisions and preventing over treatment of prostate cancer, which remains one of the most prevalent cancers among adult men. While the Gleason score, obtained from prostate biopsies, is routinely used to assess tumor aggressiveness, the biopsy procedure carries risks such as pain, infection, and, in some cases, serious complications such as sepsis. In this study, we proposed an artificial intelligence-based framework that integrates mRNA expression profiles with functional interaction networks to classify prostate cancer patients into low-, medium-, and high-risk groups defined by Gleason scores. The pipeline comprised five steps: (1) data collection from The Cancer Genome Atlas (TCGA), (2) preprocessing of gene expression data, (3) two-stage feature selection to identify informative biomarkers, (4) risk classification using a dual-branch graph neural network (GNN) that combines gene-gene interaction graphs with sample-level expression features, and (5) model interpretation using SHAP to quantify feature contributions. Differentially expressed genes were identified in the High (ASPN, GMNN, PEBP4, C2, KNCK17), Medium (C2, IGSF1, ASPN, CDKN3, AMH), and Low (TNMD, VWA5B2, ST6GALNAC5, CYP3A5, PHGR1) risk groups, underscoring the molecular heterogeneity of disease progression. On an independent held-out test set, the model achieved AUCs of 0.86, 0.88, and 0.95 for the low-, medium-, and high-risk groups, respectively, with an overall accuracy of 80%. These results suggest that combining GNN-based modeling with explainable AI can capture both global and local molecular patterns relevant to tumor aggressiveness. However, as the model was developed and evaluated solely on the TCGA cohort, the findings should be regarded as exploratory, and external validation will be required to establish generalizability. Within these limitations, the proposed framework highlights the potential of molecular profiling and graph-based deep learning to support more precise, potentially less invasive, risk assessment and individualized treatment planning in prostate cancer.

Prostatic Neoplasms↗

Relevant EEG features for the classification of spontaneous motor-related tasks.

There is a growing interest in the use of physiological signals for communication and operation of devices for the severely motor disabled as well as for healthy people. A few groups around the world have developed brain-computer interfaces (BCIs) that rely upon the recognition of motor-related tasks (i.e., imagination of movements) from on-line EEG signals. In this paper we seek to find and analyze the set of relevant EEG features that best differentiate spontaneous motor-related mental tasks from each other. This study empirically demonstrates the benefits of heuristic feature selection methods for EEG-based classification of mental tasks. In particular, it is shown that the classifier performance improves for all the considered subjects with only a small proportion of features. Thus, the use of just those relevant features increases the efficiency of the brain interfaces and, most importantly, enables a greater level of adaptation of the personal BCI to the individual user.

Brain Mapping↗

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95&#xa0;% CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma↗