PubMed HealthSearch

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Uncovering heterogeneous effects via localized feature selection.

Identifying features that interact to trigger disease, while accounting for heterogeneity across diverse populations, is essential for the development of precision and targeted medicine. Despite the availability of vast and complex health-related datasets, most existing works focus on identifying disease-associated features at the population level or within a few subpopulations, often overlooking individual-level heterogeneity within these groups. To address this limitation, we propose a framework that utilizes localized test statistics to identify disease-associated features tailored to individual profiles. Our method leverages the recently developed knockoffs methodology to control the noise level of the selection set so that the results are replicable. Moreover, it allows for the discovery of hidden heterogeneous effects within the data, as demonstrated in an application to single-cell RNA sequencing data for Alzheimer's disease. By aggregating localized feature selection results, our framework also enables powerful population-level feature selection. Our framework provides a powerful tool for exploratory studies of precision medicine, offering the potential to generate novel hypotheses for confirmatory biological experiments.

Alzheimer Disease

A novel and robust feature selection method with FDR control for omics-wide association analysis.

Omics-wide association analysis is a very important tool for medicine and human health study. However, the modern omics data sets collected often exhibit the high-dimensionality, unknown distribution response, unknown distribution features and unknown complex association relationships between the response and its explanatory features. Reliable association analysis results depend on an accurate modeling for such data sets. Most of the existing association analysis methods rely on the specific model assumptions and lack effective false discovery rate (FDR) control. To address these limitations, the paper firstly applies a single index model for omics data. The model shows robust performance in allowing the relationships between the response variable and linear combination of covariates to be connected by any unknown monotonic link function, and both the random error and the covariates can follow any unknown distribution. Then based on this model, the paper combines rank-based approach and symmetrized data aggregation approach to develop a novel and robust feature selection method for achieving fine-mapping of risk features while controlling the false positive rate of selection. The theoretical results support the proposed method and the analysis results of simulated data show the new method possesses effective and robust performance for all the scenarios. The new method is also used to analyze the two real datasets and identifies some risk features unreported by the existing finds.

Humans

engGNN: a dual-graph neural network for omics-based disease classification and feature selection.

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph neural networks offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated graph, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive representations, thereby improving predictive performance and interpretability. Through extensive simulation studies and real-world applications to three independent gene expression datasets, engGNN consistently demonstrates strong classification performance compared with competitive baselines. Beyond classification, engGNN provides feature- and source-level interpretability, enabling biologically meaningful analyses such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.

Graph Neural Networks

Elicited imitation of selected features of two american English dialects in Head Start children.

Three measures were used to check the bidialectal imitative facility of 100 black, white, and Spanish-speaking Head Start children. In general, blacks and Spanish-speaking subjects performed more accurately on black English markers than on Standard English markers and whites, the reverse. When the children did make an error on the feature marker they usually substituted the opposing dialectal marker. Blacks and Spanish-speaking subjects were more apt to be accurate on the total sentence when it was given in black English. Several explanations are offered for group similarities and differences.

Black or African American

BISON: bi-clustering of spatial omics data with feature selection.

MOTIVATION: The advent of next-generation sequencing-based spatially resolved transcriptomics (SRT) techniques has reshaped genomic studies by enabling high-throughput gene expression profiling while preserving spatial and morphological context. Understanding gene functions and interactions in different spatial domains is crucial, as it can enhance our comprehension of biological mechanisms, such as cancer-immune interactions and cell differentiation in various regions. It is necessary to cluster tissue regions into distinct spatial domains and identify discriminating genes (DGs) that elucidate the clustering result, referred to as spatial domain-specific DGs. Existing methods for identifying these genes typically rely on a two-stage approach, which can lead to the phenomenon known as double-dipping. RESULTS: To address the challenge, we propose a unified Bayesian latent block model that simultaneously detects a list of DGs contributing to spatial domain identification while clustering these DGs and spatial locations. The efficacy of our proposed method is validated through a series of simulation experiments, and its capability to identify DGs is demonstrated through applications to benchmark SRT datasets. AVAILABILITY AND IMPLEMENTATION: The R/C++ implementation of BISON is available at https://github.com/new-zbc/BISON.

Software

Shedd's formulations concerning the hyperkinetic syndrome--an empirical test of selected features.

According to Shedd's (1968) formulations, one of the distinguishing characteristics of hyperkinetic children is the score pattern which they earn on different types of ability tests. Specifically, Shedd's work suggests that IQs from a picture vocabulary test, the WISC, and a drawing test should show the following order relationship: picture vocabulary greater than WISC greater than drawing. This hypothesis was investigated for 62 overactive children (47 boys and 15 girls) enrolled in classes for learning disabled. The findings support Shedd's since IQs from the Ammons' test were significantly higher than WISC IQs which were significantly higher than IQs from the Goodenough-Harris. The magnitudes of the mean differences in scores for the three tests were within the range indicated by Shedd and 51 of the 62 children showed the order of scores he specified.

Adolescent

Two kinds of cognitive deficit associated with chronic schizophrenia.

The performance of 21 chronic schizophrenic patients was investigated on two tests of feature selection. It was found that patients with negative symptoms (muteness, withdrawal, etc.) were characterized by an extreme lack of persistence, but selected usual features; whereas patients with positive symptoms (hallucinations, delusions, etc.) had a normal degree of persistence, but selected unusual features.

Aged

Numerical evaluation of cytologic data. III. Selection of features for discrimination.

The proper selection of variables is important in assembling a profile to best describe a given group, whether of patients or cells, vis-à-vis other groups. The need often arises to determine which variables in comparable profiles best discriminate between the profiles. Three techniques for the evaluation and selection of variables on the basis of their potentiality for discrimination are discussed in this article. The Kruskal Wallis test is useful in determining if a certain feature (variable) has any statistical significance between groups. The ambiguity function after Genchi and Mori and the measure of detectability (d') are discussed as direct measurements of a feature's ability to discriminate between groups. Fully worked numerical example suitable for execution on a pocket calculator are given.

Cytological Techniques

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans

A comparative study highlights superiority of LSTM in crop genomic prediction.

We systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods, and found LSTM suitable for capturing additive and epistatic effects. Genomic prediction (GP) has been developed as an important method supporting crop breeding. By utilizing the phenotype values result from GP, breeders could make decisions in the seedling stage that consequently benefit for cost saving. In recent years, machine learning emerged as an efficient technology to solve modeling problems in many fields, including crop breeding. However, numerous modeling approaches have hindered the application of GP since breeders struggle to choose. Therefore, a comprehensively methodological research with guiding significance is extremely necessary. In the present study, we systematically evaluated three key determinants affecting prediction accuracy and the algorithm performance differences based on fifteen state-of-the-art GP methods. As for genomic feature processing, we found feature selection (SNP filtering approach) performed better than feature extraction (PCA method). Specifically, the feature relationship dependent methods (GBLUP, RNN, and LSTM) as well as DNN architecture showed superior performance with feature selection. Marker density analysis showed positive correlation with prediction accuracy in a limited threshold. Comparison on effect of population size demonstrated a positive correlation between trait genetic complexity and the optimal population size required. By testing fifteen modeling methods, we found LSTM network displayed superior performance, achieving the highest average STScore (0.967) across six datasets. Further research using all cell states or the latest cell states of LSTM inputs demonstrated its architecture particularly adept with capturing additive and epistatic QTL effects among SNPs. In conclusion, our findings provide basic principles for implementing GP in breeding project to maximize prediction accuracy while maintaining cost-effectiveness.

Plant Breeding

Sequence analysis of oligodeoxyribonucleotides by mass spectrometry. 2. Application of computerized pattern recognition to sequence determination of di-, tri-, and tetranucleotides.

A novel strategy for the sequence analysis of oligodeoxyribonucleotides has been devised which is based upon the analysis of intact underivatized oligonucleotides by mass spectrometry followed by interpretation of the mass-spectral data by computerized pattern-recognition techniques. The pyrolytic and electron-impact conditions of the mass spectrometer permit the cleavage of oligonucleotides of varying chain length and composition, yielding reproducible fragmentations and characteristic m/e values which can be used to reveal purine and/or pyrimidine base sequence information. The selection of optimum features (which are the ratios of peak heights of specific ions, or the linear combination of such ratios) has been done by an interactive feature selection method employing multidimensional k nearest-neighbor analysis and two-dimensional feature-space plots (nonlinear mappings) of the mass-spectral data. Features have been found which allow 100% classification accuracy in predicting the 5' and 3' terminus of all of the dinucleotides commonly found in DNA. Other specific features have been found which indicate adjacent nucleotides within a tetranucleotide. Knowledge of the adjacent nucleotide pairs present, in conjunction with the information as to the 3' or 5' position of the residues in each pair, permits the reconstruction of the sequence of the tetranucleotide.

Base Sequence

Deciphering microbial and metabolic influences in gastrointestinal diseases-unveiling their roles in gastric cancer, colorectal cancer, and inflammatory bowel disease.

INTRODUCTION: Gastrointestinal disorders (GIDs) affect nearly 40% of the global population, with gut microbiome-metabolome interactions playing a crucial role in gastric cancer (GC), colorectal cancer (CRC), and inflammatory bowel disease (IBD). This study aims to investigate how microbial and metabolic alterations contribute to disease development and assess whether biomarkers identified in one disease could potentially be used to predict another, highlighting cross-disease applicability. METHODS: Microbiome and metabolome datasets from Erawijantari et al. (GC: n = 42, Healthy: n = 54), Franzosa et al. (IBD: n = 164, Healthy: n = 56), and Yachida et al. (CRC: n = 150, Healthy: n = 127) were subjected to three machine learning algorithms, eXtreme gradient boosting (XGBoost), Random Forest, and Least Absolute Shrinkage and Selection Operator (LASSO). Feature selection identified microbial and metabolite biomarkers unique to each disease and shared across conditions. A microbial community (MICOM) model simulated gut microbial growth and metabolite fluxes, revealing metabolic differences between healthy and diseased states. Finally, network analysis uncovered metabolite clusters associated with disease traits. RESULTS: Combined machine learning models demonstrated strong predictive performance, with Random Forest achieving the highest Area Under the Curve(AUC) scores for GC(0.94[0.83-1.00]), CRC (0.75[0.62-0.86]), and IBD (0.93[0.86-0.98]). These models were then employed for cross-disease analysis, revealing that models trained on GC data successfully predicted IBD biomarkers, while CRC models predicted GC biomarkers with optimal performance scores. CONCLUSION: These findings emphasize the potential of microbial and metabolic profiling in cross-disease characterization particularly for GIDs, advancing biomarker discovery for improved diagnostics and targeted therapies.

Humans

Clinical and immunologic features of selective IgA deficiency.

Selective absence of serum and secretory IgA is probably the most common form of human immunodeficiency. High frequencies of recurrent sinusitis, otitis media, pneumonia, and atopy were noted among a group of 75 such patients, all but 4 of whom were Caucasian. Seven instances of familial absence of IgA were detected among 106 relatives of 34 of the group; in 1 family 1 member from each of 3 successive generations was affected. Two IgA-deficient children were later found to have normal amounts of serum IgA. Despite their humoral deficit, B lymphocytes bearing surface IgA were detected in 9/9 IgA-deficient patients in immunofluorescence studies of their peripheral blood lymphocytes. Although in vitro lymphocyte responses to 2 putative T-cell mitogens and to allogenic cells were normal, results of spontaneous rosette formation studies with sheep erythrocytes raise the possibility of a lymphocyte subpopulation deficit in this condition.

Absorption