PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Deciphering the swordtail's tale: a molecular and evolutionary quest.

The power of sexual selection to influence the evolution of morphological traits was first proposed more than 130 years ago by Darwin. Though long a controversial idea, it has been documented in recent decades for a host of animal species. Yet few of the established sexually selected features have been explored at the level of their genetic or molecular foundations. In a recent report, Zauner et al.1 describe some of the molecular features associated with one of the best characterized of sexually selected traits, the male-specific tail "sword" seen in certain species of the fish genus Xiphophorus. Zauner et al. find that the msxC gene, a gene previously implicated in fin development from work in zebrafish, is dramatically and specifically upregulated in the development of the ventral caudal fin rays, which give rise to the sword, in males. The results provide the first molecular insight into the development of this sexually selected trait while prompting new questions about the structure of the entire genetic network that underlies this trait. To fully understand the molecular-genetic and evolutionary history of this network, however, it will be essential to determine whether sword-development is a basal or derived trait in Xiphophorus.

Animals↗

Oscillatory responses in cat visual cortex exhibit inter-columnar synchronization which reflects global stimulus properties.

A fundamental step in visual pattern recognition is the establishment of relations between spatially separate features. Recently, we have shown that neurons in the cat visual cortex have oscillatory responses in the range 40-60 Hz (refs 1, 2) which occur in synchrony for cells in a functional column and are tightly correlated with a local oscillatory field potential. This led us to hypothesize that the synchronization of oscillatory responses of spatially distributed, feature selective cells might be a way to establish relations between features in different parts of the visual field. In support of this hypothesis, we demonstrate here that neurons in spatially separate columns can synchronize their oscillatory responses. The synchronization has, on average, no phase difference, depends on the spatial separation and the orientation preference of the cells and is influenced by global stimulus properties.

Animals↗

Epstein-Barr virus-negative post-transplant lymphoproliferative disorders: a distinct entity?

Post-transplant lymphoproliferative disorders (PTLDs) are usually but not invariably associated with Epstein-Barr virus (EBV). The reported incidence, however, of EBV-negative PTLDs varies widely, and it is uncertain whether they should be considered analogous to EBV-positive PTLDs and whether they have any distinctive features. Therefore, the EBV status of 133 PTLDs from 80 patients was determined using EBV-encoded small ribonucleic acid (EBER) in situ hybridization stains with or without Southern blot EBV terminal repeat analysis. The morphologic, immunophenotypic, genotypic, and clinical features of the EBV-negative PTLDs were reviewed, and selected features were compared with EBV-positive cases. Twenty-one percent of patients had at least one EBV-negative PTLD (14% of biopsies). The initial EBV-negative PTLDs occurred a median of 50 months post-transplantation compared with 10 months for EBV-positive cases. Although only 2% of PTLDs from before 1991 were EBV negative, 23% of subsequent PTLDs were EBV negative (p <0.001). Of the EBV-negative PTLDs, 67% were of monomorphic type (M-PTLD) compared with 42% of EBV-positive cases (p <0.05). The other EBV-negative PTLDs were of infectious mononucleosis-like, plasma cell-rich (n = 2), small B-cell lymphoid neoplasm, large granular lymphocyte disorder (n = 4) and polymorphic (P) types. B-cell clonality was established in 14 specimens and T-cell clonality was established in three (two patients). None of the remaining specimens were studied with Southern blot analysis and some had no ancillary studies. Rearrangement of c-MYC was identified in two M-PTLDs with small noncleaved-like features, and rearrangement of BCL-2 was found in one large noncleaved-like M-PTLD. Ten patients were alive at 3 to 63 months (only three patients received chemotherapy). Seven patients, all with M-PTLDs, are dead at 0.3 to 6 months. Therefore, EBV-negative PTLDs have distinct features, but some do respond to decreased immunosuppression, similar to EBV-positive cases, suggesting that EBV positivity should not be an absolute criterion for the diagnosis of a PTLD.

Adult↗

Recognizing plankton images from the shadow image particle profiling evaluation recorder.

We present a system to recognize underwater plankton images from the shadow image particle profiling evaluation recorder (SIPPER). The challenge of the SIPPER image set is that many images do not have clear contours. To address that, shape features that do not heavily depend on contour information were developed. A soft margin support vector machine (SVM) was used as the classifier. We developed a way to assign probability after multiclass SVM classification. Our approach achieved approximately 90% accuracy on a collection of plankton images. On another larger image set containing manually unidentifiable particles, it also provided 75.6% overall accuracy. The proposed approach was statistically significantly more accurate on the two data sets than a C4.5 decision tree and a cascade correlation neural network. The single SVM significantly outperformed ensembles of decision trees created by bagging and random forests on the smaller data set and was slightly better on the other data set. The 15-feature subset produced by our feature selection approach provided slightly better accuracy than using all 29 features. Our probability model gave us a reasonable rejection curve on the larger data set.

Algorithms↗

Predicting diagnostic gene biomarkers associated with immune infiltration in patients with diabetes.

Diabetes is a global public health problem with various complications, which can lead to disability and mortality. This study identified potential diagnostic markers for diabetes and explored the immunometabolic mechanisms in the pathological process. The gene expression of 17 diabetes cases and 16 normal controls were obtained from the Gene Expression Omnibus (GEO) database. The "limma" package was employed for screening differentially expressed genes (DEGs). Gene functions and enriched pathways of DEGs were analyzed via Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses. Candidate key genes were screened using the least absolute shrinkage and selection operator (LASSO) regression model and support vector machine recursive feature elimination (SVM-RFE) analysis. The diagnostic effectiveness of identified markers was further verified via the receiver operating characteristic (ROC) curve. The compositional patterns of immune cell infiltration and signaling pathway enrichment associated with key genes were explored via single sample Gene Set Enrichment Analysis (ssGSEA) and GSEA analysis, respectively. Possible miRNAs interacting with key genes were predicted via miRcode database. B2M, FTL, SH3BGRL3, and SOD2 were recognized as diagnostic markers for diabetes based on LASSO regression and the support vector machine recursive feature elimination (SVM-RFE) feature selection algorithm. Analysis of immune cell infiltration demonstrated that the four key genes were related to B cells, neutrophils, macrophages, and CD8+ T cells. The diagnostic value of B2M, FTL, and SOD2 for diabetes was higher than that of SH3BGRL3 according to the ROC curve. Validation experiments indicated that the mRNA expression of B2M and FTL was increased in liver tissues of diabetic mice. B2M and FTL can act as diagnostic markers for diabetes and contribute to new understandings of the disease's molecular mechanisms.

Humans↗

Selection of patient samples and genes for outcome prediction.

Gene expression profiles with clinical outcome data enable monitoring of disease progression and prediction of patient survival at the molecular level. We present a new computational method for outcome prediction. Our idea is to use an informative subset of original training samples. This subset consists of only short-term survivors who died within a short period and long-term survivors who were still alive after a long follow-up time. These extreme training samples yield a clear platform to identify genes whose expression is related to survival. To find relevant genes, we combine two feature selection methods -- entropy measure and Wilcoxon rank sum test -- so that a set of sharp discriminating features are identified. The selected training samples and genes are then integrated by a support vector machine to build a prediction model, by which each validation sample is assigned a survival/relapse risk score for drawing Kaplan-Meier survival curves. We apply this method to two data sets: diffuse large-B-cell lymphoma (DLBCL) and primary lung adenocarcinoma. In both cases, patients in high and low risk groups stratified by our risk scores are clearly distinguishable. We also compare our risk scores to some clinical factors, such as International Prognostic Index score for DLBCL analysis and tumor stage information for lung adenocarcinoma. Our results indicate that gene expression profiles combined with carefully chosen learning algorithms can predict patient survival for certain diseases.

Biomarkers, Tumor↗

Sodium dodecyl sulphate-polyacrylamide gel electrophoresis of proteins in dry-cured hams: data registration and multivariate analysis across multiple gels.

This study investigates whether dry-cured hams from two European countries can be distinguished using SDS-PAGE. Thirty-seven commercial hams (19 Spanish, 18 French) were used in the study. Four protein fractions were extracted from each sample, with sufficient material prepared to allow each fraction to be analysed in triplicate lanes. The complete extraction process was carried out in duplicate. The 24 specimens originating from each ham sample were randomly allocated to different lane positions and gels, as were at least two reference lanes (for reference proteins). In total, 118 gels were prepared. Mathematical routines were developed using a matrix language to process the gel image files. Procedures were written to carry out 'within-gel' image correction, lane extraction and normalization, 'between-gel' data registration and linear discriminant analysis (LDA) of each fraction's data to establish whether the provenance could be systematically distinguished. The between-gel registration was carried out using a genetic algorithm (GA). Feature selection was also performed using a GA, to pass subsets of features to the LDA routine. Cross-validated classification success rates were 84, 91, 81 and 85%, respectively, for the four fractions. We conclude that SDS-PAGE can be conducted in a sufficiently quantitative manner and can potentially verify the provenance of regional speciality dry-cured hams.

Electrophoresis, Polyacrylamide Gel↗

GenSo-EWS: a novel neural-fuzzy based early warning system for predicting bank failures.

Bank failure prediction is an important issue for the regulators of the banking industries. The collapse and failure of a bank could trigger an adverse financial repercussion and generate negative impacts such as a massive bail out cost for the failing bank and loss of confidence from the investors and depositors. Very often, bank failures are due to financial distress. Hence, it is desirable to have an early warning system (EWS) that identifies potential bank failure or high-risk banks through the traits of financial distress. Various traditional statistical models have been employed to study bank failures [J Finance 1 (1975) 21; J Banking Finance 1 (1977) 249; J Banking Finance 10 (1986) 511; J Banking Finance 19 (1995) 1073]. However, these models do not have the capability to identify the characteristics of financial distress and thus function as black boxes. This paper proposes the use of a new neural fuzzy system [Foundations of neuro-fuzzy systems, 1997], namely the Generic Self-organising Fuzzy Neural Network (GenSoFNN) [IEEE Trans Neural Networks 13 (2002c) 1075] based on the compositional rule of inference (CRI) [Commun ACM 37 (1975) 77], as an alternative to predict banking failure. The CRI based GenSoFNN neural fuzzy network, henceforth denoted as GenSoFNN-CRI(S), functions as an EWS and is able to identify the inherent traits of financial distress based on financial covariates (features) derived from publicly available financial statements. The interaction between the selected features is captured in the form of highly intuitive IF-THEN fuzzy rules. Such easily comprehensible rules provide insights into the possible characteristics of financial distress and form the knowledge base for a highly desired EWS that aids bank regulation. The performance of the GenSoFNN-CRI(S) network is subsequently benchmarked against that of the Cox's proportional hazards model [J Banking Finance 10 (1986) 511; J Banking Finance 19 (1995) 1073], the multi-layered perceptron (MLP) and the modified cerebellar model articulation controller (MCMAC) [IEEE Trans Syst Man Cybern: Part B 30 (2000) 491] in predicting bank failures based on a population of 3635 US banks observed over a 21 years period. Three sets of experiments are performed-bank failure classification based on the last available financial record and prediction using financial records one and two years prior to the last available financial statements. The performance of the GenSoFNN-CRI(S) network as a bank failure classification and EWS is encouraging.

Accidents↗

Somatosensory cortical mechanisms of feature detection in tactile and kinesthetic discrimination.

Neurons in somatosensory cortex of primates process sensory information from the hand by integrating information from large populations of receptors to extract specific features. Tactile neurons in areas 1 and 2 are shown to select features such as contact area, edge orientation, motion across the skin, or direction of movement. Features coded by kinesthetic neurons in areas 3a and 2 relate to joint movement, the joint angle around which the movement occurs, or coordinated postures of the hand and arm. An even higher order cortical cell integrates tactile and kinesthetic information; these "haptic neurons" respond optimally to contact of objects actively grasped in the hand. These global features are coded at the expense of loss of information concerning fine-grained spatial detail.

Animals↗

Radiomics-based gradient boosting model on contrast-enhanced MRI for non-invasive prediction of epidermal growth factor receptor expression and therapeutic response to EGFR-targeted antibody-drug conjugates in high-grade glioma organoid models.

BACKGROUND: Epidermal growth factor (EGF) and its receptor EGF(EGFR) play crucial roles in glioblastoma (GBM) prognosis. However, non-invasive assessment of their expression remains challenging. This study aimed to determine whether radiomics features extracted from contrast-enhanced MRI could predict EGFR expression in high-grade gliomas (HGG) and to explore their associations with immune infiltration and therapeutic response of EGFR-Targeted antibody drug conjugates(EGFR-ADCs). METHODS: We extracted radiomic features from contrast-enhanced MRI of 298 GBM patients from The Cancer Imaging Archive (TCIA) and matched them with RNA-seq data from The Cancer Genome Atlas (TCGA). Feature selection was performed using minimum redundancy maximum relevance (mRMR) and recursive feature elimination (RFE). Machine learning models were built to predict EGF/EGFR expression. Radiogenomic associations were validated by immune infiltration analysis. Patient-Derived Tumor-Like Cell Clusters (PTC) were used to compare the antitumor efficacy of EGFR- ADCs and temozolomide. RESULTS: Elevated EGF/EGFR expression correlated with poor prognosis and increased infiltration of M2 macrophages, regulatory T cells, and CD4&#x207a; memory T cells. Pathway analysis demonstrated significant enrichment of the mechanistic target of rapamycin (mTOR) and Mitogen-Activated Protein Kinase (MAPK) signaling cascades. Radiomics-based prediction models achieved robust performance (AUC&#x2009;>&#x2009;0.85) in stratifying EGFR expression status. In EGFR-positive tumor tissues, EGFR-ADCs exerted antitumor efficacy similar to that of temozolomide. CONCLUSIONS: EGF/EGFR expression is associated with immunosuppressive microenvironments and adverse outcomes in HGG. Radiomics may provide a non-invasive approach for estimating EGFR expression, although model performance requires external validation and EGFR-ADCs showed partial inhibitory activity within the tested range, though potency remains to be defined.These findings suggest a framework into radiogenomic stratification and targeted therapy in GBM.

Radiomics↗

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies↗

Pathomics-based machine learning models for predicting METTL5 expression and prognosis in lung adenocarcinoma.

BACKGROUND: METTL5, an N6-methyladenosine (m6A) RNA methyltransferase, has been implicated in tumor progression, but its prognostic value and non-invasive prediction in lung adenocarcinoma (LUAD) remain unclear. This study aimed to develop a pathomics-based machine learning model to predict METTL5 expression from histopathological images and evaluate its prognostic significance in LUAD. METHODS: A total of 327 LUAD patients from The Cancer Genome Atlas (TCGA) with matched hematoxylin and eosin (H&E) slides, transcriptomic, and clinical data were included and randomly divided into training and validation sets (7:3). Quantitative histopathological features were extracted using PyRadiomics. Feature selection was performed via maximum relevance minimum redundancy (mRMR) and recursive feature elimination (RFE), followed by construction of a Gradient Boosting Machine (GBM) model. A pathomics score (PS) was generated to assess prognostic relevance. Survival analyses, gene set variation analysis (GSVA), tumor mutational burden (TMB), immune infiltration analysis, and in vitro functional assays were conducted. RESULTS: METTL5 overexpression was independently associated with poor overall survival [hazard ratio (HR) =1.637, P=0.007]. The model achieved good predictive performance [area under the curve (AUC) =0.847 in the training set and 0.752 in the validation set]. High PS was significantly associated with worse survival and remained an independent prognostic factor (HR =1.563, P=0.03). Elevated PS correlated with altered metabolic pathways, increased TMB, and immune microenvironment changes. METTL5 knockdown reduced proliferation, migration, invasion, and epithelial-mesenchymal transition (EMT) in A549 cells. CONCLUSIONS: The pathomics-based model accurately predicts METTL5 expression and provides prognostic stratification in LUAD, supporting its potential as a practical imaging-derived biomarker.

Methyltransferase-like 5↗

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer↗

Computer-aided characterization of mammographic masses: accuracy of mass segmentation and its effects on characterization.

Mass segmentation is used as the first step in many computer-aided diagnosis (CAD) systems for classification of breast masses as malignant or benign. The goal of this paper was to study the accuracy of an automated mass segmentation method developed in our laboratory, and to investigate the effect of the segmentation stage on the overall classification accuracy. The automated segmentation method was quantitatively compared with manual segmentation by two expert radiologists (R1 and R2) using three similarity or distance measures on a data set of 100 masses. The area overlap measures between R1 and R2, the computer and R1, and the computer and R2 were 0.76 +/- 0.13, 0.74 +/- 0.11, and 0.74 +/- 0.13, respectively. The interobserver difference in these measures between the two radiologists was compared with the corresponding differences between the computer and the radiologists. Using three similarity measures and data from two radiologists, a total of six statistical tests were performed. The difference between the computer and the radiologist segmentation was significantly larger than the interobserver variability in only one test. Two sets of texture, morphological, and spiculation features, one based on the computer segmentation, and the other based on radiologist segmentation, were extracted from a data set of 249 films from 102 patients. A classifier based on stepwise feature selection and linear discriminant analysis was trained and tested using the two feature sets. The leave-one-case-out method was used for data sampling. For case-based classification, the area Az under the receiver operating characteristic (ROC) curve was 0.89 and 0.88 for the feature sets based on the radiologist segmentation and computer segmentation, respectively. The difference between the two ROC curves was not statistically significant.

Algorithms↗

Object and scene analysis by saccadic eye-movements: an investigation with higher-order statistics.

Based on an information theoretical approach, we investigate feature selection processes in saccadic object and scene analysis. Saccadic eye movements of human observers are recorded for a variety of natural and artificial test images. These experimental data are used for a statistical evaluation of the fixated image regions. Analysis of second-order statistics indicates that regions with higher spatial variance have a higher probability to be fixated, but no significant differences beyond these variance effects could be found at the level of power spectra. By contrast, an investigation with higher-order statistics, as reflected in the bispectral density, yielded clear structural differences between the image regions selected by saccadic eye movements as opposed to regions selected by a random process. These results indicate that nonredundant, intrinsically two-dimensional image features like curved lines and edges, occlusions, isolated spots, etc. play an important role in the saccadic selection process which must be integrated with top-down knowledge to fully predict object and scene analysis by human observers.

Adult↗

Entropy-based gene ranking without selection bias for the predictive classification of microarray data.

BACKGROUND: We describe the E-RFE method for gene ranking, which is useful for the identification of markers in the predictive classification of array data. The method supports a practical modeling scheme designed to avoid the construction of classification rules based on the selection of too small gene subsets (an effect known as the selection bias, in which the estimated predictive errors are too optimistic due to testing on samples already considered in the feature selection process). RESULTS: With E-RFE, we speed up the recursive feature elimination (RFE) with SVM classifiers by eliminating chunks of uninteresting genes using an entropy measure of the SVM weights distribution. An optimal subset of genes is selected according to a two-strata model evaluation procedure: modeling is replicated by an external stratified-partition resampling scheme, and, within each run, an internal K-fold cross-validation is used for E-RFE ranking. Also, the optimal number of genes can be estimated according to the saturation of Zipf's law profiles. CONCLUSIONS: Without a decrease of classification accuracy, E-RFE allows a speed-up factor of 100 with respect to standard RFE, while improving on alternative parametric RFE reduction strategies. Thus, a process for gene selection and error estimation is made practical, ensuring control of the selection bias, and providing additional diagnostic indicators of gene importance.

Adenocarcinoma↗

Anatomy, histology, and pathology of coronary arteries: a review relevant to new interventional and imaging techniques--Part II.

In the last 15 years, intense interest has focused on various interventional, pharmacologic, and mechanical forms of therapy for the treatment of atherosclerotic coronary artery disease. Many techniques and devices (dilating balloons, perfusion catheters, thermal probes and balloons, lasers, atherectomy devices, stents, intravascular ultrasound) have been used or are under study for future use. Many of these techniques and devices require an understanding of histologic and pathologic features of the coronary arteries and diseases which affect them. This article reviews selective areas of anatomy, histology, and pathology relevant to the use of various new interventional techniques. Part II of this four-part review will focus on aging changes seen in the epicardial coronary arteries and will review selected features of atherosclerotic plaque, including fissure and topography.

Angioplasty, Balloon, Coronary↗

Evolutionary weighting of image features for diagnosing of CNS tumors.

This paper concerns an application of evolutionary feature weighting for diagnosis support in neuropathology. The original data in the classification task are the microscopic images of ten classes of central nervous system (CNS) neuroepithelial tumors. These images are segmented and described by the features characterizing regions resulting from the segmentation process. The final features are in part irrelevant. Thus, we employ an evolutionary algorithm to reduce the number of irrelevant attributes, using the predictive accuracy of a classifier ('wrapper' approach) as an individual's fitness measure. The novelty of our approach consists in the application of evolutionary algorithm for feature weighting, not only for feature selection. The weights obtained give quantitative information about the relative importance of the features. The results of computational experiments show a significant improvement of predictive accuracy of the evolutionarily found feature sets with respect to the original feature set.

Algorithms↗