PubMed HealthSearch

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Applications of quantum AI in brain disorder diagnosis: A systematic review.

BACKGROUND AND OBJECTIVE: Brain disorder diagnosis and prediction remain challenging because neuroimaging, electrophysiological, behavioral, and multimodal data are high-dimensional, noisy, heterogeneous, and limited by small clinical cohorts. This systematic review synthesised applications of quantum artificial intelligence (QAI) for brain disorder diagnosis, prediction, detection, and monitoring. METHODS: Following PRISMA guidelines, studies published from 2016 to 13 January 2026 were retrieved from Scopus, Web of Science, and IEEE Xplore. After screening, 36 studies met the eligibility criteria and were qualitatively analysed according to disorder category, data modality, QAI method, implementation setting, validation strategy, and performance. RESULTS: At the broader disease-group level, neurodegenerative disorders were the most frequently investigated, followed by mental health and psychiatric disorders. At the individual level, Parkinson's disease and schizophrenia were the leading applications, followed by depression, anxiety, Alzheimer's disease, and stress-related tasks. MRI-based modalities were the most frequently used data source, followed by multimodal data and EEG. Methodologically, primary QAI approaches were dominated by quantum neural and QDL architectures, followed by quantum-inspired optimization or feature-selection methods and quantum-kernel/conventional QML classifiers. Qiskit/IBM Quantum and PennyLane were the most frequently reported quantum software frameworks. However, most studies relied on simulators, classical quantum-inspired implementations, or unclear implementation settings, with limited real-hardware evaluation. CONCLUSIONS: QAI shows emerging potential for brain disorder analysis, particularly through hybrid quantum-classical learning, quantum neural architectures, quantum-kernel methods, and quantum-inspired optimization. Nevertheless, current evidence remains preliminary and requires larger datasets, subject-level and external validation, fair classical benchmarking, noise-resilient circuits, real quantum hardware evaluation, explainability, and clinical validation.

Humans

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n = 549) and a validation set (n = 236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60 mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60 mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans

Integrated salivary proteomic and metabolomic analyses reveal molecular characterization and novel biomarker panels of chronic obstructive pulmonary disease.

Chronic obstructive pulmonary disease (COPD) is a respiratory disorder characterized by chronic inflammation, oxidative stress, and metabolic dysregulation. The lack of convenient and easily-accessible non-invasive diagnostic approaches remains a major clinical challenge. This study applied an integrated saliva-based proteomic and untargeted metabolomic strategy to identify potential biomarkers for COPD classification. Comprehensive multi-omics analyses identified 225 differentially abundant proteins and 60 differentially abundant metabolites between patients with COPD and healthy controls, including 24 biologically relevant endogenous metabolites. Functional enrichment analyses revealed pronounced dysregulation of mitochondrial energy metabolism, redox homeostasis, lipid remodeling, and inflammatory-related pathways in COPD. By integrating salivary proteomic and metabolomic biomarkers, a stepwise feature selection combined with LASSO logistic regression was used to construct diagnostic models, yielding an optimized biomarker panel consisting of 11 proteins and 2 endogenous metabolites. This integrated model achieved excellent diagnostic performance, with an area under the ROC curve of 0.96. Collectively, these findings demonstrate that integrated salivary proteomic and metabolomic profiling provides a robust, non-invasive approach for COPD classification and offers a promising foundation for the development of biosensor-based diagnostic platforms and early disease detection. SIGNIFICANCE: Chronic obstructive pulmonary disease (COPD) remains a major global health burden. Current diagnostic approaches rely largely on spirometry and clinical assessment, which are limited in sensitivity for early-stage disease and unsuitable for large-scale screening. This study employs an integrated saliva-based proteomic and metabolomic strategy to identify non-invasive biomarkers for COPD classification. Our findings reveal coordinated dysregulation of mitochondrial energy metabolism, redox homeostasis, and lipid remodeling in COPD, highlighting the interconnected roles of metabolic reprogramming, oxidative stress, and inflammation in disease pathophysiology. Notably, a robust diagnostic panel comprising 11 proteins and 2 endogenous metabolites was established, achieving excellent classification performance (AUC of 0.96). To our knowledge, the integrated application of salivary proteomics and metabolomics for COPD diagnosis remains largely unexplored, underscoring the significance and translational potential of our findings.

Humans

Multiomics Integration Identifies a Molecular Subtype of Intrahepatic Cholangiocarcinoma With Enhanced Benefit From Adjuvant Therapy.

Intrahepatic cholangiocarcinoma (iCCA) is a molecularly heterogeneous liver cancer with a poor prognosis. Improved stratification is needed to guide postoperative therapy. In this study, we applied integrative multiomics analysis to classify iCCA and identify biomarkers predictive of adjuvant treatment benefit. Using publicly available datasets (including whole exome sequencing, RNA sequencing, proteomics, and phosphoproteomics from FU-iCCA cohort and a transcriptomic cohort GSE244807), we defined 3 robust molecular subtypes of iCCA. These subtypes exhibited distinct genomic alterations, pathway activation, and immune microenvironments, with significant differences in overall survival (OS). Through protein-protein interaction network analysis and consensus feature selection using 10 clustering algorithms, we prioritized 8 marker genes distinguishing the subtypes. A Cox proportional-hazards model constructed from these markers stratified patients into high- and low-risk groups. High-risk iCCA, characterized by elevated expression of markers such as CLDN18, MUC1, and MUC5AC, had significantly worse OS in the absence of adjuvant therapy. Notably, in an independent validation of 174 patients with iCCA who underwent resection (single-center cohort), high expression of any of these 3 markers were associated with markedly prolonged OS in patients who received adjuvant chemotherapy or chemoembolization, compared with those who did not. In contrast, marker-negative patients showed no clear benefit from adjuvant therapy. In conclusion, our multiomics approach identified a high-risk, mucin-enriched subtype of iCCA. CLDN18, MUC1, and MUC5AC emerge as candidate predictive biomarkers for adjuvant chemotherapy benefit in iCCA, warranting prospective validation to improve personalized postoperative management.

Humans

Unveiling the power of TIIC: A prognostic tool for esophageal adenocarcinoma.

BACKGROUND: Esophageal adenocarcinoma (EAC) remains a lethal malignancy with limited prognostic tools for guiding immunotherapy. Tumor-infiltrating immune cells (TIICs) play a critical role in EAC prognosis and treatment response. METHODS: We integrated single-cell RNA sequencing and bulk transcriptome data from TCGA and GEO databases. TIIC-specific RNAs were identified via tissue specificity index calculation combined with machine learning feature selection. Twenty machine learning algorithms were benchmarked to construct an optimal TIIC signature score (TIIC-Score) based on the comprehensive C-index. Immunotherapy response, genomic mutation, and copy number variation were analyzed. Summary-data-based Mendelian randomization (SMR) and two-sample Mendelian randomization (MR) were performed to explore genetic associations. Core prognostic TIIC-related genes were functionally validated in esophageal cancer cell lines through loss-of-function assays. RESULTS: The TIIC-Score demonstrated robust prognostic value for 1-, 2-, and 3-year overall survival across multiple cohorts, outperforming 22 published models. High TIIC-Score was associated with poor survival and increased chromosomal instability. Mutation profiling revealed high frequencies of TP53 (78.2%), TTN (48.7%), and SYNE1 (30.8%). MR analysis identified a significant association between gastro-oesophageal reflux and EAC risk at SNP rs8130507. Functionally, CCNI was upregulated in esophageal cancer cells, and its knockdown suppressed malignant phenotypes while promoting apoptosis, supporting its pro-tumorigenic role. CONCLUSION: The TIIC-Score provides a novel prognostic framework for EAC that effectively stratifies patient risk and may help identify individuals most likely to benefit from immunotherapy.

Esophageal adenocarcinoma

Development and validation of a serum peptidomic signature for early detection of asymptomatic ovarian cancer: A multi-center prospective study.

Early detection of asymptomatic ovarian cancer (asym-OC) remains a critical challenge, the failure of which underlies its high mortality. Performing serum peptidomic profiling of 843 participants in the cohort SOCFCP, we distill 1,081 initial features into a 7-marker panel for asym-OC detection via a biology-informed machine-learning (ML)-based feature selection strategy. Three markers significantly revert toward non-OC levels after surgery. Integrating the panel with age, CA125, and HE4, we develop and externally validate (n = 159) a LightGBM model, ProMS+. For early-stage OC detection, ProMS+ shows a specificity of 92.6% at 95.0% sensitivity, outperforming CA125 (44.7%), HE4 (11.2%), and Risk of Ovarian Malignancy Algorithm (ROMA) (24.0%), with an area under the curve (AUC) of 0.993. In a simulated high-risk population (n = 100,000; OC prevalence = 1%), ProMS+ yields a high AUC (0.983) and a higher positive predictive value than CA125, HE4, and Age + CA125 + HE4 combined model (0.201 vs. 0.027, 0.090, and 0.064). ProMS+ offers a promising, non-invasive, and interpretable approach for the early detection of asym-OC.

Humans

Diffusion coefficients of hemoglobin by intensity fluctuation spectroscopy: effects of varying pH and ionic strength.

Measurements of the mutual diffusion coefficients (D) of the liganded human hemoglobins (Hb) oxy-HbA and oxy-HbS were performed as a function of Hb concentration (CHb), pH, and ionic strength (tau) by intensity fluctuation spectroscopy (IFS). Average diffusion coefficients, (D), and normalized variances, ((D/(D) - 1)2), were recorded. Results are reported and select features are discussed quantitatively. (a) for tau = 0.15 M, the shape of the (d) vs. CHb curve is found to vary with pH. We developed a precise description of this effect in the form of an algebraic relationship between (D), CHb, and Z, the titration charge. (b) only slight differences between the (D) values of oxy-HbS and oxy-HbA are observed, at tau = 0.15 M, for CHb Less Than or Equal To 10 g%. These differences are explained by the theory of part a. (c) No evidence of aggregation is found in solutions of oxy-HbA or oxy-HbS, at tau = 0.15 M, for CHb Less Than or Equal To 10 g%. (d) Indications of aggregation appear in oxy-HbA solutions at very low concentrations of salt. An estimate is made of the extent of aggregation, and the average radius of a cluster is determined.

Diffusion

A statewide characterization of hospital infection control practices and practitioners.

Selected features of infection control programs among the 163 general hospitals in Tennessee were surveyed in 1976 and 1979. Each hospital but one had a designated infection control practitioner. Three-fourths of the hospitals had fewer than 200 beds and most were in rural areas. The practitioners in these small hospitals worked in an isolated professional milieu: few (4%) had attended a basic training course or were members of a national (11%) or local (16%) infection control association. They also had significantly less access to standard infection control resource publications than did practitioners in large hospitals. Use of aqueous quaternary ammonium compounds for disinfection was reported by 37% of all hospitals in 1979; 68% of hospitals routinely performed bacteriologic cultures of personnel or the environment. In contrast, only 3% of hospitals did not have a policy specifying the use of sterile closed-system drainage of indwelling bladder catheters. Although these practices varied somewhat by hospital size, the differences were not statistically significant. Modest improvement in each parameter was noted since 1976. Pathology was the most common medical specialty (34%) among chairman of infection control committees; internal medicine and pediatrics accounted for only 13%. The practice of routine microbiologic monitoring was significantly more common among hospitals with chairmen who were pathologists. The implications of these findings for national priorities in hospital infection control are discussed.

Bacteria

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

Dual-Matrix Platform for Highly Specific Multi-Omics Profiling of Renal Cell Carcinoma.

Multiomics interrogation provides complementary information beyond single-omics approaches for improved disease characterization. To enable such multilayer profiling, we expanded the rapid functionalized mesoporous nanoparticle-coupled laser desorption/ionization mass spectrometry (fMNPLDI-MS) platform by designing two structurally homologous but functionally tailored fMNPs. This design enables efficient acquisition of both serum metabolic and peptide fingerprints from a total of only 2.05 μL of serum, with an LDI MS analysis time of approximately 90 s per sample, while addressing the limitation of single-matrix systems in simultaneously optimizing analytical performance for different biomolecular species. Through statistical analysis and machine learning-based feature selection, an integrated multiomics biomarker panel was established, comprising 5 peptides and 4 metabolites. Notably, this integrated panel outperformed both single-omics panels across all evaluation metrics in the validation set, improving the area under curve from 0.985 to 1.000 and increasing the classification accuracy from 0.947 (metabolites) and 0.930 (peptides) to 0.965, while showing consistent improvements in F1-score, precision, and recall. Collectively, these results demonstrate the robust performance of the dual-matrix design and multiomics integration for renal cell carcinoma classification, with potential relevance for broader applications in complex disease profiling.

Carcinoma, Renal Cell

From Variability to Consensus: Rescoring Harmonizes Peptide Identification across Diverse Search Engines and Data Sets.

Peptide-spectrum match (PSM) rescoring has become standard in proteomics workflows, improving peptide identification accuracy across diverse search engines. Despite the availability of multiple rescoring strategies, systematic comparisons spanning several search engines, data sets, and database configurations remain limited. Here, we benchmarked seven publicly available search engines, evaluating standard target-decoy-based false discovery rate (FDR) estimation alongside Percolator, MS2Rescore, and Oktoberfest across four data sets acquired on different mass spectrometry platforms in data-dependent mode and searched against protein databases of varying size and composition. Rescoring substantially increased identification consensus and reduced variability between search engines, with prediction-based approaches yielding the largest gains. While database size had limited impact for human data sets, it significantly affected identification rates on a metaproteomic data set. Entrapment-based evaluation indicated generally adequate FDR control across methods, although prediction-based rescoring exhibited a higher tendency toward FDR underestimation in specific configurations. Overall, advanced rescoring strategies harmonize peptide identification outcomes across search engines, thereby enhancing robustness and comparability in proteomics analyses. However, careful feature selection and appropriate database choice remain essential to ensure reliable FDR control and optimal performance across diverse experimental settings.

Search Engine

Isolation and characterization of distinct domains of sarcolemma and T-tubules from rat skeletal muscle.

1. Several cell-surface domains of sarcolemma and T-tubule from skeletal-muscle fibre were isolated and characterized. 2. A protocol of subcellular fractionation was set up that involved the sequential low- and high-speed homogenization of rat skeletal muscle followed by KCl washing, Ca2+ loading and sucrose-density-gradient centrifugation. This protocol led to the separation of cell-surface membranes from membranes enriched in sarcoplasmic reticulum and intracellular GLUT4-containing vesicles. 3. Agglutination of cell-surface membranes using wheat-germ agglutinin allowed the isolation of three distinct cell-surface membrane domains: sarcolemmal fraction 1 (SM1), sarcolemmal fraction 2 (SM2) and a T-tubule fraction enriched in protein tt28 and the alpha 2-component of dihydropyridine receptor. 4. Fractions SM1 and SM2 represented distinct sarcolemmal subcompartments based on different compositions of biochemical markers: SM2 was characterized by high levels of beta 1-integrin and dystrophin, and SM1 was enriched in beta 1-integrin but lacked dystrophin. 5. The caveolae-associated molecule caveolin was very abundant in SM1, SM2 and T-tubules, suggesting the presence of caveolae or caveolin-rich domains in these cell-surface membrane domains. In contrast, clathrin heavy chain was abundant in SM1 and T-tubules, but only trace levels were detected in SM2. 6. Immunoadsorption of T-tubule vesicles with antibodies against protein tt28 and against GLUT4 revealed the presence of GLUT4 in T-tubules under basal conditions and it also allowed the identification of two distinct pools of T-tubules showing different contents of tt28 and dihydropyridine receptors. 7. Our data on distribution of clathrin and dystrophin reveal the existence of subcompartments in sarcolemma from muscle fibre, featuring selective mutually exclusive components. T-tubules contain caveolin and clathrin suggesting that they contain caveolin- and clathrin-rich domains. Furthermore, evidence for the heterogeneous distribution of membrane proteins in T-tubules is also presented.

Animals

Opposing effects of estradiol and progesterone on oxytocin receptors in rabbit uterus.

Estradiol-17beta administration to young (10- to 12-week-old) rabbits to produce the "estrogen-dominated" uterus increased the uterine contractile response to both oxytocin and methacholine in vitro. In "progesterone-dominated" uteri, obtained from rabbits that received progesterone for 4 days after estrogen pretreatment, the contractile response to oxytocin in vitro was selectively abolished; the response to methacholine was unaffected. Parallel changes were observed in the concentration (but not affinity) of specific sites in uterine microsomal membranes that bind [(3)H]oxytocin with selectivity features expected for oxytocin receptors. Thus, estrogen-dominated uteri have an increased number of specific [(3)H]oxytocin binding sites per mg of membrane protein relative to untreated controls, whereas specific oxytocin binding sites are reduced to barely detectable levels in the progesterone-dominated uterus. Similar results are obtained when binding sites are measured in membranes from the myometrium of estrogen- or progesterone-dominated uteri. Short-term (24-hr) progesterone administration to estrogen-pretreated rabbits decreased, but did not abolish, specific [(3)H]oxytocin binding; the concentration of specific [(3)H]oxytocin binding sites was reduced without influence on the affinity of these sites. A sublethal dose of actinomycin D, administered over a 24-hr period to rabbits pretreated with estradiol for 4 days, likewise reduced specific oxytocin binding; additive effects were not observed when progesterone and actinomycin D were administered together. These results suggest that the regulatory effects of estrogens and progesterone upon the rabbit uterine contractile response to oxytocin are achieved, at least in part, by the opposing actions of these steroids in regulating the number of oxytocin receptors in smooth muscle cells. Estradiol increased the concentration of uterine oxytocin receptors; the maintenance of high receptor levels appears to depend upon the continuous de novo synthesis of oxytocin receptors. In contrast, progesterone, like actinomycin D, appears to act at the nuclear locus to repress synthesis of oxytocin receptors.

Animals

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans

Stage-specific ROMO1 in rheumatoid arthritis: predictive immune insights into the MIF pathway and HLA-DR/IL2RA axis via integrated GWAS, transcriptomic, single-cell, and spatial profiling.

Emerging evidence links reactive oxygen species modulator 1 (ROMO1), a key mitochondrial ROS regulator, to rheumatoid arthritis (RA) pathogenesis. However, its exact mechanism remains elusive given the conflicting evidence about its specific function. We used a four-level integrative framework combining multi-omics data and literature‑supported mechanistic inference. At the genetic level, Mendelian randomization (MR) was performed to explore potential causal relationships between ROMO1, IL2RA, HLA-DR, MIF, and RA risk, followed by differential expression analysis and machine learning-based feature selection to identify key mROS genes. The temporal expression dynamics of ROMO1 were assessed in RA progression. At the cellular and tissue levels, we integrated single-cell RNA sequencing and spatial transcriptomics to map cell-type-specific expression and synovial localization of ROMO1-related immune cells and pathways. Finally, our multi-omics findings were contextualized with literature-supported mechanistic inference. (1) MR results were consistent with a potential protective effect of ROMO1 on RA (OR = 0.52) and its potential regulation of risk factors IL2RA (OR = 0.46) and HLA-DR (OR = 0.40). Conversely, IL2RA (OR = 1.42), HLA-DR (OR = 1.88), and MIF (OR = 1.17) were positively associated with RA risk. Additionally, ROMO1 was identified as a top candidate diagnostic predictor with stage-specific dynamics: downregulated in the early but upregulated in the late/remission stages. (2) Single-cell RNA sequencing showed ROMO1's cell-specific expression in CD14+ HLA-DR+ CD74+ monocytes and CD4+ IL2RA+ T cells. Cell communication analysis further suggested that these cells may participate in MIF pathway regulation. Spatial transcriptomics subsequently identified that ROMO1-related cells localized to synovial pathological regions, with MIF pathway changes correlated with RA progression. (3) Finally, literature-supported mechanistic inference suggests that ROMO1 may modulate mROS levels to promote anti-inflammatory M2 macrophage polarization, which could theoretically contribute to reduced systemic inflammation and the alleviation of multi-organ decline in RA. This integrated multi-omics investigation, supported by literature-based mechanistic inference, suggests ROMO1 as a stage-dependent biomarker candidate and potential immune regulator in RA.

Humans

Estimating population structure using epigenome-wide methylation data.

Population stratification is one of the source of inflation in epigenome-wide association studies (EWAS) when not properly accounted for. To address this, we developed methylation population scores (MPSs) to predict genetic principal components (GPCs) using a feature selection approach. We used multi-ethnic DNA methylation data from Illumina EPIC arrays across five cohorts, including MESA (n&#xa0;=&#xa0;929), CARDIA (n&#xa0;=&#xa0;1123), JHS (n&#xa0;=&#xa0;1365), ARIC (n&#xa0;=&#xa0;2338), and HCHS/SOL (n&#xa0;=&#xa0;1475), randomly splitting participants into training (85%) and test (15%) sets. Within each cohort, associations between GPCs and CpG sites were estimated using linear regression adjusting for age, sex, smoking and alcohol use, race/ethnicity, body mass index, and cell type proportions, followed by meta-analysis and selection of CpGs with FDR <0.05. We then applied a two-stage weighted least squares Lasso regression to construct MPSs, adjusting for the aforementioned covariates. In the test dataset, MPSs showed strong correlation with GPCs, with R&#xb2; ranging from 0.27 (MPS7 vs. GPC7) to 0.98 (MPS1 vs. GPC1). Visualization demonstrated that MPSs recapitulated the pattern shown by GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups and outperformed methylation-based principal components constructed using alternative published methods. Additionally, MPSs showed comparable performance to GPCs in reducing inflation in EWAS. Overall, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations, and provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent.

Humans

CpGene: a web application for epigenetic signature identification from DNA methylation arrays.

MOTIVATION: DNA methylation (DNAme) is the best studied epigenetic mechanism that plays pivotal role in tissue differentiation and epigenetic disruption has been correlated to diverse disease types (e.g. cancer, metabolic disorders). While various DNAme array platforms have been discovered, data analysis remains a challenging task which often requires in-depth bioinformatic expertise. Here, we developed a user-friendly web-based application for data analysis and visualization that accommodates users ranging from early-career basic/translational researchers to experienced bioinformaticians. RESULTS: CpGene is a web application for analyzing DNA methylation array data. It supports Illumina 450K, EPIC, and EPICv2 methylation array platforms and processes .idat files with integrated preprocessing, normalization, and quality control. Biomarker discovery is available through either classic differential methylation point analysis or machine learning-based feature selection as well as gene enrichment analysis. Results are summarized with clear visualizations, to aid interpretation. By combining these functions in a unified interface, CpGene streamlines methylation analysis and helps identify CpG sites and genes with biological and clinical relevance. AVAILABILITY AND IMPLEMENTATION: CpGene is openly accessible as a web service through http://cpgene.duckdns.org:8001/ and it's source code is available on https://github.com/kostaslazaros/cpgenene.

DNA Methylation

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans