PubMed HealthSearch

Biomedical subjects

Hajeong Lee

Publications and source records attributed to Hajeong Lee.

2 recordsLinked to original sources

Development and external validation of an explainable machine learning model for predicting chronic kidney disease progression in the Korean population.

BACKGROUND: Current risk stratification models, such as the Kidney Failure Risk Equation (KFRE), exhibit variable performance across ethnic groups and fail to capture dynamic clinical trajectories. This study aimed to develop and validate a Korean-specific machine learning (ML) model for predicting chronic kidney disease (CKD) progression using an ensemble approach. METHODS: We used electronic health records from Seoul National University Hospital for model development (n = 28,209) and the Korean Genome and Epidemiology Study (KoGES) CKD cohort for external validation (n = 3,960). The primary outcome was a composite of ≥40% decline in estimated glomerular filtration rate (eGFR) or progression to end-stage renal disease within 2 years. A soft-voting ensemble of four ML algorithms (XGBoost, LightGBM, CatBoost, and Random Forest) was developed. RESULTS: The ensemble model demonstrated robust discrimination in internal validation (area under the receiver operating characteristic curve [AUROC], 0.939; 95% confidence interval [CI], 0.934-0.944), significantly exceeding the KFRE (AUROC, 0.879-0.884). External validation in the KoGES cohort showed comparable discrimination (AUROC, 0.859; 95% CI, 0.798-0.914) versus KFRE (four-variable AUROC, 0.882; 95% CI, 0.818-0.935). Shapley Additive exPlanations (SHAP) analysis identified baseline eGFR, serum creatinine, eGFR slope, albumin, and hemoglobin as key prognostic features, supporting a complementary framework using KFRE for community screening and the ML model for hospital-based risk stratification. CONCLUSION: The ensemble ML model accurately predicts short-term CKD progression in Korean patients. By incorporating longitudinal features and ensemble learning, it provides a precise alternative to Western-derived equations, particularly in tertiary care settings.

Chronic kidney failure

Systemic Proteome Profiling to Differentiate Primary Glomerular Diseases.

KEY POINTS: Plasma proteome profiling identified distinct signatures across biopsy-proven primary glomerular disease subtypes. An elastic net model using 93 proteins classified primary glomerular disease subtypes and controls, with external validation. Integrating proteomics with machine learning yields biologically interpretable insights in primary glomerular diseases. BACKGROUND: Primary GN is a heterogeneous group of kidney disorders where understanding of their pathophysiology remains incomplete. Despite the diagnostic potential of high-throughput proteomics, constrained proteomic depth and a reliance on binary comparisons have left the feasibility of using systemic signatures to differentiate multiple GN subtypes largely unexplored. METHODS: To identify protein signatures that noninvasively differentiate major primary glomerular disease subtypes and provide mechanistic insights, we performed large-scale systemic proteome profiling of 5416 plasma proteins via Olink Explore HT in a discovery cohort ( n =147) and an external validation cohort ( n =85) of Korean participants (mean age, 41±13 years; 46% female). The study population included patients with four GN subtypes-focal segmental glomerulosclerosis, IgA nephropathy, minimal change disease, and membranous nephropathy-alongside healthy controls. We developed a machine learning (ML) model using logistic regression with elastic net regularization to classify disease groups based on proteomic profiles and evaluated its performance in the independent validation cohort. RESULTS: Plasma proteome profiles were distinct among disease subtypes, emerging as a significant source of data variation independent of conventional markers such as eGFR or proteinuria levels. The ML model performed robustly in both the discovery and validation cohorts, achieving an area under the receiver operating characteristic curve >0.8 for differentiating minimal change disease, membranous nephropathy, and IgA nephropathy. The model, even without clinical information, correctly identified 93% of minimal change disease cases (14 of 15) and 63% of IgA nephropathy cases (20 of 32), but its performance was limited for focal segmental glomerulosclerosis, with only 21% of cases (three of 14) correctly classified. Functional analysis of key proteins highlighted distinct biologic pathways, such as hemostasis in minimal change disease. CONCLUSIONS: We identified distinct systemic proteome signatures for primary glomerular diseases, where disease subtype served as a major determinant of proteomic variance alongside conventional clinical markers. ML models demonstrated robust discriminatory performance for minimal change disease, membranous nephropathy, and IgA nephropathy, underscoring the potential for proteome-based classification.

Humans