PubMed HealthSearch

PubMed · 40577789

CSGL: chemical synthesis graph learning for molecule representation.

Abstract

MOTIVATION: Molecule representation learning (MRL) translates molecules into a real vector space, serving as input to downstream tasks in biology, chemistry, and computer science. This article introduces a chemical synthesis graph learning (CSGL) framework, which enhances MRL by considering both the atomic structures of molecules and their roles in chemical reactions through a hierarchical graph representation. Specifically, molecules are first modeled based on their molecular graphs, which capture atomic-level structural information. They are then further refined using a chemical synthesis graph, where nodes represent reactant and product molecule sets, and edges encode chemical transformations between reactants and products (e.g. changes in molecular structures). CSGL optimizes molecular embeddings of reactant and product nodes in a fashion that ensures the embeddings conform to a chemical balance constraint. RESULTS: Experimental results show that our method CSGL achieves strong performance on a variety of tasks, including product prediction, reaction classification, and molecular property prediction. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-2023/CSGL.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Anchen Li, Elena Casiraghi, Juho Rousu. 2025-07-01. CSGL: chemical synthesis graph learning for molecule representation.. https://doi.org/10.1093/bioinformatics%2Fbtaf355

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Enhancing detection of polygenic adaptation: a comparative study of machine learning and statistical approaches using simulated evolve-and-resequence data.

BACKGROUND: Detecting signals of polygenic adaptation remains a significant challenge in population genomics, as traditional methods often struggle to identify the associated subtle, multi-locus allele-frequency shifts. Here, we introduced and tested several novel approaches combining machine learning techniques with traditional statistical tests to detect polygenic adaptation patterns in time-series of allele frequency changes from whole genome data. We implemented a Naive Bayesian Classifier (NBC) and One-Class Support Vector Machines (OCSVM), and compared their performance against the classical Fisher's Exact Test (FET). Furthermore, we combined machine learning and statistical models (OCSVM-FET and NBC-FET), resulting in 5 competing approaches. The framework is mainly designed and validated for evolve-and-resequence (EaR) experimental designs, where defined selection pressures and temporal sampling are feasible, but might be applicable for certain natural experiments as well. RESULTS: Using a simulated dataset based on empirical C. riparius Pool-Seq data, we evaluated methods across evolutionary scenarios varying in generation, selection strength, and number of loci under selection. Our results demonstrate that the combined OCSVM-FET approach consistently outperformed competing methods, achieving the lowest false positive rate, highest area under the curve, and high accuracy. The performance peak aligned with what we term the 'late dynamic phase' of adaptation - the period after initial selection has occurred but before fixation - highlighting the method's sensitivity to ongoing selective processes. CONCLUSIONS: Furthermore, we emphasize the critical role of parameter tuning, balancing biological assumptions with methodological rigor. While broader applicability remains an important direction for future work, the present benchmarking is intentionally scoped to EaR experimental contexts.

Machine Learning

Improving insurance deduction identification: a hybrid artificial intelligence model using machine learning and expert systems.

PURPOSE: Financial challenges in healthcare systems worldwide, especially in low- and middle-income countries like Iran, have increased hospitals' reliance on insurance reimbursements. Unrecognized insurance deductions often cause severe financial shortages, making efficient deduction management crucial. This study aimed to design a hybrid intelligent system for identifying and predicting insurance deductions by combining machine learning and expert system frameworks. DESIGN/METHODOLOGY/APPROACH: A mixed-methods design was applied in four stages. First, a scoping review identified the causes and patterns of insurance deductions. Second, interviews with 15 insurance experts produced a validated checklist and a dataset from inpatient billing records. Third, using the CRISP-DM methodology, machine learning algorithms were developed and tested in SPSS Modeler alongside a fuzzy expert system developed in MATLAB. Finally, the model was validated using the holdout method. FINDINGS: Four categories of deduction drivers were identified: service provision, registration errors, document submission issues, and revenue conversion processes. The CHAID decision tree outperformed other algorithms with a 99% precision rate and the lowest Mean Absolute Error (9.43). A brief assessment of potential overfitting was conducted to ensure that the CHAID model's high accuracy was interpreted cautiously and supported by the validation results. The fuzzy expert system with validated rules was adaptable for deduction classification, especially for cases unsuitable for quantitative modeling. ORIGINALITY/VALUE: The hybrid model improves detection and prevention of deductions, offering actionable insights for hospital administrators, insurers, and policymakers. Its implementation can enhance hospital information systems, streamline claims processing, and optimize revenue management amid financial constraints.

Machine Learning

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning