PubMed HealthSearch

Biomedical subjects

Yixin Chen

Publications and source records attributed to Yixin Chen.

4 recordsLinked to original sources

Primer design through submodular function estimation.

MOTIVATION: Multiplex PCR-based enrichment is widely used in viral genome sequencing and pathogen surveillance. However, designing large sets of primers that maximize genome coverage while minimizing primer-primer interactions remains a major computational challenge. Existing methods such as SADDLE and Olivar use heuristics to optimize a Badness score for primer dimers but lack theoretical guarantees on solution quality. RESULTS: We introduce PRISM, a new framework that formulates multiplex primer design as a constrained submodular maximization problem. Our method defines an objective that balances genome coverage and dimer risk, and applies a local search algorithm with a constant-factor approximation guarantee. Evaluations on viral genome datasets demonstrate that PRISM consistently achieves lower Badness scores compared to PrimalScheme, Olivar, and primerJinn. These results highlight the scalability and theoretical rigor of submodular optimization in primer design. AVAILABILITY: PRISM is open-source and available at https://github.com/yhhan19/PRISM-new. The experimental data, scripts, and results used in this paper are archived on Figshare at https://doi.org/10.6084/m9.figshare.32806499.

Algorithms

BioMedGraphica: an all-in-one platform for joint textual biomedical prior knowledge and numeric graph generation.

MOTIVATION: Multiomics data analysis is essential for scientific discovery in precision medicine. However, translating analysis results of omics data analysis into novel scientific hypotheses remains a significant challenge. Human experts must manually review analysis results and generate new hypotheses based on extensive and interconnected biomedical prior knowledge, which is subjective and not scalable. While large language models can accelerate the discovery, their reasoning improves when grounded in structured, auditable, and comprehensive biomedical prior knowledge. However, biomedical knowledge is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of artificial intelligence systems to fully leverage biomedical data for scientific discovery. RESULTS: We developed BioMedGraphica, a novel all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified textual prior knowledge graph containing 2 306 921 entities and 27 232 091 relations. In addition, we present a novel textual-numeric graph (TNG) data structure concept, where textual information captures prior biological knowledge (e.g. transcription start sites, functions, mechanisms), numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data structure for developing novel graph analysis models. AVAILABILITY AND IMPLEMENTATION: The code is available at: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph database can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

Humans

Safety, tolerability, and efficacy of RIPK1 inhibitor, SAR443820, in amyotrophic lateral sclerosis (HIMALAYA): a multicentre, randomised, double-blind, placebo-controlled, phase 2 trial.

BACKGROUND: RIPK1, a protein regulating inflammatory signalling and cell death, is implicated in amyotrophic lateral sclerosis (ALS) pathophysiology. SAR443820 is a selective, oral, CNS-penetrant, reversible RIPK1 inhibitor. We aimed to evaluate the safety, tolerability, and efficacy of SAR443820 in participants with ALS. METHODS: This multicentre, randomised, double-blind, placebo-controlled, phase 2 trial was conducted at 63 clinical sites in 13 countries (Belgium, Canada, China, France, Germany, Italy, Japan, the Netherlands, Poland, Spain, Sweden, the UK, and the USA). Adults (aged 18-80 years) with a diagnosis of possible ALS, clinically probable ALS, clinically probable laboratory-supported ALS, or clinically definite ALS, in accordance with the revised El Escorial World Federation of Neurology criteria, were randomly assigned (2:1) by use of a stratified block design (blocks of three) to receive either 20 mg SAR443820 orally twice per day or matching placebo in the 24-week double-blind period. Randomisation was done centrally using interactive response technology and stratified by geographical region of trial site, region of ALS onset, use of riluzole, use of edaravone, and use of the combination of sodium phenylbutyrate and taurursodiol. Participants, care providers, investigators, and outcomes assessors were masked to trial intervention. The primary outcome was a change in ALS Functional Rating Scale Revised (ALSFRS-R) total score from baseline to week 24 and was calculated for all participants who had an ALSFRS-R total score available at baseline and at week 24. Safety analyses included all randomly assigned participants receiving one dose or more of trial intervention. This trial is registered with ClinicalTrials.gov (NCT05237284) and was terminated early. FINDINGS: Between April 13, 2022, and July 17, 2023, 397 participants were screened and 305 randomly assigned to SAR443820 (n=203) or placebo (n=102); six were excluded from the primary analysis due to missing baseline ALSFRS-R values. Mean age was 56·9 years (SD 11·5); 183 (60%) participants were male and 122 (40%) were female. Least squares mean change in ALSFRS-R from baseline to week 24 was -6·73 (95% CI -7·48 to -5·98) for SAR443820 group (n=169) and -6·32 (-7·36 to -5·27) for placebo group (n=87). There was no statistically significant difference between the study groups (least squares mean difference -0·41 [95% CI -1·71 to 0·88]). Participants in the SAR443820 group had higher incidence of adverse events (171 [85%] of 202 vs placebo 80 [78%] of 102) and treatment discontinuations (28 [14%] of 202 vs placebo five [5%] of 102), with elevated hepatic enzymes being the most common cause. Nine deaths occurred in the double-blind period (seven [3%] of 202 in the SAR443820 group and two [2%] of 102 in the placebo group); none was attributed to SAR443820. INTERPRETATION: SAR443820 did not show clinical benefit and was associated with higher hepatic enzyme increase, indicating that further clinical development of SAR443820 in ALS is not warranted. FUNDING: Sanofi.

Humans

BioMedGraphica: An All-in-One Platform for Joint Textual Biomedical Prior Knowledge and Numeric Graph Generation.

Multi-omic data analysis is essential for scientific discovery in precision medicine. However, translating statistical results of omic data analysis into novel scientific hypothesis remains a significant challenge. Human experts must manually review analysis results and generate new hypothesis based on extensive and inter-connected biomedical prior knowledge, which is subjective and not scalable. While large language models (LLMs) can accelerate the discovery, their reasoning improves when grounded in structured, auditable and comprehensive biomedical prior knowledge. Biomedical knowledge, however, is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of AI systems to fully leverage biomedical data for scientific discovery. To address these challenges, we developed BioMedGraphica , an all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified knowledge graph containing 2,306,921 entities and 27,232,091 relations. In addition, to the best of our knowledge, this is the first work to propose a novel Textual-Numeric Graph (TNG) data-structure for multi-omics data analysis. In TNG, textual information captures prior biological knowledge (e.g., transcription start sites, functions, mechanisms), while numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data-structure for the development of graph foundation models, with the potential to improve prediction performance and interpretability, while also augmenting LLMs by supplying graph-structured mechanistic context to strengthen reasoning. The details for BioMedGraphica code can be accessed by github link: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph data can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

biomedical knowledge graph