PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Chromatin structures from integrated AI and polymer physics model.

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure from indirect measures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Chromatin

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged ≥65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans

Digital pathology and spatial omics in steatohepatitis: Clinical applications and discovery potentials.

Steatohepatitis with diverse etiologies is the most common histological manifestation in patients with liver disease. However, there are currently no specific histopathological features pathognomonic for metabolic dysfunction-associated steatotic liver disease, alcohol-associated liver disease, or metabolic dysfunction-associated steatotic liver disease with increased alcohol intake. Digitizing traditional pathology slides has created an emerging field of digital pathology, allowing for easier access, storage, sharing, and analysis of whole-slide images. Artificial intelligence (AI) algorithms have been developed for whole-slide images to enhance the accuracy and speed of the histological interpretation of steatohepatitis and are currently employed in biomarker development. Spatial biology is a novel field that enables investigators to map gene and protein expression within a specific region of interest on liver histological sections, examine disease heterogeneity within tissues, and understand the relationship between molecular changes and distinct tissue morphology. Here, we review the utility of digital pathology (using linear and nonlinear microscopy) augmented with AI analysis to improve the accuracy of histological interpretation. We will also discuss the spatial omics landscape with special emphasis on the strengths and limitations of established spatial transcriptomics and proteomics technologies and their application in steatohepatitis. We then highlight the power of multimodal integration of digital pathology augmented by machine learning (ML)algorithms with spatial biology. The review concludes with a discussion of the current gaps in knowledge, the limitations and premises of these tools and technologies, and the areas of future research.

Humans

Improving risk indexes for Alzheimer's disease and related dementias for use in midlife.

Knowledge of a person's risk for Alzheimer's disease and related dementias (ADRDs) is required to triage candidates for preventive interventions, surveillance, and treatment trials. ADRD risk indexes exist for this purpose, but each includes only a subset of known risk factors. Information missing from published indexes could improve risk prediction. In the Dunedin Study of a population-representative New Zealand-based birth cohort followed to midlife (N = 938, 49.5% female), we compared associations of four leading risk indexes with midlife antecedents of ADRD against a novel benchmark index comprised of nearly all known ADRD risk factors, the Dunedin ADRD Risk Benchmark (DunedinARB). Existing indexes included the Cardiovascular Risk Factors, Aging, and Dementia index (CAIDE), LIfestyle for BRAin health index (LIBRA), Australian National University Alzheimer's Disease Risk Index (ANU-ADRI), and risks selected by the Lancet Commission on Dementia. The Dunedin benchmark was comprised of 48 separate indicators of risk organized into 10 conceptually distinct risk domains. Midlife antecedents of ADRD treated as outcome measures included age-45 measures of brain structural integrity [magnetic resonance imaging-assessed: (i) machine-learning-algorithm-estimated brain age, (ii) log-transformed volume of white matter hyperintensities, and (iii) mean grey matter volume of the hippocampus] and measures of brain functional integrity [(i) objective cognitive function assessed via the Wechsler Adult Intelligence Scale-IV, (ii) subjective problems in everyday cognitive function, and (iii) objective cognitive decline measured as residualized change in cognitive scores from childhood to midlife on matched Weschler Intelligence scales]. All indexes were quantitatively distributed and proved informative about midlife antecedents of ADRD, including algorithm-estimated brain age (β's from 0.16 to 0.22), white matter hyperintensities volume (β's from 0.16 to 0.19), hippocampal volume (β's from -0.08 to -0.11), tested cognitive deficits (β's from -0.36 to -0.49), everyday cognitive problems (β's from 0.14 to 0.38), and longitudinal cognitive decline (β's from -0.18 to -0.26). Existing indexes compared favourably to the comprehensive benchmark in their association with the brain structural integrity measures but were outperformed in their association with the functional integrity measures, particularly subjective cognitive problems and tested cognitive decline. Results indicated that existing indexes could be improved with targeted additions, particularly of measures assessing socioeconomic status, physical and sensory function, epigenetic aging, and subjective overall health. Existing premorbid ADRD risk indexes perform well in identifying linear gradients of risk among members of the general population at midlife, even when they include only a small subset of potential risk factors. They could be improved, however, with targeted additions to more holistically capture the different facets of risk for this multiply determined, age-related disease.

Alzheimer’s disease

Inferring Gene Regulatory Networks in Stem Cells: Methods and Applications.

Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.

Gene Regulatory Networks

Spatiotemporal genomic analysis and risk assessment of the plasmids carrying blaOXA-48-like genes based on a large-scale international dataset.

BACKGROUND: The spread of OXA-48-like carbapenemases represents a major public health challenge. Although previous studies have investigated OXA-48-like carbapenemases risk factors, nosocomial dissemination, and plasmid dynamics, an integrated plasmid-centered framework combining complete plasmid mining, transmission-unit analysis, phylogenetic reconstruction, and machine learning-based risk assessment remains limited. METHODS: We systematically collected 747 complete plasmid sequences carrying blaOXA-48-like genes from the NCBI database, establishing the largest collections of complete plasmid sequences to date. Using an integrative framework of population genomics, phylogenetic dating, and machine learning, this study aimed to characterize the dissemination patterns, plasmid replicon diversity, transmission units, mobile genetic elements, co-resistance profiles, and risk classification of these plasmid. RESULTS: Plasmids carrying blaOXA-48-like genes were detected across 50 countries on six continents, with blaOXA-48 predominating in Europe, blaOXA-181 in South Asia, and blaOXA-232 largely in Asia. IncL and ColKP3/IncX3 replicons, together with Tn1999.2 and other MGEs, were central drivers of plasmid maintenance and spread. Sixteen transmission units were defined, with AA068_Cluster3 estimated to have originated in the Netherlands around 2005 before expanding to Europe, the Middle East, Asia, and North America. Co-resistance analyses revealed frequent modules involving aminoglycoside and quinolone resistance, with qnrS1 and aph(3'')-Ib most prevalent. Notably, high-risk transposon structures were often identified in non-clinical environments, underscoring their cross-ecological transmission potential. Machine learning-based classification models showed good internal performance for predefined composite-risk categories, with plasmid mobility, clinical/non-clinical source composition, and host background contributing to the classification results. CONCLUSIONS: This study provides a large-scale plasmid-centered genomic analysis of publicly available complete plasmid sequences carrying blaOXA-48-like genes, integrating transmission-unit inference, phylogeographic reconstruction, mobile genetic element and co-resistance profiling, and composite genomic risk stratification. This gene-centered framework may support future One Health-oriented antimicrobial resistance surveillance and prioritization of plasmids with higher dissemination and resistance potential.

Plasmids

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies

Causal associations between hormone replacement therapy and brain structure: Evidence from large-scale Mendelian randomization and double machine learning.

BACKGROUND: Hormone replacement therapy (HRT) is widely prescribed for the management of hormone deficiency, particularly during menopause, yet its causal effects on human brain structure remain incompletely understood. Observational studies have reported heterogeneous associations, underscoring the need for robust causal inference. METHODS: We applied an integrated causal framework combining two-sample Mendelian Randomization (MR) and Double Machine Learning (DML) to evaluate the effects of four HRT-related exposures-age at initiation, age at cessation, ever-use of HRT, and a composite medication-based phenotype-on 1366 brain imaging-derived phenotypes from the UK Biobank. Genetic instruments were derived from large-scale GWAS summary statistics, and causal estimates were validated using non-parametric DML models with cross-fitting and performance evaluation. RESULTS: Genetic instruments for age at HRT initiation, age at cessation, and ever-use of HRT were strong (median F-statistics 16.29-36.66). MR analyses identified a causal association between later initiation of HRT and lower orientation dispersion in the right inferior cerebellar peduncle (ubm-a-542; primary finding, no pleiotropy detected). An additional association with the left tapetum FA (ubm-a-243) was identified but exhibited significant directional horizontal pleiotropy (MR-Egger intercept P = 0.001) and is excluded from primary conclusions (Supplementary Note S2). Later cessation of HRT was associated with increased cortical thickness in the left middle occipital gyrus, reduced surface area in the left frontopolar cortex, and increased orientation dispersion in the splenium of the corpus callosum. Ever-use of HRT was causally linked to larger volumes of the right inferior frontal gyrus and right nucleus accumbens. These associations were corroborated by independent DML validation, which provided causally debiased estimates robust to high-dimensional confounding. Results for ukb-b-8080 (median F = 1.45) are provided in Supplementary Note S1 only; weak-instrument bias precludes causal inference. CONCLUSIONS: This study provides genetic-instrument-based and machine-learning-validated evidence for causal associations between HRT exposure-particularly its timing and lifetime use-and specific features of human brain structure, including white-matter microarchitecture, cortical thickness, and regional brain volume. These findings are FDR-controlled within exposures and independently replicated by DML, but require replication in external neuroimaging GWAS cohorts to establish definitive causal conclusions. They highlight the neurobiological relevance of sex steroid exposure and inform future research on brain aging and personalized hormone-based interventions.

Humans

Serum Proteomic Profiling Implicates a Dysregulated Neurohormonal-Inflammatory Axis in Post-Fontan Sinus Tachycardia.

BACKGROUND: Postoperative sinus tachycardia is a poorly understood complication following the Fontan procedure. The molecular signaling cascades triggering acute tachycardia remain uncharacterized, limiting therapeutic innovation. Here, we present a retrospective study leveraging serum proteomics and machine learning to identify the molecular drivers of postoperative Fontan sinus tachycardia. METHODS: We integrated a clinically relevant ovine Fontan model with continuous telemetric heart rate monitoring and human patient data. Serum proteomics coupled with least absolute shrinkage and selection operator and Boruta machine learning algorithms were used to identify protein panels predictive of postoperative sinus tachycardia. Cross-species validation was performed by comparing proteomic signatures from sheep and pediatric patients undergoing Glenn or Fontan surgery. RESULTS: Ovine Fontan animals demonstrated significant heart rate elevation beginning on postoperative day 1, peaking at postoperative day 3 (159.4±11.7 bpm versus preoperative, 105.3±10.5 bpm; P=0.0002), before trending toward baseline by postoperative day 10. This pattern was mirrored in human patients with a more modest magnitude. Surgical controls did not exhibit tachycardia. The principal component most correlated with heart rate (principal component 1: r=0.78, P=2.2×10-4) was enriched for inflammatory and neural pathways. The Boruta algorithm identified an 11-protein panel with strong predictive power (area under the receiver operating characteristic curve, 0.963). Cross-species comparison demonstrated that angiotensinogen, angiotensin-converting enzyme, and pentraxin 3 were similarly dysregulated in both species postoperatively. CONCLUSIONS: This study provides molecular evidence implicating a dysregulated neurohormonal-inflammatory axis in acute postoperative Fontan sinus tachycardia and establishes a foundation for developing targeted diagnostics and therapeutics for this complication.

Animals

Synthetic DNA barcodes identify singlets in scRNA-seq datasets and evaluate doublet algorithms.

Single-cell RNA sequencing (scRNA-seq) datasets contain true single cells, or singlets, in addition to cells that coalesce during the protocol, or doublets. Identifying singlets with high fidelity in scRNA-seq is necessary to avoid false negative and false positive discoveries. Although several methodologies have been proposed, they are typically tested on highly heterogeneous datasets and lack a priori knowledge of true singlets. Here, we leveraged datasets with synthetically introduced DNA barcodes for a hitherto unexplored application: to extract ground-truth singlets. We demonstrated the feasibility of our framework, "singletCode," to evaluate existing doublet detection methods across a range of contexts. We also leveraged our ground-truth singlets to train a proof-of-concept machine learning classifier, which outperformed other doublet detection algorithms. Our integrative framework can identify ground-truth singlets and enable robust doublet detection in non-barcoded datasets.

Algorithms

Challenges and future directions in AI-driven biomaterials for microbiome-associated oral infectious diseases: A systematic review.

Oral biofilm-induced antimicrobial resistance is the core pathogenic mechanism of microbiome-associated oral infectious diseases (dental caries, periodontitis, peri-implantitis, and endodontic infection). Traditional therapies and biomaterials are limited by poor biofilm penetration, drug resistance induction, single functionality, and inadequate adaptation to dynamic oral microenvironmental changes (e.g., pH fluctuations, salivary rinsing, masticatory stimulation). Artificial intelligence (AI) has transformed the field by integrating materials science, microbiology, and stomatology data. Via machine learning, deep learning, and multi-physics simulation, AI optimizes biomaterial physicochemical properties, decodes microenvironmental signals, constructs precise sensing-response loops, and supports the full chain of material design, performance prediction, and action simulation, advancing treatment from empirical intervention to precision regulation. This systematic review retrieved literature from PubMed, Embase, and Web of Science (January 2016-January 2026) using keywords across three dimensions: AI, biomaterials, and oral microbiome. Following inclusion/exclusion criteria, 99 articles were included. It elaborates on five core mechanisms of AI-driven oral biomaterials (precise oral microbiome analysis, targeted material design/optimization, performance prediction/simulation, targeted delivery/intervention, effect evaluation/dynamic regulation), analyzes their applications in microbiome-targeted biomaterial research and development (R&D) and clinical practice for the four major oral infectious diseases, addresses technical bottlenecks (insufficient targeting specificity and precision of biomaterials, poor stability and durability in complex oral microenvironments, inadequate biofilm disruption capacity, and clinical translation obstacles), and proposes future directions (multimodal design to enhance targeting specificity, structural and component optimization to improve stability/durability, development of multi-mechanism synergistic biofilm disruption strategies, strengthening translational research for clinical application, and deep integration of AI in the full chain of biomaterial R&D). This work provides comprehensive theoretical and practical support for the R&D, optimization, and clinical translation of AI-driven microbiome-targeted oral biomaterials.

Humans

Deep learning-based multimodal pathogenomics integration for precision cancer prognosis.

BACKGROUND: Recent studies have revealed valuable prognostic insights in haematoxylin and eosin (H&E)-stained histological sections and transcriptomic profiles, suggesting potential applications in machine learning. However, existing methods lack sufficient intra- and inter-modal interactions, and face challenges in clinical validation due to incomplete multimodal data. METHODS: We proposed PathoGems (PathoGenomics-based integrative survival prediction), a weakly-supervised, interpretable multimodal learning framework that integrates histology and genomic profiles for precise cancer prognosis prediction. To evaluate the robustness of PathoGems, we initially curated a dataset of 1965 cases across four cohorts from The Cancer Genome Atlas (TCGA), including breast, colorectal, glioblastoma, and esophageal cancers. For external validation, PathoGems was further evaluated on four independent cohorts, consisting of 76 breast cancer and 41 esophageal squamous cell carcinoma cases from Zhejiang Cancer Hospital, as well as 102 colorectal cancer and 58 glioblastoma cases from the Clinical Proteomic Tumor Analysis Consortium (CPTAC). RESULTS: PathoGems effectively stratified patients into favorable and unfavorable risk groups, revealing significant differences in histological patterns, genomic features, and overall survival (log-rank test, p&#x2009;<&#x2009;0.05). Moreover, the model&#x2019;s predictions are further supported by visualization and transcriptomic analysis, enhancing interpretability and reliability. CONCLUSIONS: By fusing histological and clinicogenomic multimodal models, PathoGems will provide a solid foundation for developing an innovative tool that aids clinicians in making informed decisions and selection personalized treatment strategies for cancer patients.

Humans

A three-metabolite microbiota-associated signature for early risk stratification of gestational diabetes mellitus.

BACKGROUND: Gestational diabetes mellitus (GDM) is associated with adverse pregnancy outcomes and long-term metabolic and cardiovascular risk. However, oral glucose tolerance testing at 24-28 gestational weeks limits early risk stratification. Gut microbiota-associated metabolites may reflect early metabolic abnormalities, including those relevant to cardiometabolic health, but robust early-pregnancy biomarkers remain limited. METHODS: We conducted a multicenter nested case-control and prospective study involving 2,693 pregnant women. Untargeted metabolomics and metagenomics were integrated to identify GDM-associated metabolites and gut microbial alterations. Three consistently dysregulated metabolites, 3-hydroxydecanoic acid, &#x3b3;-Glu-Leu, and propionic acid, were quantified by targeted LC-MS/MS. Candidate algorithms were compared using repeated 10-fold cross-validation, and a final generalized linear model was externally and prospectively validated. RESULTS: Women who later developed GDM showed an adverse early-pregnancy metabolic profile, including higher BMI, triglycerides, and platelet count. Untargeted metabolomics identified 14 persistently altered metabolites enriched in energy, oxidative stress, and amino acid metabolism pathways. Metagenomics revealed taxonomic restructuring and coordinated microbiota-metabolite associations. The three-metabolite model achieved AUCs of 0.838 (95% CI, 0.791-0.885) in training, 0.840 (95% CI, 0.769-0.911) in internal validation, 0.955 (95% CI, 0.925-0.985) and 0.917 (95% CI, 0.875-0.958) in two external cohorts, and 0.969 (95% CI, 0.937-1.000) in the prospective cohort. CONCLUSION: Early microbiota-associated metabolic dysregulation is detectable before routine GDM diagnosis. This compact three-metabolite panel may support early GDM risk stratification and provides metabolic evidence relevant to broader cardiometabolic risk assessment in pregnancy.

Humans

Bridging genotype, phenotype, and clinical insight: the role of multi-omics in cardiovascular disease.

INTRODUCTION: It is increasingly evident that the multifactorial nature of cardiovascular disease requires the combination of different omics approaches for improving our mechanistic understanding, identifying novel drug targets, and developing accurate diagnostic, predictive, and prognostic biomarker panels. AREAS COVERED: We review the current state and the potential of multi-omics in cardiovascular disease, with a specific focus on plasma-, spatial-, and single-cell approaches. We discuss lipidomics as a genotype&#x2011;to&#x2011;phenotype bridge, the utility of remote longitudinal monitoring via microsampling/dried blood spots, and emerging clinical&#x2011;trial integrations of multi-omics approaches. We outline critical gaps in standardization and how to overcome these, pre&#x2011;analytical challenges and constraints that are often neglected, and data&#x2011;integration methods spanning from canonical correlation analysis to modern machine learning approaches. EXPERT OPINION: Multi&#x2011;omics can shape cardiovascular care by identifying drug targets in diseased tissue and by yielding small, usable biomarker panels.

Humans

A Multi-omics Regulated Cell Death Framework Defines Immune Phenotypes and Guides Precision Therapy in Colorectal Cancer.

Colorectal cancer (CRC) is molecularly and immunologically heterogeneous, contributing to variable treatment response. Because regulated cell death (RCD) intersects with tumor metabolism, immune regulation, and therapeutic susceptibility, we built an RCD-centered framework for CRC stratification. Multi-cohort transcriptomic data were used to infer RCD subtypes with non-negative matrix factorization (NMF) and non-negative least squares (NNLS). Genomic, bulk RNA-seq, single-cell RNA-seq, and spatial transcriptomic datasets were integrated to characterize subtype-associated biology. Machine-learning models were developed for immunotherapy response and survival-risk estimation. Candidate compounds were screened by GDSC2-based drug-sensitivity modeling and molecular docking, and FSTL3 was functionally assessed in vitro. The framework separated CRC samples into two RCD-related phenotypes resembling immune-hot and immune-cold states. RCD1 showed immune activation and higher mutational burden, whereas RCD2 showed immune-suppressed features, intratumoral heterogeneity, and aggressive biology. RCD-associated signatures showed potential for predicting immunotherapy response and survival risk. Dasatinib was prioritized for immune-cold, high-risk tumors, with preliminary evidence supporting its activity in CRC cells, while functional assays suggested a role for FSTL3 in growth, invasion, epithelial-mesenchymal transition, and apoptosis regulation. These findings suggest that RCD-based multi-omics analysis may refine CRC stratification and help generate therapeutic hypotheses.

Colorectal cancer

Efforts towards a precision medicine approach in juvenile idiopathic arthritis.

Juvenile idiopathic arthritis (JIA) is the commonest group of childhood arthritides. Despite the availability of advanced therapeutics, many children and young people (CYP) with JIA experience disease flares, and in some, chronic joint damage. Tailoring treatment based on unique biological profiles would benefit CYP with JIA given their variable clinical presentation and disease course. To date, biomarkers to predict treatment response are lacking. With advances in single cell technologies, we are now able to profile the genes and proteins of target tissues at unprecedented resolution to define the biological basis of disease and guide novel treatment approaches. The complex analyses and combination of biological and clinical outcome data from large datasets across disease phenotypes have become possible with the development of computational and machine learning methods. Here, we summarize the strategies to integrate data through multimodal based approaches to maximize precision medicine and research priorities for CYP with JIA.

Humans

ALPAR: automated learning pipeline for antimicrobial resistance.

SUMMARY: The field of machine learning in antimicrobial resistance (AMR) research has experienced rapid growth, fueled by advancements in high-throughput genome sequencing and the growing capacity of computational resources. However, the complexity and lack of standardized data preparation and bioinformatic analyses present significant challenges, especially for newcomers to the domain. In response to these challenges, we introduce ALPAR (Automated Learning Pipeline for Antimicrobial Resistance), a comprehensive AMR data analysis tool covering the entire process from processing of raw genomic data to training machine learning models to interpretation of results. Our method relies on a reproducible pipeline that integrates widely used bioinformatics tools, presenting a simplified, automatic workflow specifically tailored for single-reference AMR analysis. Accepting genomic data in the form of FASTA files as input, ALPAR facilitates the generation of machine learning-ready data tables and both the training of machine learning and the execution of genome-wide association studies (GWAS) experiments. Additionally, our tool offers supplementary functionalities such as phylogeny-based analysis of the distribution of mutations, enhancing its utility for researchers. The tool has also proven its performance in competitive benchmarks, winning the 2024 CAMDA Anti-Microbial Resistance Prediction Challenge and placing third in the 2025 edition. AVAILABILITY AND IMPLEMENTATION: ALPAR is open-source and freely accessible via GitHub (https://github.com/kalininalab/ALPAR). The pipeline is fully reproducible and can be easily installed as a Conda package (https://anaconda.org/kalininalab/ALPAR).

Machine Learning

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning