PubMed HealthSearch

SEARCH · PubMed Health

Results for “data curation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Evaluating the need for late angiography for complete obliteration of AVMs after endovascular treatment: collaborative AVM center experience and a systematic review.

Brain arteriovenous malformations (bAVMs) can be treated curatively by endovascular embolization. However, limited data exist on late recanalization rates. We investigated the rate of late bAVM recanalization following complete endovascular obliteration, confirmed by primary control angiography.We performed a single-center retrospective cohort study on late recurrences in adult patients with bAVMs after complete endovascular obliteration at a Radboud - Isala - MUMC+ (RIM) collaborative AVM center in the Netherlands between 2014 and 2022.Additionally, we conducted a systematic review to evaluate the rate of late recanalization following complete endovascular obliteration, confirmed on primary control angiography, in adults with bAVMs. The protocol for this review was registered in PROSPERO (CRD42024546875).Our retrospective study revealed 42 adult patients treated by endovascular means only; 90.5% (38 patients with mean age of 48.1 years) had confirmed complete obliteration. Both primary (6 months post-treatment) and secondary (more than 1 year post-treatment) angiographic control confirmed complete obliteration in 21 of 23 patients with complete follow-up. Mean follow-up was 47.8 months. Two late recurrences (9.5%) were detected at 5- and 6-years' follow-up imaging.Our systematic review included two studies encompassing a total of 19 patients with mean angiographical follow-up of 20.8 months. There were no late recurrences.Late bAVM recurrence after endovascular treatment with proven complete obliteration may be underestimated owing to limited long-term follow-up. Our findings suggest that after a 6-month angiographic confirmed obliteration, a 5-year angiographic imaging control is justified.

Humans

A corpus of GA4GH phenopackets: Case-level phenotyping for genomic diagnostics and discovery.

The Global Alliance for Genomics and Health (GA4GH) Phenopacket Schema was released in 2022 and approved by ISO as a standard for sharing clinical and genomic information about an individual, including phenotypic descriptions, numerical measurements, genetic information, diagnoses, and treatments. A phenopacket can be used as an input file for software that supports phenotype-driven genomic diagnostics and for algorithms that facilitate patient classification and stratification for identifying new diseases and treatments. There has been a great need for a collection of phenopackets to test software pipelines and algorithms. Here, we present Phenopacket Store. Phenopacket Store v.0.1.19 includes 6,668 phenopackets representing 475 Mendelian and chromosomal diseases associated with 423 genes and 3,834 unique pathogenic alleles curated from 959 different publications. This represents the first large-scale collection of case-level, standardized phenotypic information derived from case reports in the literature with detailed descriptions of the clinical data and will be useful for many purposes, including the development and testing of software for prioritizing genes and diseases in diagnostic genomics, machine learning analysis of clinical phenotype data, patient stratification, and genotype-phenotype correlations. This corpus also provides best-practice examples for curating literature-derived data using the GA4GH Phenopacket Schema.

Humans

HoloFoodR: a statistical programming framework for holo-omics data integration workflows.

SUMMARY: Holo-omics is an emerging research area that integrates multi-omic datasets from the host organism and its microbiome to study their interactions. Recently, curated and openly accessible holo-omic databases have been developed. The HoloFood database, for instance, provides nearly 10 000 holo-omic profiles for salmon and chicken under controlled treatments. However, bridging the gap between holo-omic data resources and algorithmic frameworks remains a challenge. Combining the latest advances in statistical programming with curated holo-omic data sets can facilitate the design of open and reproducible research workflows in the emerging field of holo-omics. AVAILABILITY AND IMPLEMENTATION: HoloFoodR R/Bioconductor package and the source code are available under the open-source Artistic License 2.0 at the package homepage https://doi.org/10.18129/B9.bioc.HoloFoodR.

Software

MelanoDB: A dataset of clinical and molecular features of patients with advanced melanoma treated with MAPK inhibitors.

MAPK inhibitors (MAPKi) have revolutionized the treatment of patients with advanced melanoma. However, primary and acquired resistance mechanisms limit their efficacy. Predicting MAPKi response from the tumor baseline features remains challenging due to the limited size of patient cohorts. Therefore, we collected data from nine different patient cohorts (total n = 417 patients with advanced melanoma treated with MAPKi) to identify clinical and molecular features. Our curated dataset, named MelanoDB, includes whole or partial exome sequencing data for 191 patients, copy number alteration information for 66 patients, and gene expression data for 132 patients. We provide a web application to explore the integrated dataset and data distribution across the collected studies, and we share this dataset with the scientific community according to the Findable, Accessible, Interoperable, Reusable (FAIR) principles.

Humans

Multi-omics identification of therapeutic targets of compound sappan decoction in hepatocellular carcinoma.

BACKGROUND: Compound sappan decoction (CSD) is a multi-herbal traditional Chinese medicine formulation with clinical relevance in hepatocellular carcinoma (HCC). However, its therapeutic mechanisms remain unclear. METHODS: Bioactive compounds of CSD were identified and standardized using pharmacological and chemical databases. Potential targets were predicted via multiple target inference platforms. HCC-related genes were curated from comprehensive disease databases. Summary-data-based Mendelian randomization (SMR) was conducted to infer causal relationships between compound targets and HCC risk using large-scale quantitative trait loci (QTL) datasets and HCC genome-wide association study data. Colocalization analysis, protein-protein interaction (PPI) network construction, and GO/KEGG enrichment were performed on SMR-identified targets. Molecular docking evaluated binding affinities of representative compounds to prioritized targets. RESULTS: A total of 784 overlapping genes between predicted CSD targets and HCC-related genes were subjected to SMR analysis. Among these, 22 targets were significantly associated with HCC risk based on transcriptomic or proteomic QTLs and showed colocalization evidence. Notably, four targets (ADRB2, APOE, SYK, and PGF) were supported by both replication in an independent cohort and strong colocalization. These 22 targets were enriched in apoptosis, PI3K-Akt signaling, redox metabolism, and detoxification pathways. PPI analysis revealed central hubs including MMP9, BCL2, CASP1, and MCL1. Molecular docking demonstrated strong binding of APOE to quercetin, PGF to luteolin-7-olate, and SYK to kaempferol. CONCLUSIONS: CSD may exert therapeutic effects on HCC through modulation of genetically validated targets involved in tumor progression, inflammation, and metabolic reprogramming, supporting its potential clinical utility as an adjunctive treatment strategy. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s12672-026-04740-8.

Caesalpinia

Multi-omics integration and colocalization analyses prioritize candidate molecular loci associated with hypothermia.

BACKGROUND: Hypothermia is a life-threatening condition lacking specific pharmacological treatments. This study aimed to prioritize genetically supported molecular loci associated with hypothermia and to explore their pharmacological tractability using multi-omics data. METHODS: Initially, 2532 druggable genes were curated from the Drug-Gene Interaction Database and established literature. These were cross-referenced with cis-eQTL and cis-pQTL datasets, encompassing 870,655 and 114,281 SNPs for blood, respectively, alongside 2379 shared SNPs across adipose, skeletal muscle, and heart tissues. Matched instrumental variables were integrated with hypothermia GWAS summary statistics for two-sample Mendelian randomization (MR) and Bayesian colocalization. Transcriptomic differential expression analysis (DEA) was subsequently conducted as an exploratory analysis of cold-exposure-associated expression changes. Database-derived compound annotations were systematically re-evaluated according to target specificity, established pharmacological mechanism, and concordance with the direction of the MR estimates. RESULTS: Among 671 gene-level MR tests, 36 genes reached nominal significance, whereas only ABCC8 remained significant after FDR correction. Colocalization was evaluable for 8 of these 36 genes, and 4 loci (COL18A1, SLC1A7, ADIPOQ, and MERTK) met the prespecified PP.H4>0.90 threshold. The remaining 28 loci were not evaluable because sufficient overlapping regional variants were unavailable after harmonization. Transcriptomic analysis identified altered expression of SLC1A3 and SLCO4A1 under cold exposure, although these findings did not directly validate the colocalization-supported loci. Re-evaluation of database-derived compound annotations did not identify any direct, selective, and directionally concordant drug-repurposing candidate for hypothermia. CONCLUSIONS: COL18A1, SLC1A7, ADIPOQ, and MERTK showed colocalization support among the 8 evaluable nominal MR-associated loci. Because colocalization coverage was limited, these genes should be regarded as preliminary candidate loci rather than established therapeutic targets. The pharmacological annotations were indirect, non-selective, unsupported, or directionally inconsistent and should be interpreted solely as hypothesis-generating information.

Bayesian colocalization

Sec and Tat Mediated Secretion Safeguards Mycobacterium tuberculosis Membrane Homeostasis.

Protein secretion is essential for the growth and virulence of Mycobacterium tuberculosis, yet the organization and function of its secretion pathways remain poorly understood. We reviewed the existing literature, combined it with systematic queries, and finalized annotations based on experimental data and computational predictions to compile a curated list of 92 secretory components and 198 reactions involved in Sec, twin-arginine translocation (Tat), and ESX pathways. Using CRISPRi, targeted depletion of SecA1 or TatAC impaired both in vitro growth and ex vivo survival. Label-free quantitative secretome analysis revealed decreased export of substrates dependent on SecA1 and TatAC, with enrichment of cytosolic proteins in culture filtrates, indicating increased membrane dysbiosis. Membrane proteomics showed elevated levels of proteins engaged in intermediary and lipid metabolism, while proteins associated with the cell wall and cell processes decreased, suggesting weakened membrane integrity. Loss of SecA1 or TatAC increased membrane permeability, with the effect being more pronounced in the case of TatAC, and caused structural abnormalities seen under electron microscopy. Overall, our integrated multi-omics and functional genetics studies demonstrate that the SecA1 and Tat pathways are essential for maintaining membrane homeostasis in Mycobacterium tuberculosis. These results suggest that essential secretory proteins may be promising targets for therapeutic intervention.

Mycobacterium tuberculosis

Programmatic access to ICTV virus taxonomy through a public ontology API.

BACKGROUND: The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. FINDINGS: To address this, we developed a public and sustainable solution leveraging ontology-based APIs. All available ICTV Master Species List (MSL) releases, from MSL1 to MSL41, were transformed into a unified, semantically structured ontology comprising more than 195,000 current and historical entities and deployed through the Ontology Lookup Service (OLS). The ontology is automatically rebuilt and republished whenever a new MSL release becomes available. Complementary ICTV-NCBI mappings and helper libraries support integration into downstream systems. CONCLUSIONS: Together, these resources enable, for the first time, public programmatic retrieval of current and historical ICTV taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints, including resolution of former taxonomic terms to their current accepted taxon or taxa and retrieval of taxon histories across releases. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API

Programmatic access to ICTV virus taxonomy through a public ontology API.

The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API

CholeraSeq: a comprehensive genomic pipeline for cholera surveillance and near real-time outbreak investigation.

SUMMARY: Next Generation Sequencing is widely deployed in cholera-endemic regions, yet an end-to-end reproducible pipeline that unifies read QC, filtering, reference mapping, variant calling/annotation, recombination screening, and extraction of parsimony informative sites/variant codons, phylogenetic inference for downstream phylodynamic and epidemiological analyses have been lacking, slowing outbreak investigation and public health response. CholeraSeq is a high-throughput genomics pipeline for cholera genomic surveillance. It ingests consensus genomes, short read sequence data, draft assemblies, and scales seamlessly from local to cloud environments. To accelerate epidemiological context placement of new outbreak strains, we provide a curated ready-to-use core genome alignment compiled from public data, enabling flexible, fast, integration of new samples for outbreak investigations. AVAILABILITY AND IMPLEMENTATION: CholeraSeq is freely available on the GitHub platform https://github.com/CERI-KRISP/CholeraSeq. CholeraSeq is implemented in Nextflow with a modular design building upon the nf-core community standards.

Cholera

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors

Project Sickle Cure: A Prospective, International Observational Study of Hematopoietic Cell Transplantation for Sickle Cell Disease.

BACKGROUND: Sickle cell disease (SCD) is a chronic and life-limiting hemoglobin and systemic vascular disease. While over 1000 people have undergone hematopoietic cell transplantation (HCT) over the last 40 years, long-term disease-specific and health-related quality of life data are lacking. The American Society of Hematology 2021 Guidelines for SCD emphasized the need for more detailed registry data to inform patients and providers with decision-making and practice recommendations. PROCEDURES: In January 2021, the Sickle Cell Transplant Advocacy and Research Alliance (STAR) launched Project Sickle Cure (PSC). This multi-center, prospective study of patients who have undergone HCT for SCD includes baseline demographics and SCD-specific post-HCT outcomes, serial neurocognitive testing, health-related quality of life measures, health equity evaluations, a neuroimaging bank, detailed evaluation of neurologic status pre- and post-transplant, and chronic pain evaluation. A biorepository is in the planning stage of development. RESULTS: As of November 2025, 115 participants have enrolled at 18 STAR sites with enrollment ongoing. CONCLUSIONS: PSC is a STAR prospective study which will address a major gap in our understanding of outcomes post-HCT specific to SCD. WeDecide, a larger study comparing HCT health-related quality of life outcomes to those who receive non-transplant disease modifying therapy (NT-DMT) is in development, and PSC will provide the HCT comparator data. These data will also be highly relevant as other curative and transformative therapies, such as gene therapy, become more widely used.

Humans

Landscape of retron diversity across the SPIRE microbial metagenome resource reveals candidate novel type XI-like lineages.

Retrons are bacterial genetic elements encoding a specialized reverse transcriptase (RT) that synthesizes multicopy single-stranded DNA and are increasingly recognized as components of bacterial anti-phage defense systems. However, their diversity and ecological distribution across large-scale genomic resources remain poorly characterized. Here, we surveyed retron RTs across the SPIRE representative metagenome collection, a non-redundant, species-level data set spanning diverse microbial habitats. Using a curated panel of type-specific hidden Markov models, we identified retrons representing all canonical types together with additional divergent lineages. Retron distribution showed strong taxonomic and ecological structuring, with some groups restricted to specific bacterial phyla, whereas others were broadly distributed across environmental categories. Systematic novelty assessment identified two candidate type XI-like lineages, TXI_C2like and TXI_noncan_h, characterized by protease-independent architectures and distinct accessory modules associated with WYL- and DnaB_C-containing proteins, respectively. De novo covariance-based analyses further identified candidate msr/msd-like non-coding RNA structures in both lineages, supporting conservation of the canonical RT-ncRNA organizational framework despite extensive sequence divergence. Together, these findings expand the known diversity of retron systems and identify type XI-like retrons as a dynamic and previously underexplored evolutionary group.IMPORTANCERetrons are bacterial genetic elements that are increasingly exploited as programmable tools for genome editing, molecular recording, and biosensing in addition to their natural role in anti-phage defense. Despite this growing biotechnological interest, the true diversity of retrons across the bacterial world has remained largely unmapped. By mining a resource of over 100,000 processed microbial metagenomes, we uncovered thousands of retron sequences spanning known types as well as previously unrecognized lineages and found that their distribution is strongly shaped by both bacterial taxonomy and ecological niche. Among these, we identified two candidate new lineages related to type XI retrons that lack the protease domain typical of this group but instead carry distinct accessory proteins, expanding the known architectural diversity of these systems. These findings broaden the catalog of retron diversity available for functional characterization and biotechnological engineering and provide a framework for prioritizing candidate lineages for future experimental validation.

effectors

Clinical Variant Interpretation with the Integrative Genomics Viewer (IGV) for Molecular Pathologists.

The integrative genomics viewer (IGV) is a pivotal tool in clinical genomics, enabling the visualization and interpretation of complex sequencing data. Bringing clinical knowledge to bear with visual evaluation of sequencing results is the primary means by which molecular pathologists and other professionals assess and finalize cases. A variety of software tools can assist, but their relationship to the underlying data must be understood and applied systematically. This study includes essential background on next-generation sequencing (NGS) data file types (e.g., FASTQ, BAM, VCF) with a discussion of their format and purpose. We then describe features of IGV that derive nuances from these files. We utilize a series of curated practical cases based on clinical vignettes through which the reader will interact with clinical NGS sequencing data using the IGV software to review various types of clinically relevant variants relative to the human reference genome. These clinical vignettes have been curated to describe examples of some of the complexities of interpretation of genomic data, and how utilizing IGV as part of a routine workflow can provide additional interpretive information for variants beyond routine bioinformatic software algorithm variant calls. The visual inspection of genomic variants utilizing the tools within IGV can unmask subtle contextual cues (i.e., variant allele frequency, strand bias, tissue-specific context) that can influence the interpretation of genomic variants. Although this study focuses on using IGV for the detection and interpretation of somatic variants, the provided applications can be extrapolated for use in the germline setting, including analysis of complex variants and detection of mosaicism.

Humans

Anatomical rationale of ablative surgery for temporal lobe seizures and dyscontrol: suggested stereo-chemode chelate-blockade alternative.

Anatomical data now strongly suggest that the common factor in curative ablative operations for the commonest (i.e. ammonshorn-sclerosis) form of temporal-lobe epilepsy is the cutting of the ipsilateral temporoammonic perforant path's "nozzle" where it leaves the entorhinal cortex to "spray" along the length of the ammonshorn. This substantially deafferences ipsilateral dentate granule-cells and hence the unsclerosed pyramidal neurons notably in "resistant sector" CA2, which are probably the source of the seizures. Stereo-chemoding of long-lasting (experimentally tested) chelates along the zone of peculiarly zinc-rich synapses of the mossy fibre system should block the commissural as well as the ipsilateral inputs to these residual neurons, to give higher percentage cures, and could probably be performed bilaterally (where indicated, in adults) without endangering memory function.

Chelating Agents

RBC-GEM: A genome-scale metabolic model for systems biology of the human red blood cell.

Advancements with cost-effective, high-throughput omics technologies have had a transformative effect on both fundamental and translational research in the medical sciences. These advancements have facilitated a departure from the traditional view of human red blood cells (RBCs) as mere carriers of hemoglobin, devoid of significant biological complexity. Over the past decade, proteomic analyses have identified a growing number of different proteins present within RBCs, enabling systems biology analysis of their physiological functions. Here, we introduce RBC-GEM, one of the most comprehensive, curated genome-scale metabolic reconstructions of a specific human cell type to-date. It was developed through meta-analysis of proteomic data from 29 studies published over the past two decades resulting in an RBC proteome composed of more than 4,600 distinct proteins. Through workflow-guided manual curation, we have compiled the metabolic reactions carried out by this proteome to form a genome-scale metabolic model (GEM) of the RBC. RBC-GEM is hosted on a version-controlled GitHub repository, ensuring adherence to the standardized protocols for metabolic reconstruction quality control and data stewardship principles. RBC-GEM represents a metabolic network is a consisting of 820 genes encoding proteins acting on 1,685 unique metabolites through 2,723 biochemical reactions: a 740% size expansion over its predecessor. We demonstrated the utility of RBC-GEM by creating context-specific proteome-constrained models derived from proteomic data of stored RBCs for 616 blood donors, and classified reactions based on their simulated abundance dependence. This reconstruction as an up-to-date curated GEM can be used for contextualization of data and for the construction of a computational whole-cell models of the human RBC.

Humans

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer

Inference of differential kinase interaction networks with KINference.

MOTIVATION: Differential kinase interaction networks (DKINs) are networks containing kinase-substrate links that are differentially active between two conditions. Existing methods are either able to predict condition-agnostic kinase-substrate links or condition-specific differential kinase activity, but do not provide differential kinase-substrate links. Moreover, existing methods for predicting kinase-substrate links usually rely on curated biochemical knowledge. Thus, there is a lack of data-driven DKIN inference methods that are also applicable when prior knowledge is scarce. RESULTS: To address this need, we present KINference. KINference combines computation of a baseline KIN representing the space of all possible kinase-substrate links with filters applied to nodes and edges to identify differentially active subnetworks that are relevant in the context of a specific phosphoproteomics dataset. For the node filters, we rely on functional relevance and differential phosphorylation scores; for the edge filters, we make use of prize-collecting Steiner trees and correlations between phosphorylation sites of kinases and their target proteins. Tests on two phosphoproteomics datasets (kinase inhibition in breast cancer cells, SARS-CoV-2 infection in Calu-3 cells) show that the proposed filters produce significant results in terms of overlap with known interactions between kinases and phosphorylation sites. Furthermore, a case study on the SARS-CoV-2 infection data, suggests a potential host pathway linked to virus replication, showcasing the process of hypothesis generation utilizing DKINs computed by KINference. AVAILABILITY AND IMPLEMENTATION: KINference is available as an R package at https://github.com/bionetslab/KINference and https://doi.org/10.5281/zenodo.15411150. Scripts to reproduce the results are available at https://github.com/bionetslab/KINference-Evaluation-Scripts and https://doi.org/10.5281/zenodo.15424599.

Humans