PubMed HealthSearch

SEARCH · PubMed Health

Results for “data curation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics

Antibiotic-impregnated bone graft to prevent infection after total hip arthroplasty (ABOGRAFT): protocol for a randomised, double-blind, placebo-controlled trial.

INTRODUCTION: Studies have shown promising results using bone graft as a carrier for local administration of antibiotics to reduce the risk of prosthetic joint infection (PJI). The objective of this clinical trial is to determine if tobramycin and vancomycin-impregnated bone graft is safe and effective in reducing the rate of PJI after total hip arthroplasty (THA). METHODS AND ANALYSIS: This study is an international, randomised, double-blinded, placebo-controlled clinical drug trial. Patients scheduled for THA (n=1100) requiring bone grafting (excluding revisions due to an ongoing infection) are randomised in a 1:1 ratio to prophylactic treatment with tobramycin and vancomycin or placebo-impregnated bone graft.The primary outcome is the time to reoperation due to infection or diagnosis of PJI, expressed as a relative risk difference between the two groups. A risk reduction of at least 50% is considered clinically relevant. Secondary outcomes are time to and reason for reoperation and implant revision, type of micro-organism and antibiotic susceptibility pattern within 2 and 5 years after surgery. Safety outcomes are the number of adverse events and revision rate due to aseptic loosening. The primary analysis will be performed using proportional hazard models. ETHICS AND DISSEMINATION: The study has been approved under the Clinical Trial Regulation No 536/2014 (EU CT; 2024-510921-25-00). Results will be published in open-access peer-reviewed journals and disseminated to patient organisations and the media, and de-identified individual participant data will be curated and shared on reasonable request in accordance with the Findability, Accessibility, Interoperability and Reuse principles, subject to the laws and regulations governing data protection in each participating country. TRIAL REGISTRATION NUMBER: NCT05169229.

Humans

Whole genome sequencing reveals a specific microbiota in subglottic stenosis C. acnes may contribute to inflammation.

PURPOSE: Subglottic stenosis (SGS) progressively reduces the airway below the vocal folds. The cause is not known and there is a recurrent need of surgical treatment. Including all phenotypes, SGS affects 1/400 000/yr, with a female dominance. Previous studies have revealed a possible role of the Mycobacterium complex in SGS development. Our hypothesis is that microbiota is associated with the inflammation in SGS, if true it might affect the prevailing treatment options. METHODS: This prospective cross-sectional study included biopsies from 34 patients with subglottic stenosis, collected between 2020 and 2023. Nucleic acids were extracted from the tissue samples and analysed using whole genome sequencing. Microbial composition was characterized using taxonomic profiling of sequencing data. Species with sufficient read counts were selected for further validation using sequence alignment methods to ensure accuracy of identification. RESULTS: Using the most comprehensive form of genomic testing currently in clinical use, we present curated and stable data on the presence of Cutibacterium acnes in 28 out of the 34 cases. CONCLUSION: Cutibacterium acnes may serve as a driver of the inflammation characterizing SGS and should be considered in therapeutically oriented future studies.

Cutibacterium acnes

Meta2DB: curated shotgun metagenomic feature sets and metadata for health state prediction.

SUMMARY: Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health. AVAILABILITY: Data is available at https://gdo-meta2db.llnl.gov/ and https://zenodo.org/records/17315984.

Metadata

Predictive design of tissue-specific mammalian enhancers that function in the mouse embryo.

Enhancers control tissue-specific gene expression across animals1. Although deep learning2,3 has enabled enhancer prediction and design in mammalian cell lines and non-mammalian model organisms4-10 (reviewed in a previous publication11), it remains unclear whether such approaches can operate within the regulatory complexity of mammalian genomes and tissues in vivo. Here we present a general strategy for designing tissue-specific enhancers that function reliably in mice. We use deep learning to train compact convolutional neural networks on curated chromatin accessibility data and fine-tune them by transfer learning on validated human and mouse enhancers. Guided by these models, we design 15 synthetic enhancers for the heart, limb and central nervous system in mouse embryos, all of which are active in their intended target tissue. These results demonstrate that mammalian enhancer function can be reliably inferred from DNA sequence alone, enabling the predictive de novo design of tissue-specific synthetic enhancers from modest training sets. This work establishes a generalizable framework for programmable control of mammalian gene expression in vivo, opening new avenues in functional genomics, synthetic biology and gene therapy.

Animals

Benchmarking Assembly-Free K-mer Methods for Species Identification in Complex Plant Groups: A Case Study in Populus.

Species identification in taxonomically complex plant groups is frequently limited by the inadequacy of organellar markers, whose phylogenetic signal is disrupted by cytonuclear discordance and chloroplast capture. Using the taxonomically complex genus Populus as a model, we evaluated an assembly-free k-mer workflow against a curated SNP reference benchmark. Whole-genome resequencing data from 235 Populus individuals were curated to a 202-individual, 34-species reference dataset in which all retained species are strictly monophyletic in a genome-wide SNP analysis. Independent maximum likelihood analyses further confirmed that the 31 non-hybrid backbone species each maintained high-support monophyly, while taxa of documented reticulate origin showed placement patterns consistent with their reticulate histories. ABBA-BABA D-statistics detected widespread residual allele sharing within the backbone, though the strongest signals did not correspond to the species pairs responsible for the few k-mer identification failures. Against this benchmark, complete plastomes showed limited resolution, recovering only 3.0% species monophyly and 71.1% nearest-neighbor assignment. The optimized k-mer workflow, operating directly on raw reads without assembly or alignment, recovered 91.2% species monophyly, 99.0% nearest-neighbor assignment, and 98.0% group-average assignment. K-mer length was the primary accuracy-controlling parameter, with k = 31 falling within a stable accuracy plateau. Distance-based metrics reached near-saturation at 0.2× sequencing depth, indicating that low-coverage genome skimming can support scalable nuclear genome-based identification with standard computational resources. K-mer distance heatmaps also flagged unusual genomic affinities in hybrid-origin and outlier samples, providing a rapid screen for subsequent population genomic analyses. These results support assembly-free k-mer distances as an efficient tool for reference-based species identification and sample screening in complex plant groups, with residual limitations concentrated near recently diverged species boundaries. Model-based phylogenomic, coalescent, and network analyses remain necessary for resolving deeper species relationships and detailed introgression histories.

Populus

The Lipid Interactome: an interactive and open access platform for exploring cellular lipid-protein interactions.

SUMMARY: Lipid-protein interactions play essential roles in cellular signaling and membrane dynamics, yet their systematic characterization has long been hindered by the inherent biochemical properties of lipids. Recent advances in functionalized lipid probes-equipped with photoactivatable crosslinkers, affinity handles, and photocleavable protecting groups-have enabled proteomics-based identification of lipid interacting proteins with unprecedented specificity and resolution. Despite the growing number of published lipid interactomes, there remains no centralized effort to harmonize, compare, or integrate these datasets. The Lipid Interactome addresses this gap by providing a structured, interactive web portal that adheres to FAIR data principles-ensuring that lipid interactome studies are Findable, Accessible, Interoperable, and Reusable. Through standardized data formatting, interactive visualizations, and direct cross-study comparisons, this resource enables researchers to systematically explore the protein-binding partners of diverse bioactive lipids. By consolidating and curating lipid interactome proteomics data from multiple studies, the Lipid Interactome database serves as a critical tool for deciphering the biological functions of lipids in cellularsystems. AVAILABILITY AND IMPLEMENTATION: This site can be viewed at LipidInteractome.org. All data are available for download. No user information is collected or necessary for data navigation, interaction, or download.

Proteins

Selective toxicity of anticancer drugs: Presidential Address.

In the chemotherapy of infectious diseases, selective toxicity has been achieved by designing curative regimens based on pharmacokinetic data. Selective toxicity of antitumor drugs has been demonstrated for rapidly growing large growth fraction tumors occurring in patients under age 30. In these tumors curative schedules have been achieved by application of animal data relating to cellular and drug kinetics. The attempts to improve chemotherapy of large and small growth fraction tumors by kinetic observations in vivo in humans have been disappointing. Recent evidence suggests that the heterogeneity of cells within tumors has prevented precise observations on the relation of cellular and drug kinetics to improved selective toxicity. The availability of xenografts, flow cytometry, and tumor markers presents an opportunity to isolate subpopulations of tumor cells; to characterize their cellular and drug kinetics; and to correlate these with values obtained in vivo in humans. It should then be possible at long last to examine the potential role of cellular and drug kinetics in devising drug schedules with greater selective toxicity for human cancer.

Antineoplastic Agents

Exploring shotgun metagenomic data to detect microeukaryotic pathogens in wildlife.

BACKGROUND: Microeukaryotic parasites of the intestinal tract are an understudied group of organisms that infect humans and many other animals. Targeted sequencing methods focused on individual loci are usually employed for detection of these parasites, making comprehensive studies of microeukaryotic parasite diversity within hosts or other systems difficult. Exploratory approaches such as shotgun metagenomic sequencing to survey the diversity of microeukaryotic parasites in new and existing datasets are not well developed. RESULTS: Utilizing existing datasets from 12 goose fecal samples, we explored some of the benefits and challenges of using shotgun metagenome sequencing to detect microeukaryotic parasites. We demonstrated the importance of careful curation of read classification data to avoid erroneously linking pathogens to hosts or environments as unsupported classifications were common in the data and varied widely depending on analysis parameters. However, we were able to establish strong support for the presence of sequences of Eimeria and Enterocytozoon bieneusi. In addition, examination of trichomonad reads indicated that parasite reads mapping to human pathogens unlikely to colonize geese may in fact represent cryptic microeukaryotic species that are not included in existing curated databases opening new potential avenues of study. CONCLUSIONS: Taken together these findings support the idea that exploring microeukaryotic parasite diversity within shotgun metagenomic datasets can be beneficial to our understanding of the presence and diversity of these organisms in wildlife hosts.

Animals

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

Integrating Next-Generation Sequencing into von Willebrand Disease Diagnostics: Insights from the PCM-EVW-ES Multicenter Project.

Von Willebrand disease (VWD) is the most common inherited bleeding disorder, caused by quantitative or qualitative defects in von Willebrand factor (VWF). Diagnosis is challenging and requires integrating bleeding history, VWF antigen and activity measurements, FVIII assays, and specialized phenotyping. Genetic testing is increasingly recognized as a key component. Here, we review current concepts in VWD diagnostics and highlight the Spanish Clinical and Molecular Profile of von Willebrand Disease (PCM-EVW-ES) project as a model for genomics-enabled precision medicine. PCM-EVW-ES is a multicenter initiative involving 48 hospitals, centralized phenotypic testing, and next-generation sequencing of the VWF coding region, enabling definitive classification in 730 individuals with VWD to date. Harmonized recruitment criteria and standardized workflows improve subtype assignment, uncover complex genotypes, refine genotype-phenotype correlations, and facilitate the identification of asymptomatic carriers. The PCM-EVW-ES variant spectrum highlights recurrent disease-causing variants in Spain and underscores the value of coordinated national registries for variant curation. Building on these data, we propose a diagnostic algorithm in which bleeding assessment and first-line VWF/FVIII assays, combined with, early VWF molecular testing increases diagnostic accuracy and guides targeted second-line investigations to confirm and refine VWD subtype classification. We also outline persisting challenges, including the interpretation of variants of uncertain significance and patients without identifiable pathogenic VWF variants, and future directions integrating third-generation sequencing, expanded gene panels, functional studies, and artificial-intelligence-driven multiomic approaches. Together, these advances illustrate how robust multicenter studies can bridge the gap between complex diagnostics and clinical practice in VWD.

Humans

CircExor enables interpretable prediction of circRNA localization into extracellular vesicles.

Certain circular RNAs (circRNAs) are selectively enriched in extracellular vesicles (EVs), in which they contribute to intercellular communication and represent promising biomarkers, yet the sequence determinants of their sorting remain unclear. Existing computational predictors are optimized mainly for linear RNAs and rarely address circRNA localization into EVs. Here we introduce circExor, the first framework specifically designed for circRNA EV localization. We curate a dedicated benchmark data set of 2102 circRNAs and implement a variable-length end-to-end concatenation strategy together with k-mer frequency encoding to accommodate circular topology, long sequence length, and length heterogeneity. Using a tree-based classifier, circExor achieves superior performance compared with RNAlocate-v3 and ExoGRU, reaching an AUROC of 0.743 on the internal test set and an average AUROC of 0.680 on the held-out test set. SHAP-based analysis, sequence perturbation analysis, motif mapping, and cell-based experimental validation support the predicted EV tendency and identify YBX1, HNRNPK, HNRNPL, and NOVA2 as candidate RBPs potentially associated with circRNA sorting. CircExor therefore provides a predictive and interpretable framework that links in silico modeling to mechanistic hypotheses, and supports biomarker discovery and candidate prioritization for downstream studies of EV-associated circRNAs.

Journal Article

Multi-omics Mendelian Randomization Prioritizes Neutrophil Extracellular Trap-related Genes Associated with Atrial Fibrillation Risk.

BACKGROUND: Neutrophil extracellular traps (NETs) participate in thrombosis, inflammation, and cardiovascular remodeling, yet whether NET-related genes (NRGs) are associated with atrial fibrillation (AF) risk across multiple molecular layers remains unclear. This study used a multiomics Mendelian randomization framework to prioritize NRGs supported by methylation, expression, and protein quantitative trait loci (QTL) data. METHODS: Genome-wide significant cis instruments (P < 5 &#xd7; 10-8) were obtained for 90 methylation QTLs (mQTLs), 100 expression QTLs (eQTLs), and 38 protein QTLs (pQTLs) mapped to 137 literature- curated NRG entries. Summary-data-based Mendelian randomization (SMR) coupled with the heterogeneity in dependent instruments (HEIDI) test was applied using whole-blood mQTL data (n = 1,980), eQTLGen blood eQTL data (n = 31,684), and deCODE plasma pQTL data (n = 35,559). AF outcome data were obtained from a meta-analysis including 60,620 cases and 970,216 controls of European ancestry. RESULTS: At the methylation level, 21 CpG-feature associations across 13 genes remained significant after HEIDI filtering and false discovery rate (FDR) correction. Expression-level analysis identified eight significant gene-AF associations, whereas protein-level analysis identified seven significant features representing five unique proteins. Cross-omics integration prioritized C3, MAPK3, and STAT3 as Tier 1 genes, CTSC, LPAR3, and THBD as Tier 2 genes, and fourteen additional genes as Tier 3 candidates. C3 showed risk-increasing protein-level associations together with multiple significant CpG signals, whereas MAPK3 and STAT3 showed directionally protective expression/protein or methylation/protein patterns. DISCUSSION: The cross-omics convergence on C3, MAPK3, and STAT3 is consistent with complement activation, immune-fibrotic signaling, and cytokine-regulatory pathways implicated in AF biology, but the findings should be interpreted as genetic prioritization rather than definitive intervention-ready causality. CpG-level heterogeneity at the C3 locus and the blood/plasma origin of the QTL resources further support a cautious interpretation. Modest colocalization support and the unresolved possibility of pQTL sample overlap further support this cautious, hypothesis-generating interpretation. CONCLUSION: Multi-omics SMR prioritizes C3, MAPK3, and STAT3 as the most consistently supported NET-related genes associated with AF risk. These findings provide a framework for atrialtissue replication and mechanistic validation of NET-related pathways in AF.

Atrial fibrillation

CERTOMICS: trusted single-cell multiomics pipeline for high-resolution profiling of adoptive cellular immunotherapies.

SUMMARY: Adoptive cellular immunontherapies, such as chimeric antigen receptor (CAR) T cell therapy, have transformed cancer treatment, yet challenges such as resistance, relapse, and high costs limit their efficacy and accessibility. A comprehensive understanding of cellular heterogeneity and molecular profiles is essential to improve these therapies. Advanced single-cell multiomics technologies have the power to analyze the complex interactions between CAR-engineered cells, immune cells, and tumor cells. However, standardized single-cell multiomics computational pipelines specifically tailored to CAR-engineered cell products are lacking. Due to the synthetic nature of CAR transgenes, additional steps for reliable identification and characterization of CAR-positive cells are required but not included in existing data-processing workflows. To address this, we present CERTOMICS, a Nextflow-based, CAR-aware pipeline offering enhanced CERTainty in immunophenotyping and data interpretation, tailored for single-cell multiOMICSprofiling of adoptive cellular immunotherapies. The pipeline standardizes processing 10x Genomics single-cell multiomics data and integrates CAR-specific identification and quality control. Additionally, a curated repository of CAR construct sequences and annotation data is provided, serving as an extensible resource to support the analysis and development of CAR T cell therapies. AVAILABILITY AND IMPLEMENTATION: Detailed documentation of this pipeline, along with a resource on latest FDA-approved CAR therapies is available on our website: https://fraunhofer-izi.github.io/Living-Drugs-Wiki/. The data underlying this article are available on GitHub at https://github.com/fraunhofer-izi/CERTOMICS. The code is also published on Zenodo at https://doi.org/10.5281/zenodo.18709693.

Multiomics

TaxTriage: an open-source metagenomic sequencing data analysis pipeline enabling putative pathogen detection.

MOTIVATION: TaxTriage is a comprehensive pathogen identification workflow designed for both short- and long-read untargeted DNA and RNA sequencing data. Combining read classification, mapping, and de novo assembly approaches, putative pathogens are identified through comparisons to curated pathogens and abundance expectations from healthy cohort data. Flexible installation options are enabled using Nextflow&#x2122; (NF), including cloud deployment via NF Tower (Seqera Platform) and local installation on a variety of systems, including standalone installations without external internet access. Final analysis summaries are compiled into an Organism Discovery Report, which lists likely pathogens and supporting data, including a custom confidence score. RESULTS: Evaluation of published in silico, clinical, and outbreak datasets identified performance comparable to alternative cloud-based processing pipelines for expected pathogen and co-infection detection with similar sensitivity and increased specificity. To support both public health and veterinary diagnostics communities, customization options have been incorporated to enable improved performance for host species of interest. AVAILABILITY AND IMPLEMENTATION: Source code for TaxTriage is freely available at https://github.com/jhuapl-bio/taxtriage. TaxTriage v2.1.1 has been archived on Zenodo at https://zenodo.org/records/17081354 to permit reproducible analysis as described in this manuscript.

Software

engGNN: a dual-graph neural network for omics-based disease classification and feature selection.

Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph neural networks offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated graph, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive representations, thereby improving predictive performance and interpretability. Through extensive simulation studies and real-world applications to three independent gene expression datasets, engGNN consistently demonstrates strong classification performance compared with competitive baselines. Beyond classification, engGNN provides feature- and source-level interpretability, enabling biologically meaningful analyses such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.

Graph Neural Networks

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural

Circulating immune complexes in patients following clinically curative resection of colorectal cancer.

Sixty-nine patients have been followed prospectively after curative resection of Dukes-Kirklin B-2 or C colorectal cancer. Serial plasma samples were studied in selected patients to determine changes in circulating immune complex concentrations (CIC) following primary tumor resection, and to compare serial plasma CIC and carcinoembryonic antigen (CEA) levels. CIC was determined in an average of seven serial samples per patient by inhibition of antibody-dependent cell-mediated cytotoxicity (ADCC). CEA assays were performed by the Hanson Z-gel method. Two distinct patterns of serial CIC have emerged. In seven patients with no known tumor recurrences, serial CEA levels and CIC oscillated regularly and were inversely related. In seven of eight patients whose tumors recurred, both CEA and CIC rose together. In three patients with elevated plasma CEA levels due to inflammatory bowel disease, serial Ag-Ab complex concentrations did not vary, nor did separated Ag or Ab fractions inhibit ADCC. These data suggest that, in patients following curative resection of colorectal cancer, serial changes in circulating immune complexes may discriminate between transient CEA elevations which occur despite no known tumor recurrence and tumor recurrence which is beyond the capacity of adequate host antitumor defense.

Adenocarcinoma