PubMed HealthSearch

SEARCH · PubMed Health

Results for “data curation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Microbiology Galaxy Lab: The first community-driven gateway for reproducible and FAIR analysis of microbial data.

The explosion of microbial omics data has outpaced the ability of many researchers to analyze it, with complex tools and limited computational resources creating barriers to discovery. To address this gap, we present the Microbiology Galaxy Lab: a free, globally accessible, community-supported platform that combines state-of-the-art analytical power with user-friendly accessibility. Supported by the Galaxy and global microbiology communities, this platform integrates over 315 tool suites and 115 curated workflows, enabling comprehensive metabarcoding, (meta)genomic, (meta)transcriptomic, and (meta)proteomic data analysis within a FAIR-aligned environment. It also supports research in the health and infectious disease sectors, as well as in environmental microbiology. The platform's utility is exemplified through various use cases, including antimicrobial resistance tracking, biomarker prediction, microbiome classification, and functional annotation of key microbes. Built on reproducibility and community engagement, it supports creation, sharing, and updating of best-practice workflows. Over 35 tutorials and learning paths empower scientists, fostering an ecosystem that keeps resources at the forefront of microbial science. The Microbiology Galaxy Lab enables collective analysis, democratising research, thereby accelerating discovery across the global microbiology community (microbiology.usegalaxy.org, .eu, .org.au, .fr).

Journal Article

Chemotherapy as an adjuvant to surgery for colorectal cancer. A follow-up report.

An adjuvant program of fluorouracil for patients undergoing "curative" resection for adenocarcinoma of the colon and rectum was initiated as a randomized clinical trial in January 1968. Patients were randomly assigned to an intraluminal fluorouracil or intraluminal control (saline) group and were so treated at the time of surgical resection if findings at operation indicated that all gross neoplastic disease could be resected. Those patients receiving intraluminal fluorouracil (30 mg/kg) received intravenous fluorouracil (10 mg/kg) on each of the first two postoperative days and five subsequent postoperative courses of oral fluorouracil (90 mg/kg) in each 18-day course over a one-year period. By July 1, 1975, there were 203 patients undergoing curative resection entered into the study. Survival and disease-free data, as of Dec 31, 1976, revealed no benefit from this adjuvant course of fluorouracil. These data support the need for continued randomized clinical trials of new and innovative adjuvant therapy compared with an untreated control group.

Adenocarcinoma

RP3Net: a deep learning model for predicting recombinant protein production in Escherichia coli.

MOTIVATION: Recombinant protein expression can be a limiting step in the production of protein reagents for drug discovery and other biotechnology applications. We introduce RP3Net (Recombinant Protein Production Prediction Network), an AI model of small-scale heterologous soluble protein expression in Escherichia coli. RP3Net utilizes the most recent protein and genomic foundational models. A curated dataset of internal experimental results from AstraZeneca and publicly available data from the Structural Genomics Consortium was used for training, validation and testing of RP3Net. RESULTS: RP3Net achieves an increase in area under the receiver operator curve (AUROC) of 0.15, compared to a baseline model. When experimentally validated on an independent, prospective, manually selected set of 97 constructs, RP3Net outperformed currently available models, with an AUROC of 0.83, delivering accurate predictions in 77% of the cases, and correctly identifying successfully expressing constructs in 92% of cases. AVAILABILITY AND IMPLEMENTATION: The model, along with installation and running instructions, is available under an MIT licence at https://github.com/RP3Net/RP3Net, DOI 10.5281/zenodo.17243498.

Escherichia coli

Markers of microvascular instability predict severity and survival in idiopathic pulmonary fibrosis.

INTRODUCTION: Most research on idiopathic pulmonary fibrosis (IPF) has focused on the interplay among fibroblasts, the immune system and epithelial cells. There is growing evidence that microvascular dysfunction also plays a role in disease progression, but large human translational studies are lacking. In this research, we aim to identify a proteomic signature of microvascular instability and assess the impact of current therapeutics on the microvasculature. METHODS: Olink proteomic data from patients with IPF were obtained from the Pulmonary Fibrosis Foundation Patient Registry (PFF-PR) (n=914) and an independent validation cohort (n=366). Among the PFF-PR, 640 patients also have whole-blood RNA sequencing data available. A subset of 79 microvascular-associated proteins was curated, and their associations with disease severity and transplant-free survival were examined. An adaptive least absolute shrinkage and selection operator was used to generate a novel microvascular risk score. RESULTS: Higher plasma levels of five microvascular-associated proteins (SDC1, MMP10, THBS2, HGF and SERPINA5) were associated with lung function and survival in both cohorts. Whole-blood RNA sequencing of patients with microvascular risk revealed enrichment of immune-mediated processes. Patients with higher microvascular risk who were subsequently put on nintedanib in the following year had significantly better 3-year transplant-free survival compared with patients who did not receive antifibrotic intervention (HR 0.56, 95% CI 0.35 to 0.89, p=0.0142). DISCUSSION: Integrative multi-omics analyses suggest that perturbations to microvascular remodelling contribute to disease severity and progression in IPF. This analysis offers a framework for a precision medicine approach for IPF.

Idiopathic pulmonary fibrosis

PSIA: A Comprehensive Knowledgebase of Plant Self-incompatibility.

Self-incompatibility (SI) is an important genetic mechanism in angiosperms that prevents inbreeding and promotes outcrossing, with significant implications for crop breeding, including genetic diversity, hybrid seed production, and yield optimization. In eudicots, SI is typically governed by a single S-locus containing tightly linked pistil and pollen S-determinant genes. Despite major advances in SI research, a centralized, comprehensive resource for SI-related genomic data remains lacking. To address this gap, we developed the Plant Self-Incompatibility Atlas (PSIA), a systematically curated knowledgebase providing an extensive compilation of plant SI, including genomic resources for SI species, S gene annotations, molecular mechanisms, phylogenetic relationships, and comparative genomic analyses. The current release of PSIA includes over 500 genome assemblies from 469 SI species. Using known S genes as queries, we manually identified and rigorously curated 3700 S genes. PSIA provides detailed S-locus information from assembled genomes of SI species and offers an interactive platform for browsing, BLAST searches, S gene analysis, and data retrieval. Additionally, PSIA serves as a unique platform for comparative genomic studies of S-loci, facilitating exploration of the dynamic processes underlying the origin, loss, and regain of SI. As a comprehensive and user-friendly resource, PSIA will greatly advance our understanding of angiosperm SI and serve as a valuable tool for crop breeding and hybrid seed production. PSIA is freely available at http://www.plantsi.cn.

Self-Incompatibility in Flowering Plants

Proteomic insights into Helicobacter pylori infection in stomach cells, revealing host response and host-targeted therapeutics repurposing.

BACKGROUND: Helicobacter pylori (H. pylori) is a globally prevalent gastric pathogen strongly associated with chronic gastritis, peptic ulcers, and gastric cancer. While bacterial factors have been extensively studied, host proteomic responses and their therapeutic potential remain largely underexplored. RESEARCH DESIGN AND METHODS: Current analyses employed a systematic proteomics-based data integration and harmonization approach (retrospective qualitative cohort study) to identify important differentially regulated host proteins. Proteomic datasets were curated from in vitro studies and analyzed for functional enrichment, protein-protein interaction networks, and hub protein identification. To explore therapeutic repurposing, drug repositioning was performed using the DrugBank database. RESULTS: Data summation describing protein differential regulation in human gastric cells as a result of the infection revealed 1672 perturbed host proteins. Bioinformatics analysis revealed 11 proteins including CSK, MET, RELA, MARK2, GRB2, FTO, PLCG1, CRKL, RPS5, RPS9, and RPS27A to be ideal host targets for therapeutic repurposing. Clinically approved drugs such as Dasatinib (targeting CSK) and Crizotinib (targeting MET) emerged as promising candidates due to favorable pharmacokinetics and known bioactivity. CONCLUSIONS: Host-directed therapeutics could offer alternative strategies to conventional antibiotic therapy, addressing challenges such as resistance and infection recurrence, providing a foundation for future experimental validation and development of host-targeted interventions for infection control.

Humans

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals

Apollo: a sequence annotation editor.

The well-established inaccuracy of purely computational methods for annotating genome sequences necessitates an interactive tool to allow biological experts to refine these approximations by viewing and independently evaluating the data supporting each annotation. Apollo was developed to meet this need, enabling curators to inspect genome annotations closely and edit them. FlyBase biologists successfully used Apollo to annotate the Drosophila melanogaster genome and it is increasingly being used as a starting point for the development of customized annotation editing tools for other genome projects.

Animals

[Analysis of the quality of clinical diagnosis from generalized findings of the pathologoanatomic service].

A statistical analysis of generalized data of the pathoanatomical service on quality of clinical diagnosis in curative-prophylactic institutions in 54 administrative territories of the RSFSR was carried out. The structure (extensive indices) and frequency (intensive indices)of erroneous clinical diagnoses referring to the most important classes of diseases were identified. As to the structure of indices and frequency of clinico-anatomic disparities the first place was occupied by oncological diseases (20.1+/-0.11 and 14.2+/-0.22%), the second--by infectious diseases (16.5+/-0.1 and 13.0+/-0.34%), the third--by diseases of the digestive system (14.6+/-0.09 and 13.0+/-0.33%), the forth--by diseases of the urogenital system (14.0+/-0.09 and 12.2+/-0.49%), the fifth--by disease of the respiratory system (12.7+/-0.09 and 10.6+/-0.24%), the sixth--by diseases of the cardiovascular system (11.1+/-0.08 and 8.0+/-0.14%). The recommendation is put forward to carry on annually a complex satistical analysis of extensive and intensive indices of erroneous clinical diagnoses demonstrating the quality of clinical diagnosis in therapeutic institutions of a given administrative territory.

Diagnostic Errors

CancerTrialMatch: a computational resource for the management of biomarker-based clinical trials at a community cancer center.

MOTIVATION: The widespread implementation of next-generation sequencing in cancer care has enabled routine use of molecular and biomarker profiling. At our cancer center, as with many others, biomarker-based clinical trials are increasingly available to oncologists as potential treatment options via molecular tumor boards. To better support this effort, we developed CancerTrialMatch, a systematic approach to capture structured clinical trial data and match patients to trials based on their disease characteristics and sequencing profiles. RESULTS: CancerTrialMatch is an open-source application designed to streamline clinical trial curation and patient trial matching, while also enabling an institution's curated trial portfolio to be distributed across the institution for easy access to providers, care teams and researchers. It facilitates curating, updating, and searching for trials through a semi-automated interface built using R Shiny, MongoDB, and Docker. While much of the trial data is retrieved via the clinicaltrials.gov Application Programming Interface, certain items like biomarkers and disease subtypes are entered manually. The user inputs disease type using the OncoTree classification, and provides relevant biomarker details, such as mutations, copy numbers, fusions, and other disease-specific markers. This resource reduces the time required for institutional trial management and helps to identify potential clinical trials for patients, ultimately supporting larger clinical trial enrollment and enhancing the clinical application of precision oncology. AVAILABILITY AND IMPLEMENTATION: CancerTrialMatch was implemented and tested on Windows 11 (64-bit, 32 GB RAM) using WSL2 with Ubuntu 22.04. Docker 27.0.3 and Docker Compose 2.28.1 were used to build images and containers. Users can build it by cloning the repo and following the README instructions and supplemental file (cancertrialmatchsupplemental.pdf) . The source code and example data are available in GitHub and Figshare at https://github.com/AveraSD/CancerTrialMatch and 10.6084/m9.figshare.28447367 respectively.

Humans

Results of pulmonary arterial banding in infancy. Survey of 5 years' experience in the New England Regional Infant Cardiac Program.

The results of pulmonary arterial banding in 238 infants, 12 percent of the infants admitted to the New England Regional Infant Cardiac Program, is reviewed. Overall survival to age 1 year was 63 percent. Survival was least likely (37 percent) in those who required banding within the 1st month of life. Additional surgery decreased the survival rate in those operated on after 1 month of age. Infants with anomalies for which no corrective surgical procedure is available (23 of 238) have only a 30 percent chance of survival. Those with lesions correctable within the 1st year (133 of 238) have a 74 percent survival rate; 52 percent (82 of 238) of those for whom a curative operation is available after the 1st year survive. These pulmonary arterial banding data coupled with results of primary correction should provide the data base required for an intelligent decision in respect to appropriate surgical treatment of infants with critical heart disease.

Heart Defects, Congenital

Functional mapping of the Trypanosoma cruzi serinome by fluorophosphonate activity-based protein profiling.

Serine hydrolases (SHs) constitute one of the largest enzyme superfamilies in eukaryotes, yet their roles in Trypanosoma cruzi, the causative agent of Chagas disease, remain largely uncharacterized. Here, we report an activity-based chemoproteomic map of the T. cruzi epimastigote serinome by combining genome-informed in silico curation with whole-cell activity-based protein profiling (ABPP) using a panel of cell-permeable fluorophosphonate (FP)-alkyne probes. Whole-cell labelling followed by label-free quantitative proteomics (LFQ-MS) identified 37 enriched SH-like proteins, including 35 with conserved or partially conserved catalytic triad/dyad features, spanning lipases, peptidases, esterases, and previously uncharacterized hydrolases. The 35 SHs represent approximately 63% of the 56 predicted SHs retained after catalytic-site curation. Domain architecture analysis revealed broad structural diversity, while orthologue-based localization data suggested association with multiple subcellular compartments, including glycosomal, mitochondrial, and endosomal localizations. Gene Ontology enrichment highlighted lipid metabolic and catabolic processes as dominant functional themes, and protein-protein interaction network analysis supported functional connectivity among the captured enzymes. Several identified SHs, including oligopeptidase B, prolyl oligopeptidase Tc80, serine carboxypeptidase CPB1, and phospholipase A1 (PLA1) have previously been characterized in trypanosomatids, with roles linked to parasite virulence or host-pathogen interactions. Together, these findings establish a fluorophosphonate-based chemoproteomic resource for the kinetoplastid community and prioritize probe-accessible active T. cruzi SHs for future functional validation and antiparasitic inhibitor discovery.

Activity-based protein profiling

[Treatment of rhabdomyosarcoma in mice C3H/He by Co 60 and hyperbaric oxygen (author's transl)].

UNLABELLED: Experimental study of rhabdomyosarcoma with successive transplantation upon C3H/He mice, treated by irradiation (Co 60) and combined irradiation-hyperbaric oxygen (HBO), dating from 3, 14 and 15 days after transplantation. The data (tumor volume evolution, histological modifications, pulmonary metastases) are compared with controls. CONCLUSIONS: curative radiotherapy depends on starting treatment as soon as possible with or without HBO. After the 14th day, sensitisation to combined HBO and C60 is seen. The extension of pulmonary metastases is a function of tumor growth. Paradoxically metastases were less frequent after HBO only and more frequent after HBO-Co 60.

Animals

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts ∼50 000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein

Assessment of Gene Set Enrichment Analysis using curated RNA-seq-based benchmarks.

Pathway enrichment analysis is a ubiquitous computational biology method to interpret a list of genes (typically derived from the association of large-scale omics data with phenotypes of interest) in terms of higher-level, predefined gene sets that share biological function, chromosomal location, or other common features. Among many tools developed so far, Gene Set Enrichment Analysis (GSEA) stands out as one of the pioneering and most widely used methods. Although originally developed for microarray data, GSEA is nowadays extensively utilized for RNA-seq data analysis. Here, we quantitatively assessed the performance of a variety of GSEA modalities and provide guidance in the practical use of GSEA in RNA-seq experiments. We leveraged harmonized RNA-seq datasets available from The Cancer Genome Atlas (TCGA) in combination with large, curated pathway collections from the Molecular Signatures Database to obtain cancer-type-specific target pathway lists across multiple cancer types. We carried out a detailed analysis of GSEA performance using both gene-set and phenotype permutations combined with four different choices for the Kolmogorov-Smirnov enrichment statistic. Based on our benchmarks, we conclude that the classic/unweighted gene-set permutation approach offered comparable or better sensitivity-vs-specificity tradeoffs across cancer types compared with other, more complex and computationally intensive permutation methods. Finally, we analyzed other large cohorts for thyroid cancer and hepatocellular carcinoma. We utilized a new consensus metric, the Enrichment Evidence Score (EES), which showed a remarkable agreement between pathways identified in TCGA and those from other sources, despite differences in cancer etiology. This finding suggests an EES-based strategy to identify a core set of pathways that may be complemented by an expanded set of pathways for downstream exploratory analysis. This work fills the existing gap in current guidelines and benchmarks for the use of GSEA with RNA-seq data and provides a framework to enable detailed benchmarking of other RNA-seq-based pathway analysis tools.

Humans

Rapid assessment of clinical severity for salmonellosis cases via protein family domain analysis and machine learning.

Salmonella is a common pathogen, infecting more than a million people yearly. Rapid assessment of clinical case severity is essential for improving patient outcomes and optimizing healthcare resources. Advancements in genome sequencing technologies have enabled the analysis of bacterial genomes from many clinical cases, opening up new opportunities for precise and timely diagnosis. This study proposes a genome-based framework for identifying critical Salmonella cases before the onset of critical symptoms and facilitating early medical intervention. By leveraging protein family (Pfam) domains as the representation for genomic data, the complex genetic profiles of Salmonella cases are simplified into interpretable features. The severity levels of cases were investigated through rigorous data analysis, resulting in a set of 70 Pfam domains that could be potentially used as biomarkers. Machine Learning was employed to assess the predictive power of the curated Pfam biomarkers, achieving high accuracy (~93%) in sorting cases into critical, moderate, and mild categories. The results demonstrate the efficacy of the proposed approach. This framework highlights the potential of using bacterial genomic data in clinical decision-making, opening the window for timely personalized interventions for Salmonella infection management.

Domains of unknown function (DUFs)

ELISA (Embedding-Linked Interactive Single-cell Agent): an interpretable hybrid generative Artificial Intelligence agent for expression-grounded discovery in single-cell genomics.

Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language. Here, we introduce ELISA (Embedding-Linked Interactive Single-cell Agent), an interpretable framework that unifies single-cell generative pretrained transformer expression embeddings with biomedical bidirectional encoder representations from transformers-based semantic retrieval and large-language model (LLM)-mediated interpretation for interactive single-cell discovery. An automatic query classifier routes inputs to gene marker scoring, semantic matching, or reciprocal rank fusion pipelines depending on whether the query is a gene signature, natural language concept, or mixture of both. Integrated analytical modules perform pathway activity scoring across 60+ gene sets, ligand-receptor interaction prediction using 280+ curated pairs, condition-aware comparative analysis, and cell-type proportion estimation, all operating directly on embedded data without access to the original count matrix. Benchmarked across six diverse scRNA-seq datasets spanning inflammatory lung disease, pediatric and adult cancers, organoid models, healthy tissue, and neurodevelopment, ELISA significantly outperforms CellWhisperer, a classical lexical retriever (BM25), and a random baseline in cell type retrieval (combined permutation test, $p < 2\times 10^{-5}$ for each), with particularly large gains on gene-signature queries (Cohen's $d = 5.98$ for mean reciprocal rank). ELISA replicates published biological findings (mean composite score 0.88), and generates candidate hypotheses through grounded LLM reasoning, bridging the gap between transcriptomic data exploration and biological discovery.

Generative Artificial Intelligence