PubMed HealthSearch

SEARCH · PubMed Health

Results for “data curation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

PSIA: A Comprehensive Knowledgebase of Plant Self-incompatibility.

Self-incompatibility (SI) is an important genetic mechanism in angiosperms that prevents inbreeding and promotes outcrossing, with significant implications for crop breeding, including genetic diversity, hybrid seed production, and yield optimization. In eudicots, SI is typically governed by a single S-locus containing tightly linked pistil and pollen S-determinant genes. Despite major advances in SI research, a centralized, comprehensive resource for SI-related genomic data remains lacking. To address this gap, we developed the Plant Self-Incompatibility Atlas (PSIA), a systematically curated knowledgebase providing an extensive compilation of plant SI, including genomic resources for SI species, S gene annotations, molecular mechanisms, phylogenetic relationships, and comparative genomic analyses. The current release of PSIA includes over 500 genome assemblies from 469 SI species. Using known S genes as queries, we manually identified and rigorously curated 3700 S genes. PSIA provides detailed S-locus information from assembled genomes of SI species and offers an interactive platform for browsing, BLAST searches, S gene analysis, and data retrieval. Additionally, PSIA serves as a unique platform for comparative genomic studies of S-loci, facilitating exploration of the dynamic processes underlying the origin, loss, and regain of SI. As a comprehensive and user-friendly resource, PSIA will greatly advance our understanding of angiosperm SI and serve as a valuable tool for crop breeding and hybrid seed production. PSIA is freely available at http://www.plantsi.cn.

Self-Incompatibility in Flowering Plants

Proteomic insights into Helicobacter pylori infection in stomach cells, revealing host response and host-targeted therapeutics repurposing.

BACKGROUND: Helicobacter pylori (H. pylori) is a globally prevalent gastric pathogen strongly associated with chronic gastritis, peptic ulcers, and gastric cancer. While bacterial factors have been extensively studied, host proteomic responses and their therapeutic potential remain largely underexplored. RESEARCH DESIGN AND METHODS: Current analyses employed a systematic proteomics-based data integration and harmonization approach (retrospective qualitative cohort study) to identify important differentially regulated host proteins. Proteomic datasets were curated from in vitro studies and analyzed for functional enrichment, protein-protein interaction networks, and hub protein identification. To explore therapeutic repurposing, drug repositioning was performed using the DrugBank database. RESULTS: Data summation describing protein differential regulation in human gastric cells as a result of the infection revealed 1672 perturbed host proteins. Bioinformatics analysis revealed 11 proteins including CSK, MET, RELA, MARK2, GRB2, FTO, PLCG1, CRKL, RPS5, RPS9, and RPS27A to be ideal host targets for therapeutic repurposing. Clinically approved drugs such as Dasatinib (targeting CSK) and Crizotinib (targeting MET) emerged as promising candidates due to favorable pharmacokinetics and known bioactivity. CONCLUSIONS: Host-directed therapeutics could offer alternative strategies to conventional antibiotic therapy, addressing challenges such as resistance and infection recurrence, providing a foundation for future experimental validation and development of host-targeted interventions for infection control.

Humans

Prophylactic lymph node dissection in patients with advanced gastric cancer promotes increased survival time.

BACKGROUND: There is no consensus of opinion regarding the efficacy of lymph node dissection. METHODS: Data were analyzed from 452 patients with advanced gastric cancer who underwent curative resection in the Department of Surgery II, Kyushu University Hospital, between 1970 and 1985, with special reference to the lymph node metastasis. RESULTS: Metastatic lesions were evident in the dissected lymph nodes of 300 of 452 (66.4%) patients. Survival time for patients without lymph node metastasis was longer than for those with it (P less than 0.01). In patients without lymph node metastasis, the tumor was smaller, serosal invasion was less prominent, tumor growth was less infiltrating, and the tumor stage was, therefore, less advanced. Lymphatic involvement was found in 38.9% of the patients with no evidence of lymph node metastasis. CONCLUSIONS: Because the postoperative mortality rate is low in patients with lymph node dissection, the authors advocate prophylactic lymph node dissection to prevent a recurrence.

Aged

eVOC: a controlled vocabulary for unifying gene expression data.

Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.

Animals

Apollo: a sequence annotation editor.

The well-established inaccuracy of purely computational methods for annotating genome sequences necessitates an interactive tool to allow biological experts to refine these approximations by viewing and independently evaluating the data supporting each annotation. Apollo was developed to meet this need, enabling curators to inspect genome annotations closely and edit them. FlyBase biologists successfully used Apollo to annotate the Drosophila melanogaster genome and it is increasingly being used as a starting point for the development of customized annotation editing tools for other genome projects.

Animals

[Analysis of the quality of clinical diagnosis from generalized findings of the pathologoanatomic service].

A statistical analysis of generalized data of the pathoanatomical service on quality of clinical diagnosis in curative-prophylactic institutions in 54 administrative territories of the RSFSR was carried out. The structure (extensive indices) and frequency (intensive indices)of erroneous clinical diagnoses referring to the most important classes of diseases were identified. As to the structure of indices and frequency of clinico-anatomic disparities the first place was occupied by oncological diseases (20.1+/-0.11 and 14.2+/-0.22%), the second--by infectious diseases (16.5+/-0.1 and 13.0+/-0.34%), the third--by diseases of the digestive system (14.6+/-0.09 and 13.0+/-0.33%), the forth--by diseases of the urogenital system (14.0+/-0.09 and 12.2+/-0.49%), the fifth--by disease of the respiratory system (12.7+/-0.09 and 10.6+/-0.24%), the sixth--by diseases of the cardiovascular system (11.1+/-0.08 and 8.0+/-0.14%). The recommendation is put forward to carry on annually a complex satistical analysis of extensive and intensive indices of erroneous clinical diagnoses demonstrating the quality of clinical diagnosis in therapeutic institutions of a given administrative territory.

Diagnostic Errors

CancerTrialMatch: a computational resource for the management of biomarker-based clinical trials at a community cancer center.

MOTIVATION: The widespread implementation of next-generation sequencing in cancer care has enabled routine use of molecular and biomarker profiling. At our cancer center, as with many others, biomarker-based clinical trials are increasingly available to oncologists as potential treatment options via molecular tumor boards. To better support this effort, we developed CancerTrialMatch, a systematic approach to capture structured clinical trial data and match patients to trials based on their disease characteristics and sequencing profiles. RESULTS: CancerTrialMatch is an open-source application designed to streamline clinical trial curation and patient trial matching, while also enabling an institution's curated trial portfolio to be distributed across the institution for easy access to providers, care teams and researchers. It facilitates curating, updating, and searching for trials through a semi-automated interface built using R Shiny, MongoDB, and Docker. While much of the trial data is retrieved via the clinicaltrials.gov Application Programming Interface, certain items like biomarkers and disease subtypes are entered manually. The user inputs disease type using the OncoTree classification, and provides relevant biomarker details, such as mutations, copy numbers, fusions, and other disease-specific markers. This resource reduces the time required for institutional trial management and helps to identify potential clinical trials for patients, ultimately supporting larger clinical trial enrollment and enhancing the clinical application of precision oncology. AVAILABILITY AND IMPLEMENTATION: CancerTrialMatch was implemented and tested on Windows 11 (64-bit, 32 GB RAM) using WSL2 with Ubuntu 22.04. Docker 27.0.3 and Docker Compose 2.28.1 were used to build images and containers. Users can build it by cloning the repo and following the README instructions and supplemental file (cancertrialmatchsupplemental.pdf) . The source code and example data are available in GitHub and Figshare at https://github.com/AveraSD/CancerTrialMatch and 10.6084/m9.figshare.28447367 respectively.

Humans

[Research and surgical treatment of epilepsy].

The currently available surgical procedures for the treatment of epilepsy, from fundamental data to therapeutic results, including various means of investigation are reported. The work is based on a review of the literature and on the cases studied by the two teams from the Universities of Montreal and Bordeaux who share the same concept of epilepsy surgery. The patient groups of the two teams include 316 S.E.E.G., 214 cortectomies, 39 callosotomies and 2 multiple sub-pial transsections. In the first part, the authors attempt to demonstrate that the epileptic focus corresponds to the region where the seizures arise, that this focus is not directly comparable to the region where inter-ictal spikes are recorded and sometimes becomes autonomous from the causal lesion. The epileptic phenomenon has a definite harmful effect on cerebral functions and a probable self-aggravating potential. The second chapter summarizes the clinical data on which the indications and contraindications are based. These obviously depend on whether the intervention is intended to be curative or palliative. Various non-invasive and invasive investigations are then reviewed, according to their relative importance and the experience of each team. The main points developed are: the electroclinical correlations during seizures, the symptomatological data for differentiating between temporal and frontal lobe seizures, the contribution of M.R.I. in demonstrating the epileptogenic and epileptic lesions, the electrophysiological information suggesting that S.E.E.G. remains the most informative mean of investigation. The various methods of investigation of assessing electrical, functional (cerebral blood flow, metabolism) and morphological aspects of epilepsy, supply non-redondant findings about the localisation of the epileptic focus. The chapter on surgical techniques mainly discusses the various modes of implantation of subdural and intracerebral electrodes and reports the same rate of morbidity in both cases. Orthogonal teleradiography is still perfectly suited to the implantation of intracerebral electrodes. S.E.E.G. is still the most anatomically precise technique. However, in certain conditions, extraoperative E.Co.G. is more adequate. New surgical modalities have recently appeared such as the multiple subpial transsections which allow treatment of epileptic foci unapproachable by cortectomy and such as modified techniques of hemispherectomy, which by decreasing morbidity, renew interest in them. In the chapter on surgical results, the authors emphasize the methodological problems of evaluation that partly account for their wide variability. The results obtained with the various surgical modalities are reviewed. The outcome in cortectomies is discussed at length in terms of the data from the literature as well as the results reported by both teams.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent

Morphometry of gastric carcinoma: its association with patient survival, tumour stage, and DNA ploidy.

Morphometric image analysis of nuclear features was performed on tissue from 46 patients who had had curative resections for gastric cancer. Clinical, pathological, flow cytometric, and follow-up data were available for these patients, which were drawn from a larger, previously reported series. The morphometric data were compared with patient survival, clinico-pathological status, and DNA ploidy. Univariate survival analysis revealed that morphometric parameters were not significantly related to survival, but examination of clinico-pathological data showed lymph node involvement, involvement of the resection margin, and lymphatic invasion to be significantly associated (P < 0.01) with patient prognosis. Multivariate survival analysis using the Cox model found only lymph node and resection margin involvement to be independently related to survival. Comparison of morphometric results with the clinico-pathological parameters showed various features, relating to nuclear size, and its variation to be significantly associated (P < 0.01) with the presence of lymphatic invasion, resection margin involvement, and tumour pattern (intestinal/diffuse). A comparison of morphometry with flow cytometric analysis in these cases showed that nuclear size was not significantly related to either DNA aneuploidy or the DNA proliferative index.

Aged

Results of pulmonary arterial banding in infancy. Survey of 5 years' experience in the New England Regional Infant Cardiac Program.

The results of pulmonary arterial banding in 238 infants, 12 percent of the infants admitted to the New England Regional Infant Cardiac Program, is reviewed. Overall survival to age 1 year was 63 percent. Survival was least likely (37 percent) in those who required banding within the 1st month of life. Additional surgery decreased the survival rate in those operated on after 1 month of age. Infants with anomalies for which no corrective surgical procedure is available (23 of 238) have only a 30 percent chance of survival. Those with lesions correctable within the 1st year (133 of 238) have a 74 percent survival rate; 52 percent (82 of 238) of those for whom a curative operation is available after the 1st year survive. These pulmonary arterial banding data coupled with results of primary correction should provide the data base required for an intelligent decision in respect to appropriate surgical treatment of infants with critical heart disease.

Heart Defects, Congenital

Functional mapping of the Trypanosoma cruzi serinome by fluorophosphonate activity-based protein profiling.

Serine hydrolases (SHs) constitute one of the largest enzyme superfamilies in eukaryotes, yet their roles in Trypanosoma cruzi, the causative agent of Chagas disease, remain largely uncharacterized. Here, we report an activity-based chemoproteomic map of the T. cruzi epimastigote serinome by combining genome-informed in silico curation with whole-cell activity-based protein profiling (ABPP) using a panel of cell-permeable fluorophosphonate (FP)-alkyne probes. Whole-cell labelling followed by label-free quantitative proteomics (LFQ-MS) identified 37 enriched SH-like proteins, including 35 with conserved or partially conserved catalytic triad/dyad features, spanning lipases, peptidases, esterases, and previously uncharacterized hydrolases. The 35 SHs represent approximately 63% of the 56 predicted SHs retained after catalytic-site curation. Domain architecture analysis revealed broad structural diversity, while orthologue-based localization data suggested association with multiple subcellular compartments, including glycosomal, mitochondrial, and endosomal localizations. Gene Ontology enrichment highlighted lipid metabolic and catabolic processes as dominant functional themes, and protein-protein interaction network analysis supported functional connectivity among the captured enzymes. Several identified SHs, including oligopeptidase B, prolyl oligopeptidase Tc80, serine carboxypeptidase CPB1, and phospholipase A1 (PLA1) have previously been characterized in trypanosomatids, with roles linked to parasite virulence or host-pathogen interactions. Together, these findings establish a fluorophosphonate-based chemoproteomic resource for the kinetoplastid community and prioritize probe-accessible active T. cruzi SHs for future functional validation and antiparasitic inhibitor discovery.

Activity-based protein profiling

[Treatment of rhabdomyosarcoma in mice C3H/He by Co 60 and hyperbaric oxygen (author's transl)].

UNLABELLED: Experimental study of rhabdomyosarcoma with successive transplantation upon C3H/He mice, treated by irradiation (Co 60) and combined irradiation-hyperbaric oxygen (HBO), dating from 3, 14 and 15 days after transplantation. The data (tumor volume evolution, histological modifications, pulmonary metastases) are compared with controls. CONCLUSIONS: curative radiotherapy depends on starting treatment as soon as possible with or without HBO. After the 14th day, sensitisation to combined HBO and C60 is seen. The extension of pulmonary metastases is a function of tumor growth. Paradoxically metastases were less frequent after HBO only and more frequent after HBO-Co 60.

Animals

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts &#x223c;50&#x2009;000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship

BAV-LLPS: a database of bacterial, archaea, and virus liquid-liquid phase separation proteins.

MOTIVATION: Liquid-liquid phase separation (LLPS) is a key process underlying the formation of biomolecular condensates, such as membrane-less organelles, that compartmentalize biochemical processes inside the cells. While LLPS has been extensively studied in eukaryotes, its role in bacteria, archaea, and viruses remains far less characterized. Recent studies in bacteria have revealed that LLPS-driven condensates play critical roles in RNA processing, stress response, and pathogenicity. Similarly, many viruses exploit LLPS to facilitate crucial steps in their infection cycles, including viral entry, genome replication, assembly, and host immune evasion. RESULTS: In this work, we introduce a hand-curated database of LLPS proteins from bacteria, archaea, and viruses (BAV-LLPS Database). This resource, extended through sequence similarity searches, comprises over 5000 proteins and integrates diverse data including biological annotations, sequence features, predicted disordered regions, LLPS per site probability, and AlphaFold2-based structural models. Additionally, our web server enables users to explore both the curated and homologous derived datasets, providing a platform to uncover evolutionary relationships and intrinsic and differential properties of LLPS proteins across various taxonomic groups. This work seeks to deepen our understanding of LLPS mechanisms beyond eukaryotic organisms, emphasizing their significance across diverse life forms. It also aims to foster the development of specialized predictive tools that will facilitate the exploration and characterization of LLPS processes in a wide array of living organisms, thereby contributing to advancements in both fundamental biological research and applied biomedical sciences. AVAILABILITY AND IMPLEMENTATION: BAV-LLPS DB is freely accessible at https://bav-llps-db.bioinformatica.org/. The data can be retrieved from the website. The source code of the database can be downloaded from https://bav-llps-db.bioinformatica.org/download.

Databases, Protein

Renal cell carcinoma: resection of solitary and multiple metastases.

Between 1985 and 1991, 23 patients underwent resection of pulmonary metastases from renal cell carcinoma, of whom 18 had previously received interleukin-2 based immunotherapies. Mean survival from exploration in all patients was 43 months. Survival after resection did not correlate with the number of nodules on preoperative tomograms, the number of nodules resected, or the disease-free interval. Patients who underwent complete resection of metastatic disease (n = 15), however, had a significantly longer survival (mean, 49 months; median not yet achieved) compared with patients with incomplete resection (median, 16 months) (p2 = 0.02). Two of the 15 patients who underwent curative resections are presently free of disease greater than 45 months after exploration. These data support surgical resection of isolated pulmonary metastatic disease from renal cell cancer.

Adult

Assessment of Gene Set Enrichment Analysis using curated RNA-seq-based benchmarks.

Pathway enrichment analysis is a ubiquitous computational biology method to interpret a list of genes (typically derived from the association of large-scale omics data with phenotypes of interest) in terms of higher-level, predefined gene sets that share biological function, chromosomal location, or other common features. Among many tools developed so far, Gene Set Enrichment Analysis (GSEA) stands out as one of the pioneering and most widely used methods. Although originally developed for microarray data, GSEA is nowadays extensively utilized for RNA-seq data analysis. Here, we quantitatively assessed the performance of a variety of GSEA modalities and provide guidance in the practical use of GSEA in RNA-seq experiments. We leveraged harmonized RNA-seq datasets available from The Cancer Genome Atlas (TCGA) in combination with large, curated pathway collections from the Molecular Signatures Database to obtain cancer-type-specific target pathway lists across multiple cancer types. We carried out a detailed analysis of GSEA performance using both gene-set and phenotype permutations combined with four different choices for the Kolmogorov-Smirnov enrichment statistic. Based on our benchmarks, we conclude that the classic/unweighted gene-set permutation approach offered comparable or better sensitivity-vs-specificity tradeoffs across cancer types compared with other, more complex and computationally intensive permutation methods. Finally, we analyzed other large cohorts for thyroid cancer and hepatocellular carcinoma. We utilized a new consensus metric, the Enrichment Evidence Score (EES), which showed a remarkable agreement between pathways identified in TCGA and those from other sources, despite differences in cancer etiology. This finding suggests an EES-based strategy to identify a core set of pathways that may be complemented by an expanded set of pathways for downstream exploratory analysis. This work fills the existing gap in current guidelines and benchmarks for the use of GSEA with RNA-seq data and provides a framework to enable detailed benchmarking of other RNA-seq-based pathway analysis tools.

Humans

Future directions with carboplatin: can therapeutic monitoring, high-dose administration, and hematologic support with growth factors expand the spectrum compared with cisplatin?

This paper reviews the role of pharmacokinetic methods in the clinical use of carboplatin. Published data establish that pretreatment renal function is the most significant determinant of carboplatin pharmacokinetics. Evidence suggests that the area under the plasma concentration versus time curve (AUC) correlates well with hematologic toxicity. Published data also suggest that a relationship exists between the AUC and therapeutic outcome in testicular teratoma and in ovarian cancer, although in the latter case the data are not conclusive. It is suggested that pharmacokinetically based dosing schemes may be advantageous and that randomized trials should be performed to test this hypothesis. In some clinical situations where dose prediction is not feasible, a simple therapeutic drug-monitoring strategy may prove useful. Since the toxicities of carboplatin are mainly hematologic, it has been possible to study the use of high-dose carboplatin with various forms of hematologic support. Carboplatin doses have been increased fourfold to sixfold and high response rates have been reported in ovarian cancer. An overall therapeutic advantage for this strategy has not yet been demonstrated in a randomized setting. Published data on ovarian cancer cell lines suggest that the range of sensitivities encountered is very large (30- to 100-fold). If the range of sensitivities found in vivo is equally large, then a clinical dose escalation of 10- to 100-fold may be necessary to produce a curative therapy for this disease. The investigation of the use of hematopoietic growth factors in association with carboplatin has just begun. Early data suggest that some support of the WBC count may be achieved, but the toxicities of granulocyte-macrophage colony-stimulating factor have themselves been a problem. Our own studies will attempt to achieve a substantial escalation in the administered AUC of carboplatin by increasing the frequency of carboplatin dosing, while administering granulocyte colony-stimulating factor and platelet support.

Blood Transfusion

Rapid assessment of clinical severity for salmonellosis cases via protein family domain analysis and machine learning.

Salmonella is a common pathogen, infecting more than a million people yearly. Rapid assessment of clinical case severity is essential for improving patient outcomes and optimizing healthcare resources. Advancements in genome sequencing technologies have enabled the analysis of bacterial genomes from many clinical cases, opening up new opportunities for precise and timely diagnosis. This study proposes a genome-based framework for identifying critical Salmonella cases before the onset of critical symptoms and facilitating early medical intervention. By leveraging protein family (Pfam) domains as the representation for genomic data, the complex genetic profiles of Salmonella cases are simplified into interpretable features. The severity levels of cases were investigated through rigorous data analysis, resulting in a set of 70 Pfam domains that could be potentially used as biomarkers. Machine Learning was employed to assess the predictive power of the curated Pfam biomarkers, achieving high accuracy (~93%) in sorting cases into critical, moderate, and mild categories. The results demonstrate the efficacy of the proposed approach. This framework highlights the potential of using bacterial genomic data in clinical decision-making, opening the window for timely personalized interventions for Salmonella infection management.

Domains of unknown function (DUFs)