PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data curation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

PreBIND and Textomy--mining the biomedical literature for protein-protein interactions using a support vector machine.

BACKGROUND: The majority of experimentally verified molecular interaction and biological pathway data are present in the unstructured text of biomedical journal articles where they are inaccessible to computational methods. The Biomolecular interaction network database (BIND) seeks to capture these data in a machine-readable format. We hypothesized that the formidable task-size of backfilling the database could be reduced by using Support Vector Machine technology to first locate interaction information in the literature. We present an information extraction system that was designed to locate protein-protein interaction data in the literature and present these data to curators and the public for review and entry into BIND. RESULTS: Cross-validation estimated the support vector machine's test-set precision, accuracy and recall for classifying abstracts describing interaction information was 92%, 90% and 92% respectively. We estimated that the system would be able to recall up to 60% of all non-high throughput interactions present in another yeast-protein interaction database. Finally, this system was applied to a real-world curation problem and its use was found to reduce the task duration by 70% thus saving 176 days. CONCLUSIONS: Machine learning methods are useful as tools to direct interaction and pathway database back-filling; however, this potential can only be realized if these techniques are coupled with human review and entry into a factual database such as BIND. The PreBIND system described here is available to the public at http://bind.ca. Current capabilities allow searching for human, mouse and yeast protein-interaction information.

Algorithms↗

NBA-Palm: prediction of palmitoylation site implemented in Naïve Bayes algorithm.

BACKGROUND: Protein palmitoylation, an essential and reversible post-translational modification (PTM), has been implicated in cellular dynamics and plasticity. Although numerous experimental studies have been performed to explore the molecular mechanisms underlying palmitoylation processes, the intrinsic feature of substrate specificity has remained elusive. Thus, computational approaches for palmitoylation prediction are much desirable for further experimental design. RESULTS: In this work, we present NBA-Palm, a novel computational method based on Naïve Bayes algorithm for prediction of palmitoylation site. The training data is curated from scientific literature (PubMed) and includes 245 palmitoylated sites from 105 distinct proteins after redundancy elimination. The proper window length for a potential palmitoylated peptide is optimized as six. To evaluate the prediction performance of NBA-Palm, 3-fold cross-validation, 8-fold cross-validation and Jack-Knife validation have been carried out. Prediction accuracies reach 85.79% for 3-fold cross-validation, 86.72% for 8-fold cross-validation and 86.74% for Jack-Knife validation. Two more algorithms, RBF network and support vector machine (SVM), also have been employed and compared with NBA-Palm. CONCLUSION: Taken together, our analyses demonstrate that NBA-Palm is a useful computational program that provides insights for further experimentation. The accuracy of NBA-Palm is comparable with our previously described tool CSS-Palm. The NBA-Palm is freely accessible from: http://www.bioinfo.tsinghua.edu.cn/NBA-Palm.

Acyltransferases↗

PAHG: the database of human multi-gene families.

BACKGROUND: In the early vertebrate history, gene duplications, including single-gene, segmental-gene (SSD), and whole-genome duplication (WGD), formed multigene families. Despite efforts to classify metazoan multigene families hierarchically for evolutionary insight, a gap exists in accessible, curated resources for human/vertebrate multigene families. RESULTS: Addressing this, we present the Phylogenomic Analysis of Human Genome (PAHG) database. It focuses on curated multigene families in the human genome, particularly within four paralogons: HOX-bearing (Hsa:2/7/12/17), FGFR-bearing (Hsa:4/5/8/10), MHC-bearing (Hsa:1/6/9/19), and chromosomes 1/2/8/20. CONCLUSION: The current PAHG version details the phylogenetic history of 221 human multigene families (1247 gene members) with 15,231 protein sequences from diverse metazoans. It provides insights into gene duplication timings, co-duplication events, and their relationships with human genome syntenic organization. The PAHG database addresses the lack of accessible resources, offering valuable information on human/vertebrate multigene family evolution. Access the PAHG database at: https://www.pahgncb.com/ and http://pahg.qau.edu.pk/ . This resource enriches our understanding of vertebrate genetic evolution.

Humans↗

Unveiling microbial risks in Chinese household dust: a comprehensive analysis from absolute abundance to virulence unit.

BACKGROUND: People spend the majority of their lives indoors, yet the risk and virulence potential of household microbiota remain largely unexplored, particularly in developing countries. RESULTS: Here, we conducted a nationwide survey on both dust samples and health information across 118 Chinese households. The microbiota composition and its functional units were analyzed using absolute 16S rRNA/ITS sequencing, metagenomics, and metaproteomics. Cross-domain network analysis of the core microbial communities revealed robust co-occurrence patterns in household dust. The mean absolute abundance of potentially pathogenic bacteria and fungi in households was 2.39 × 105 and 2.83 × 106 DNA copies/g dust. The potentially pathogenic community was primarily influenced by latitude, relative humidity, and average temperature. Although total absolute abundance was substantially lower in urban areas, the relative abundance of potentially pathogenic bacteria was markedly higher compared to rural environments. While urban-rural differences existed, the underlying statistical drivers were the environmental variables. The absolute abundance of potential pathogens was significantly associated with the prevalence of rhinitis, wheeze, and dermatitis in 266 participants. Children were identified as the highest-risk group from inhalation exposure of average daily dose. A total of 170 bacterial, 223 fungal virulence factors (VFs), and 370 antibiotic resistance genes (ARGs) were detected in dust and dust extracellular vesicle (EV)-associated DNA. EV-associated cargoes contributed 47.13% to the bacterial VF profiles, 11.90% to fungal VF profiles, and 44.45% to ARG profiles. Metaproteomic analysis confirmed the presence of VF profiles in dust EVs, which was further verified by curated proteomics data from 35 household pathogens. CONCLUSIONS: This study provides a comprehensive, quantitative framework linking indoor microbial exposure to health risks, highlighting EVs as a non-negligible, novel, extracellular mechanistic pathway for health impact in household environments. Video Abstract.

Child↗

WebQTL: web-based complex trait analysis.

WebQTL is a website that combines databases of complex traits with fast software for mapping quantitative trait loci (QTLs) and for searching for correlations among traits. WebQTL also includes well-curated genotype data for five sets of mouse recombinant inbred (RI) lines. Thus, to identify QTLs, users need provide only quantitative trait data from one of the supported populations. The WebQTL databases include both biological traits--neuroanatomical, pharmacological, and behavioral traits--and microarray-based gene expression data from BXD RI lines. A search function finds correlations between RNA expression and biological traits, and mapping functions find QTLs for either type of trait. The WebQTL service is available at http://www.webqtl.org/.

Animals↗

Optimizing treatment of chronic myeloid leukemia: a rational approach.

Imatinib mesylate, a novel, molecularly targeted agent for the treatment of chronic myeloid leukemia (CML), has expanded the management options for this disease and provided a paradigm for the treatment of other cancers. Imatinib is a potent, specific inhibitor of BCR-ABL, the constitutively active protein tyrosine kinase critical to the pathogenesis of CML. A randomized, phase III comparison of imatinib with interferon-alfa plus cytarabine as initial treatment for newly diagnosed chronic-phase CML, which demonstrated significantly higher rates of disease response with less toxicity, better quality of life, and a significantly longer progression-free survival time, provided the most persuasive data supporting a major role for imatinib. Currently, allogeneic stem cell transplantation is the only treatment modality with long-term data demonstrating curative potential in CML. An option for less than half of CML patients and associated with substantial morbidity and mortality, transplantation may still be appropriate initial therapy for certain patients. Busulfan and hydroxyurea have no demonstrable effect on disease natural history. The interferon-plus-cytarabine combination can induce durable cytogenetic remissions and was previously the CML pharmacotherapy standard of care, but it is often poorly tolerated. Imatinib is now indicated as first-line therapy for CML in all phases.

Antineoplastic Agents↗

Molecular epidemiology of Neisseria meningitidis.

Neisseria meningitidis poses a major disease burden on human beings. Meningococcal typing has a longstanding tradition for epidemiological surveillance of the disease. Genetic and antigenetic variability resulting from horizontal genetic exchange has been exploited for this purpose. Neisseria meningitidis served as the bacterial prototype organism for the development of multi locus sequence typing, which has replaced multi locus enzyme electrophoresis as the gold standard for meningococcal typing. Due to the rapid emergence of new porin variants serotyping by monoclonal antibodies is currently being replaced by DNA sequencing of variable regions of the porA gene. Sequence data were used to characterize the population structure of meningococci in carriage and disease. The advances in molecular epidemiology of meningococcal disease, and the rapid exchange of DNA sequence data via curated internet websites have resulted in an interactive international network, which is capable of identifying newly emerging clones within weeks.

Humans↗

The Chromosome 6 database at the Sanger Centre.

The Sanger Centre Chromosome 6 Database (6ace) has been developed as the primary means of release of annotated sequencing and mapping information for human chromosome 6 from the Sanger Centre. It is also being used to curate global data from published and unpublished external sources. The rationale behind the development of 6ace is described, together with information as to how to access the database.

Base Sequence↗

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation↗

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility↗

[Viral hepatitis in hemodialysis and renal transplantation patients].

In patients with chronic renal failure be they hemodialysed or transplanted, viral hepatitis B and C tend to progress towards chronicity, in spite of both a frequent silent clinical presentation and an atypical course of viral markers. Therefore, only liver biopsy will allow a precise diagnosis of liver disease in these patients. Immunosuppression clearly modifies natural history of B and C viral hepatitis, but their real impact at mid and long-term on patient survival is still a matter of debate. Treatment is firstly preventive (vaccination and isolation of infected patients in dialysis units) and secondly curative. Preliminary data suggest that antiviral drugs such as ARA-AMP and interferon-alpha may have the same efficacy as in non-immunosuppressed patients. It is therefore of urgent need to evaluate both efficacy and tolerance of these antiviral drugs in patients on hemodialysis and in kidney transplant recipients.

Actuarial Analysis↗

matchprobes: a Bioconductor package for the sequence-matching of microarray probe elements.

UNLABELLED: The nucleotide sequences of the probes on a microarray can be used for a variety of purposes in the analysis of microarray experiments. We describe software and a paradigm for the creation of data packages for curating, distributing and working with probe sequence data in a uniform, across-types-of-microarrays manner. While the implementation is specific to the Bioconductor project, the ideas and general strategies are more general and could be easily adopted by other projects. AVAILABILITY: The R package matchprobes is available under LGPL at http://www.bioconductor.org SUPPLEMENTARY INFORMATION: The package contains documentation in the form of a vignette and manual pages.

Algorithms↗

The Genome Sequence DataBase: towards an integrated functional genomics resource.

During 1998 the primary focus of the Genome Sequence DataBase (GSDB; http://www.ncgr.org/gsdb ) located at the National Center for Genome Resources (NCGR) has been to improve data quality, improve data collections, and provide new methods and tools to access and analyze data. Data quality has been improved by extensive curation of certain data fields necessary for maintaining data collections and for using certain tools. Data quality has also been increased by improvements to the suite of programs that import data from the International Nucleotide Sequence Database Collaboration (IC). The Sequence Tag Alignment and Consensus Knowledgebase (STACK), a database of human expressed gene sequences developed by the South African National Bioinformatics Institute (SANBI), became available within the last year, allowing public access to this valuable resource of expressed sequences. Data access was improved by the addition of the Sequence Viewer, a platform-independent graphical viewer for GSDB sequence data. This tool has also been integrated with other searching and data retrieval tools. A BLAST homology search service was also made available, allowing researchers to search all of the data, including the unique data, that are available from GSDB. These improvements are designed to make GSDB more accessible to users, extend the rich searching capability already present in GSDB, and to facilitate the transition to an integrated system containing many different types of biological data.

Animals↗

GrainGenes 2.0. an improved resource for the small-grains community.

GrainGenes (http://wheat.pw.usda.gov) is an international database for genetic and genomic information about Triticeae species (wheat [Triticum aestivum], barley [Hordeum vulgare], rye [Secale cereale], and their wild relatives) and oat (Avena sativa) and its wild relatives. A major strength of the GrainGenes project is the interaction of the curators with database users in the research community, placing GrainGenes as both a data repository and information hub. The primary intensively curated data classes are genetic and physical maps, probes used for mapping, classical genes, quantitative trait loci, and contact information for Triticeae and oat scientists. Curation of these classes involves important contributions from the GrainGenes community, both as primary data sources and reviewers of published data. Other partially automated data classes include literature references, sequences, and links to other databases. Beyond the GrainGenes database per se, the Web site incorporates other more specific databases, informational topics, and downloadable files. For example, unique BLAST datasets of sequences applicable to Triticeae research include mapped wheat expressed sequence tags, expressed sequence tag-derived simple sequence repeats, and repetitive sequences. In 2004, the GrainGenes project migrated from the AceDB database and separate Web site to an integrated relational database and Internet resource, a major step forward in database delivery. The process of this migration and its impacts on database curation and maintenance are described, and a perspective on how a genomic database can expedite research and crop improvement is provided.

Breeding↗

Curation of complex, context-dependent immunological data.

BACKGROUND: The Immune Epitope Database and Analysis Resource (IEDB) is dedicated to capturing, housing and analyzing complex immune epitope related data http://www.immuneepitope.org. DESCRIPTION: To identify and extract relevant data from the scientific literature in an efficient and accurate manner, novel processes were developed for manual and semi-automated annotation. CONCLUSION: Formalized curation strategies enable the processing of a large volume of context-dependent data, which are now available to the scientific community in an accessible and transparent format. The experiences described herein are applicable to other databases housing complex biological data and requiring a high level of curation expertise.

Allergy and Immunology↗

Metabolite coupling in genome-scale metabolic networks.

BACKGROUND: Biochemically detailed stoichiometric matrices have now been reconstructed for various bacteria, yeast, and for the human cardiac mitochondrion based on genomic and proteomic data. These networks have been manually curated based on legacy data and elementally and charge balanced. Comparative analysis of these well curated networks is now possible. Pairs of metabolites often appear together in several network reactions, linking them topologically. This co-occurrence of pairs of metabolites in metabolic reactions is termed herein "metabolite coupling." These metabolite pairs can be directly computed from the stoichiometric matrix, S. Metabolite coupling is derived from the matrix ŝŝT, whose off-diagonal elements indicate the number of reactions in which any two metabolites participate together, where ŝ is the binary form of S. RESULTS: Metabolite coupling in the studied networks was found to be dominated by a relatively small group of highly interacting pairs of metabolites. As would be expected, metabolites with high individual metabolite connectivity also tended to be those with the highest metabolite coupling, as the most connected metabolites couple more often. For metabolite pairs that are not highly coupled, we show that the number of reactions a pair of metabolites shares across a metabolic network closely approximates a line on a log-log scale. We also show that the preferential coupling of two metabolites with each other is spread across the spectrum of metabolites and is not unique to the most connected metabolites. We provide a measure for determining which metabolite pairs couple more often than would be expected based on their individual connectivity in the network and show that these metabolites often derive their principal biological functions from existing in pairs. Thus, analysis of metabolite coupling provides information beyond that which is found from studying the individual connectivity of individual metabolites. CONCLUSION: The coupling of metabolites is an important topological property of metabolic networks. By computing coupling quantitatively for the first time in genome-scale metabolic networks, we provide insight into the basic structure of these networks.

Chromosome Mapping↗

Adjuvant therapy with oral fluoropyrimidines as main chemotherapeutic agents after curative resection for colorectal cancer: individual patient data meta-analysis of randomized trials.

BACKGROUND: Oral 5-fluorouracil and its prodrugs (tegafur, carmofur) is now being studied for adjuvant chemotherapy of curatively resected colorectal cancers. To evaluate the effect of these oral fluoropyrimidines (o-FPs), an individual patient data (IPD) meta-analysis of randomized clinical trials was performed in Japan as an inter-trialist group study. METHODS: Data from the three clinical trials in which postoperative adjuvant therapy with o-FPs was compared with surgery alone in patients with colorectal cancer were sought. IPD from a total of 4960 patients with follow-up periods of at least 5 years were analyzed. RESULTS: The results of the meta-analysis on an 'intention to treat' basis demonstrated a significant benefit of o-FPs in terms of the disease-free survival (DFS) of the total patients [risk ratio (RR) 0.830, 95% confidence interval (CI) 0.742-0.929, P = 0.001]. o-FPs were also demonstrated to be effective for survival in rectal cancer (RR 0.857, 95% CI 0.734-0.999, P = 0.049) and in Dukes'C colorectal cancer (RR 0.828, 95% CI 0.711-0.965, P = 0.016). CONCLUSION: The results suggest the advantage of long term o-FPs, possibly with the injection of mitomycin C, for prognosis for curatively resected colorectal cancer patients.

Administration, Oral↗

[Factors in recurrence after curative resection of rectal cancer].

Data on 409 Patients who had a curative resection for rectal cancer at the National Taiwan University Hospital between 1977 and 1989 were analysed to determine the independent effect on recurrence. In our series, the total operative mortality rate was 1.7% and the resectability rate was 88.1%. For these cases who received curative resection, the overall 5-year survival rate was 62.6%. The 5-year survival rate varied according to the Dukes' stage: stage A, 96.4%; stage B, 73.3% and stage C, 39.6%. The total recurrence rate after curative resection was 35.9%, including local recurrence 18.3%, systemic recurrence 12.0%, and combined recurrence 5.6%. According to Dukes' staging system, the recurrence rate for stage A, B and C were 0%, 22.4% and 63.8%, respectively. We used Cox's regression model to analyse the patients characteristics and pathological variables on recurrence and to assess the independent influence of each when all other factors were held constant. The pathological stage of the cancer had the strongest association. Other variables found to have an independent yet significant importance, were CEA level and tumor size. The identification of the patient group, at a high risk of recurrence, might promote a more judicious selection for surgical procedure and trials of adjuvant therapy.

Adult↗