PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Factual”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Analysis of Variants' Dynamic Using the CLIMB Database in COVID-19 Patients Admitted to Hospitals of Barts Health NHS Trust.

The COVID-19 pandemic, caused by SARS-CoV-2, has led to significant global health challenges. This study analyzes the dynamics of SARS-CoV-2 variants among patients admitted to Barts Health National Health Service (NHS) Trust hospitals using data from the CLIMB-COVID decentralized digital infrastructure allowing precise identification of SARS-CoV-2 variants. A total of 423 patients admitted between October 2020 and March 2021 were included in the study and divided into two groups: the alpha lineage group, which comprised the B.1.1.7 variant, and the other lineages group, which included all other variants. Whole-genome sequencing of SARS-CoV-2 genomes was conducted using the COVID-CLIMB pipelines. Clinical outcomes, such as mortality rates and deterioration within 28 days, were analyzed. To ensure robust findings, analyzes were adjusted for confounding factors, including age and comorbidities. Our findings revealed a significant increase in mortality with age for the alpha lineage and other lineages. The study underscores the importance of age adjustment in clinical studies to accurately assess the impact of different variants. Consistent genomic sequencing and data completeness are crucial for obtaining reliable results and guiding public health responses. These insights are vital for improving patient outcomes and providing a truthful picture of the pandemic, informing both current and future healthcare strategies.

Humans

Covid-19 Vaccination During Pregnancy in France: a Descriptive Study of Uptake Using the National Healthcare data System.

Pregnant women and their fetus are at increased risk of complications when infected by SARS-CoV-2. This study describes COVID-19 vaccine uptake during pregnancy in France (i) according to socio-demographic characteristics, and (ii) over calendar time and pregnancy stages. Using the hospital discharge records of the national healthcare database, we identified women ending pregnancy at hospital between Dec 27, 2020, i.e. beginning of Covid-19 vaccine availability, and Dec 31, 2022. Covid-19 vaccinations and SARS-CoV-2 infections were also available on two specific linked databases. Vaccine coverage was defined as uptake of at least one dose of vaccine during the 2-year study period, or more specifically during one of five phases of pregnancy (first, second and third trimesters, pre-conceptional and post-pregnancy periods). We also defined ineligibility as the six-month period following each Covid-19 infection. We calculated weekly rates of vaccination by pregnancy phase, accounting for ineligibility. About 75% of women received at least one dose of vaccine in 2021-2022 vs. 90% in the general population with a similar age structure. About 26% received at least one dose of vaccine while pregnant. The rate was lower among younger women, deprived women, and those living in the south-east of mainland France or in overseas territories. Weekly vaccination rates according to the phase of pregnancy showed that women mostly received vaccination outside the pregnancy trimesters (after or, to a less extent, before pregnancy). Vaccinations during pregnancy mostly occurred in the second trimester. When only elective terminations were considered, weekly rates of vaccination were almost identical, whatever the pregnancy phase (before, during or after pregnancy). Overall, our results show the positive impact of national recommendations on uptake dynamics. Vaccine coverage in pregnant women could be improved with targeted campaigns.

Humans

Published Database Resources for Traditional, Complementary, and Integrative Medicine: Update of a Systematic Review.

BACKGROUND: Traditional, Complementary, and Integrative Medicine (TCIM) has been established in the academic context of universities. In recent years, strategies have been developed worldwide to strengthen the role of TCIM in supporting the health of the population. Online databases are a common way for obtaining evidence-based information. This article is an update of a former systematic review from 2010 on published databases resources for TCIM. METHODS: The databases CINAHL, CAMbase, Web of Science, MEDLINE/PubMed, and Google Scholar search engine were searched for databases related to TCIM published in peer-reviewed journals between 2010 and November 2024. All included databases were visited online, and information on the origin, content, and scope of the database was extracted. RESULTS: A total of 6579 articles were identified through the literature search. After exclusion of irrelevant articles, full-text screening of 127 articles yielded 37 new databases. Together with 16 still available old databases, these mainly contained information on herbal therapies (n = 15) and Traditional Chinese Medicine (n = 11) from 18 different countries. Newly identified medicinal plant databases offer various scientific resources such as crude drugs, indigenous plants, and structures for natural and phytochemical components with molecular biological content. CONCLUSIONS: This literature review illustrates the dynamic development in the database landscape over the last 15 years. While the number of bibliographic databases is shrinking, databases in the field of medical plants/herbal therapy content are on the rise, which might be due to advances in plant genomics and molecular biology.

Humans

Association between SGLT2 inhibitors and reporting of phimosis/paraphimosis: a comparative pharmacovigilance analysis of the WHO database.

PURPOSE: Recent data have discussed occurrence of phimosis with Sodium-glucose co-transporter-2 (SGLT2) inhibitors. However, the potential risk among the different SGLT2 inhibitors is unknown. METHODS: Using Individual Case Safety Reports (ICSRs) registered in the WHO pharmacovigilance database (01/01/2000-30/06/2025), comparisons between the different SGLT2 inhibitors and versus other drugs used in diabetes (DUD) were performed. Results are shown as Reporting Odds Ratios (ROR). RESULTS: Among 11 342 810 ICSRs, 227 were phimosis/paraphimosis with SGLT2 inhibitors, mainly between 45 and 64 years. The higher ROR value was found with empagliflozin followed by dapagliflozin and canagliflozin. ROR for SGLT2 inhibitors was higher that of all other DUD [34.72 (25.86-46.62)]. The reporting risk of phimosis/paraphimosis with SGLT2 inhibitors was also higher than that of each pharmacological class of DUD. CONCLUSION: The results suggest an association between SGLT2 inhibitors use and phimosis ICSRs. Empagliflozin had the higher reporting risk.

Humans

Network-based integration of metabolomics data from large-scale repositories.

INTRODUCTION: Public metabolomics data repositories such as MetaboLights and Metabolomics Workbench host rapidly growing volumes of raw data, processed results, and metadata. As data deposition becomes a prerequisite for funding and publication, there is an increasing need for tools that enable integration and joint reanalysis of datasets across studies to maximise reuse and reproducibility. OBJECTIVES: This study aims to enable large-scale integrative meta-analysis of public metabolomics data, exploiting harmonised metabolite annotations to identify robust multi-study metabolite and pathway signatures and to provide global visual overviews of repository content. METHODS: We developed a network-based integration framework operating at both the study (dataset) level and the metabolite or pathway level. Metabolite-level meta-networks integrate studies with shared biological context using co-occurrences of differential metabolites represented as bipartite graphs. Study-level networks compare observed metabolites for overall repository exploration. Networks can be explored interactively using a dedicated Python Dash app available at https://github.com/EloisaRL/Metabolomic-data-analysis-app/tree/main . RESULTS: As an example, the approach was applied to six COVID-19 plasma datasets from MetaboLights generated using LC-MS and NMR. Ten metabolites were identified as differential in at least three studies, including consistently up-regulated pyroglutamic acid, in agreement with the literature. Pathway-level networks provided an overview of shared biological processes across studies. A global network of 1,181 studies in Metabolomics Workbench demonstrated clustering by assay coverage and associated metadata, as expected. CONCLUSION: Network-based integration of harmonised metabolomics data enables robust cross-study analyses and highlights the critical importance of standardised annotation pipelines. Such approaches enhance the reuse, reproducibility, and impact of public metabolomics datasets, accelerating biological discovery.

Metabolomics

An adjuvant database for preclinical evaluation of vaccines and immunotherapeutics.

Adjuvants are immunostimulators used to enhance vaccine efficacy against infectious diseases. However, current methods for evaluating their efficacy and safety are limited, hindering large-scale screening. To address this, we developed a prototype Adjuvant Database (ADB) containing transcriptome data, generated using the same protocols as the widely used Open TG-GATEs (OTG) toxicogenomics database, covering 25 adjuvants across multiple species, organs, time points, and doses. This enabled cross-database integration of ADB and OTG. Transcriptomic patterns successfully distinguished each adjuvant regardless of organs or species. Using both databases, we built machine learning models to predict adjuvanticity and hepatotoxicity. Notably, we identified colchicine's adjuvant activity and FK565's liver toxicity through data-driven analysis. Overall, ADB combined with OTG offers a framework for transcriptomics-based, data-driven screening of adjuvant candidates.

Animals

Impact of homologous recombination repair gene mutations on survival in metastatic prostate cancer: A real-world analysis from an observational database.

The prognostic significance of homologous recombination repair gene (HRRg) mutations across the different metastatic prostate cancer stages remains unclear. This retrospective real-world study analyzed 162 metastatic castration-sensitive (mCSPC) and 126 castration-resistant (mCRPC) patients from the ProGène database, stratified by HRRg mutational status. Mutation prevalence was similar in both groups (16.0% in mCSPC vs. 13.5% in mCRPC). HRR-positive mCSPC patients had significantly shorter median overall survival (OS) (24.0months; 95% confidence interval [CI]: 16.0-41.0) compared to HRR-negative patients (45.0months; 95% CI: 34.0-69.0; P=0.04). Notably, BRCA2-mutated patients exhibited a reduced median OS of 24.0months (95% CI: 9.0-40.0; P=0.036) and a faster progression-free survival compared to HRR-negative patients (median PFS=8.0months; 95% CI: 0.0-14.0 vs. 17.0months; 95% CI: 12.0-20.0; P=0.006). These findings suggest that HRRg mutations - especially BRCA2 - are associated with worse prognosis in mCSPC, supporting the value of early genomic screening to guide personalized treatment strategies.

Humans

Development and validation of a comprehensive prognostic model for 28-day ICU mortality in non-traumatic subarachnoid hemorrhage: an analysis based on the MIMIC-IV database.

BACKGROUND: Due to the complex pathophysiology of non-traumatic subarachnoid hemorrhage (SAH), accurate risk prediction remains a challenge. Our aim is to develop and validate a comprehensive prognostic model that integrates demographic characteristics, vital signs, laboratory parameters, and more, to provide clinical decision-making support in real-world practice. METHODS: We conducted a retrospective cohort study of 785 Non-traumatic subarachnoid hemorrhage patients. The cohort was randomly divided into a training set (n = 549) and a validation set (n = 236). Feature selection was performed using LASSO regression, followed by backward stepwise Cox regression for optimization. A nomogram was constructed based on independent predictive factors, and model performance was assessed using discrimination, calibration, and decision curve analysis. To prevent immortal-time bias, all predictors were anchored to a fixed early (first-24-hour) measurement window, treatment variables were modelled as binary indicators rather than cumulative exposures, and a five-model sensitivity analysis with baseline-severity adjustment was performed. RESULTS: The development of our model followed a systematic approach: first, 15 potential predictive factors were selected via LASSO regression, which were then refined to 12 independent predictors using backward stepwise Cox regression. The final predictive factors included: Ventilation, AHT, Nimodipine 60 mg, Age, SAPS.II, Input amount, Calcium total, Platelet count, White blood cells, Anion gap, pH, and Chloride. The integrated model demonstrated excellent predictive ability for 7-day, 14-day, and 21-day mortality in both the training set (AUC: 0.972, 0.934, 0.898) and the validation set (AUC: 0.968, 0.948, 0.911). Calibration curves and decision curve analysis confirmed the model's reliability and clinical utility across different time points. We constructed a nomogram for individualized risk prediction. Univariate Kaplan-Meier survival analysis demonstrated significant stratification of survival outcomes by each predictor, while restricted cubic spline analysis revealed non-linear relationships between continuous variables and mortality risk. Random survival forest analysis identified the top three predictive factors (Nimodipine 60 mg, Ventilation, AHT) and compared them with our full 12-variable model, confirming superior performance of the integrated model at all time points. At the 28-day primary endpoint, the model achieved a time-dependent AUC of 0.898 (training) and 0.904 (validation); after restricting predictors to the early baseline window, the leakage-controlled model retained good discrimination (validation C-index 0.803). CONCLUSIONS: Our ICU 28-day mortality prognosis model demonstrated robust performance in predicting ICU 28-day mortality in non-traumatic subarachnoid hemorrhage. The model, through the nomogram, provides individualized risk assessment, aiding clinical decision-making and patient stratification.

Humans

Chromosomal toxin-antitoxin systems in Pseudomonas putida are rather selfish than beneficial.

Chromosomal toxin-antitoxin (TA) systems are widespread genetic elements among bacteria, yet, despite extensive studies in the last decade, their biological importance remains ambivalent. The ability of TA-encoded toxins to affect stress tolerance when overexpressed supports the hypothesis of TA systems being associated with stress adaptation. However, the deletion of TA genes has usually no effects on stress tolerance, supporting the selfish elements hypothesis. Here, we aimed to evaluate the cost and benefits of chromosomal TA systems to Pseudomonas putida. We show that multiple TA systems do not confer fitness benefits to this bacterium as deletion of 13 TA loci does not influence stress tolerance, persistence or biofilm formation. Our results instead show that TA loci are costly and decrease the competitive fitness of P. putida. Still, the cost of multiple TA systems is low and detectable in certain conditions only. Construction of antitoxin deletion strains showed that only five TA systems code for toxic proteins, while other TA loci have evolved towards reduced toxicity and encode non-toxic or moderately potent proteins. Analysis of P. putida TA systems' homologs among fully sequenced Pseudomonads suggests that the TA loci have been subjected to purifying selection and that TA systems spread among bacteria by horizontal gene transfer.

Anti-Bacterial Agents

BioMedGraphica: an all-in-one platform for joint textual biomedical prior knowledge and numeric graph generation.

MOTIVATION: Multiomics data analysis is essential for scientific discovery in precision medicine. However, translating analysis results of omics data analysis into novel scientific hypotheses remains a significant challenge. Human experts must manually review analysis results and generate new hypotheses based on extensive and interconnected biomedical prior knowledge, which is subjective and not scalable. While large language models can accelerate the discovery, their reasoning improves when grounded in structured, auditable, and comprehensive biomedical prior knowledge. However, biomedical knowledge is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of artificial intelligence systems to fully leverage biomedical data for scientific discovery. RESULTS: We developed BioMedGraphica, a novel all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified textual prior knowledge graph containing 2 306 921 entities and 27 232 091 relations. In addition, we present a novel textual-numeric graph (TNG) data structure concept, where textual information captures prior biological knowledge (e.g. transcription start sites, functions, mechanisms), numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data structure for developing novel graph analysis models. AVAILABILITY AND IMPLEMENTATION: The code is available at: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph database can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica.

Humans

MedImg: An Integrated Database for Public Medical Images.

The advancements in deep learning algorithms for medical image analysis have garnered significant attention in recent years. While several studies have shown promising results, with models achieving or even surpassing human performance, translating these advancements into clinical practice is still accompanied by various challenges. A primary obstacle lies in the availability of large-scale, well-characterized datasets for validating the generalization of approaches. To address this challenge, we curated a diverse collection of medical image datasets from multiple public sources, containing 105 datasets and a total of 1,995,671 images. These images span 14 modalities, including X-ray, computed tomography, magnetic resonance imaging, optical coherence tomography, ultrasound, and endoscopy, and originate from 13 organs, such as the lung, brain, eye, and heart. Subsequently, we constructed an online database, MedImg, which incorporates and systematically organizes these medical images to facilitate data accessibility. MedImg serves as an intuitive and open-access platform for facilitating research in deep learning-based medical image analysis, accessible at https://www.cuilab.cn/medimg/.

Humans

Sociodemographic trends in prostate cancer: insights from the All of Us Research Program.

BACKGROUND: Prostate cancer disproportionately affects vulnerable populations. The All of Us Research Program (AoURP) is a database that aims to encapsulate the diversity of the United States. To explore the utility of this dataset in assessing prostate cancer disparities, we investigated whether treatment usage, disease progression, and genomic research participation vary across sociodemographic factors among AoURP participants with prostate cancer. METHODS: We identified AoURP participants with prostate cancer. Genomic research participation in AoURP, treatment usage, time-to-treatment, and time-to-metastasis were assessed by demographics and distance from a National Cancer Institute-designated comprehensive cancer center. Multivariable logistic regression and Cox proportional hazards regression were performed to evaluate treatment usage and time-to-treatment and time-to-metastasis, respectively. RESULTS: We observed lower genomic data availability in Black vs White patients (P&#x2009;<&#x2009;.001). In multivariable analyses, patients residing more than 80 miles from an NCI-designated comprehensive cancer center were less likely to receive androgen receptor pathway inhibitors (odds ratio [OR]&#x2009;=&#x2009;0.30, 95% CI = 0.14 to 0.66; P&#x2009;=&#x2009;.002) and bone targeting agents (OR&#x2009;=&#x2009;0.46, 95% CI = 0.30 to 0.70; P&#x2009;<&#x2009;.001) but more likely to undergo prostatectomy (OR&#x2009;=&#x2009;1.97, 95% CI = 1.43 to 2.71; P&#x2009;<&#x2009;.001) than those&#x2009;residing less than&#x2009;40 miles away. These patients also initiated treatment faster (hazard ratio [HR]&#x2009;=&#x2009;1.54, 95% CI = 1.27 to 1.87; P&#x2009;<&#x2009;.001) and developed metastasis slower (HR&#x2009;=&#x2009;0.58, 95% CI = 0.40 to 0.86; P&#x2009;=&#x2009;.006). Black patients were less likely to receive radiation (OR&#x2009;=&#x2009;0.45, 95% CI = 0.23 to 0.88; P&#x2009;=&#x2009;.020), prostatectomy (OR&#x2009;=&#x2009;0.65, 95% CI = 0.44 to 0.96; P&#x2009;=&#x2009;.028), and bone targeting agents (OR&#x2009;=&#x2009;0.65, 95% CI = 0.45 to 0.93; P&#x2009;=&#x2009;.018) than White patients. CONCLUSIONS: Prostate cancer treatment usage, disease progression, and genomic research participation varied between demographic populations. As AoURP matures, additional studies may leverage future data releases to confirm these findings.

Aged

EMPIAR: the Electron Microscopy Public Image Archive.

Public archiving in structural biology is well established with the Protein Data Bank (PDB; wwPDB.org) catering for atomic models and the Electron Microscopy Data Bank (EMDB; emdb-empiar.org) for 3D reconstructions from cryo-EM experiments. Even before the recent rapid growth in cryo-EM, there was an expressed community need for a public archive of image data from cryo-EM experiments for validation, software development, testing and training. Concomitantly, the proliferation of 3D imaging techniques for cells, tissues and organisms using volume EM (vEM) and X-ray tomography (XT) led to calls from these communities to publicly archive such data as well. EMPIAR (empiar.org) was developed as a public archive for raw cryo-EM image data and for 3D reconstructions from vEM and XT experiments and now comprises over a thousand entries totalling over 2 petabytes of data. EMPIAR resources include a deposition system, entry pages, facilities to search, visualize and download datasets, and a REST API for programmatic access to entry metadata. The success of EMPIAR also poses significant challenges for the future in dealing with the very fast growth in the volume of data and in enhancing its reusability.

Imaging, Three-Dimensional

IMAGE cDNA clones, UniGene clustering, and ACeDB: an integrated resource for expressed sequence information.

In this study we describe a new information resource that provides integrated access to information on IMAGE (integrated molecular analysis of genomes and their expression) cDNA library clones and derived expressed sequence tags (ESTs). We have developed an automated procedure that collates data from various public sources into a single ACeDB database. This database is a valuable tool for electronic cloning experiments and gene expression studies. It allows researchers to find information about cDNA libraries, plate addresses, insert sizes, and sequence data for IMAGE clones, the assignment of ESTs to UniGene clusters, and the chromosomal location of those genes in an efficient, graphically oriented manner.

Cloning, Molecular

Unveiling the BMI Risk Threshold for Osteoarthritis: Multi-Database Causal and Nonlinear Evidence.

OBJECTIVE: To characterize the nonlinear relationship between BMI and osteoarthritis (OA), and to identify BMI thresholds that inform precise prevention strategies. METHODS: This multi-database study integrated Global burden of disease&#xa0;2021, National Health and Nutrition Examination Survey 2007-2018, and Genome-Wide Association Studies. A generalized additive model was performed to visualize the BMI-OA relationship, adjusting for multiple confounders. We applied segmented logistic regression models to identify potential threshold effects and used Mendelian randomization to estimate the causal effects of BMI on OA subtypes. RESULTS: From 1990 to 2021, the age-standardized prevalence and years lived with disability rates for OA were highest in regions with high SDI. OA prevalence rose nonlinearly with BMI, with breakpoints at 24.00 and 41.58&#x2009;kg/m2. Each unit increase in BMI was associated with higher odds of OA between 24.00 and 41.58&#x2009;kg/m2 (OR&#x2009;=&#x2009;1.022, 95% CI: 1.003-1.041) and above 41.58&#x2009;kg/m2 (OR&#x2009;=&#x2009;1.055, 95% CI: 1.022-1.090). Women and individuals aged &#x2265;&#x2009;45&#x2009;years exhibited a higher susceptibility to knee osteoarthritis. BMI was causally associated with knee osteoarthritis (OR&#x2009;=&#x2009;1.63, 95% CI 1.50-1.77) and hip osteoarthritis (OR&#x2009;=&#x2009;1.54, 95% CI 1.40-1.70). CONCLUSIONS: These findings suggest that OA risk awareness and weight-management strategies should begin before BMI reaches the high range, particularly among individuals with BMI exceeding 24.00&#x2009;kg/m2.

Humans

Metformin Adherence and Risk of Polyneuropathy in Type 2 Diabetes Mellitus: An International Matched Cohort Study with Independent Validation.

BACKGROUND: Metformin is a popular first-line glucose-lowering medication for type 2 diabetes mellitus (T2DM). Although metformin reduces the risks of various complications of diabetes, its potential to cause polyneuropathy by depleting vitamin B12 levels is concerning. This study investigated whether the adherence or discontinuation of metformin after adding-on a second-line antiglycemic agent increases the risk of polyneuropathy in patients with T2DM. METHODS: Data from TriNetX were obtained, and patients with T2DM who were receiving second-line antiglycemic agents were divided into metformin-adherent and metformin-nonadherent groups based on prescription claims data. Neuropathy incidence was evaluated using diagnostic claims and nerve conduction examinations. For independent confirmation and external validation of the primary findings, we used data from the National Health Insurance Research Database (NHIRD) of Taiwan. RESULTS: After matching, 58,027 patients were included in each group. Compared with metformin adherent patients, metformin nonadherent patients had a higher risk of polyneuropathy (adjusted hazard ratios [aHR] 1.26; 95% confidence interval [CI] 1.23-1.29; P < 0.001). Risks of diabetic foot ulcer, amputation, neuropathy-related medication use, and bone fracture were also higher among nonadherent patients. Sensitivity analyses confirmed the robustness of findings. In the validation NHIRD cohort (31,384 matched pairs), metformin nonadherence remained associated with increased polyneuropathy risk (aHR 1.25; 95% CI 1.10-1.42; P < 0.001). CONCLUSIONS: Metformin adherence in patients with T2DM who require second-line treatment may reduce the risk of polyneuropathy; vitamin B supplementation may enhance this benefit.

Humans

The Open Microscopy Environment (OME) Data Model and XML file: open tools for informatics and quantitative analysis in biological imaging.

The Open Microscopy Environment (OME) defines a data model and a software implementation to serve as an informatics framework for imaging in biological microscopy experiments, including representation of acquisition parameters, annotations and image analysis results. OME is designed to support high-content cell-based screening as well as traditional image analysis applications. The OME Data Model, expressed in Extensible Markup Language (XML) and realized in a traditional database, is both extensible and self-describing, allowing it to meet emerging imaging and analysis needs.

Computational Biology

Worldwide Innovative Network Consortium: Building a Common Global Cancer Database.

This review shares the ongoing work of the global Worldwide Innovative Network (WIN) Consortium for Precision Medicine to synthesize emerging cancer treatment data and to define the requirements for a common global cancer database that can truly support precision oncology. We performed a narrative review of emerging cancer treatment data, molecular profiling technologies, and existing clinicogenomic databases, focusing on how tumors are characterized, how subgroups are defined, and how demographic, lifestyle, and environmental factors are captured. The growth in molecular profiling technologies and the development of new targeted therapies are transforming cancer care. Tumors, regardless of tissue origin, are increasingly defined as composites of multiple, often rare, subgroups, each with distinct biology and likely response to specific therapies, based on multidimensional profiling of the tumor and its microenvironment. The solution lies in building vast databases that capture racial and ethnic diversity, reflected in genomic data, as well as diet and lifestyle factors that may have epigenetic impact on gene expression and post-translational modifications. A truly inclusive and informative data set must reflect global diversity, and there are multiple examples of demography-dependent differences in genomic signals. With members caring for and studying patients with cancer across five continents, WIN is actively exploring pathways to create a global cancer database, rich in clinical and molecular detail, granular enough for precise analysis, and large enough to power artificial intelligence-driven insights, provided appropriate data quality, validation, and governance frameworks are in place. This review surveys the current landscape and outlines practical paths forward to achieve this goal.

Humans