PubMed HealthSearch

SEARCH · PubMed Health

Results for “Cancer genomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

In silico generation of synthetic cancer genomes using generative AI.

Understanding how genomic alterations drive cancer is key to advancing precision oncology. To detect these alterations, accurate algorithms are used; however, due to privacy concerns, few deeply sequenced cancer genomes can be shared, limiting benchmarking and representing a major obstacle to the improvement of analytic tools. To address this, we developed OncoGAN, a generative AI model combining adversarial networks and variational autoencoders to create realistic synthetic cancer genomes. Trained on large-scale genomic datasets, OncoGAN accurately reproduces somatic mutations, copy number alterations, and structural variants across cancer types while preserving donors' privacy. The synthetic genomes reflect tumor-specific mutational signatures and positional mutation patterns. Using DeepTumour, we validated the synthetic data's fidelity, showing high concordance between generated and predicted tumors. Moreover, augmenting the training data with synthetic genomes improved DeepTumour's accuracy, underscoring OncoGAN's potential to generate shareable datasets with known ground truths for benchmarking and enhancement of cancer genome analysis tools.

Humans

Optimizing participant and community engagement in cancer genomic sequencing research.

PURPOSE: We describe strategies implemented across research centers of the Participant Engagement and Cancer Genome Sequencing (PE-CGS) Network to optimize engagement of participants and communities in cancer genomics research. We also present consensus definitions of engagement and engagement optimization, informed by our shared experiences in the Network. METHODS: Key informant interviews and a document review identified engagement and optimization strategies across PE-CGS research centers. Findings were synthesized using qualitative content analysis. Consensus on definitions of engagement and optimization were developed through iterative review by PE-CGS members. RESULTS: PE-CGS research centers adopted tailored strategies based on community needs and scientific gaps. Engagement strategies included community-based efforts (eg, advisory boards and newsletters) and participant-focused approaches (eg, enhanced informed consent and decision support tools). Optimization strategies leveraged scientific methods (eg, randomized controlled trials and surveys) to evaluate engagement. Engagement was described as the sustained and meaningful interactions between researchers, participants, and communities. Optimization was described as the application of scientific methods to refine and improve engagement and research processes and outcomes. CONCLUSION: Engagement and optimization strategies have informed research planning, conduct, and dissemination across PE-CGS. These approaches and definitions provide a foundation for developing evidence-based practices to strengthen participant and community involvement in cancer genomics research.

Humans

Development of a Computational Histology Artificial Intelligence-Powered Prognostic Biomarker in Colorectal Cancer in The Cancer Genome Atlas.

BACKGROUND: Risk stratification in colorectal cancer (CRC) plays an important role in treatment decision-making. As such, prognostic biomarkers that can augment risk stratification have clinical value. Quantitative histologic features from routine hematoxylin and eosin (H&E)-stained whole slide images (WSIs) provide a novel avenue for biomarker discovery. In this study, we explored the potential for a computational histology artificial intelligence (CHAI) platform to develop and validate a prognostic biomarker in CRC. METHODS: The Cancer Genome Atlas Colorectal Adenocarcinoma project was utilized for this study, with inclusion of all subjects (stage I-IV) with available digitized H&E specimens. The cohort was split into development and validation cohorts by a stratified random split. The previously developed CHAI platform was applied in the development cohort to construct a continuous risk score from histologic features associated with progression-free interval (PFI) that was dichotomized based on an optimized cutpoint for distinguishing PFI into a high risk CHAI (+) and lower risk CHAI (-). PFI was compared between CHAI (+) and CHAI (-) patients in the validation cohort in multivariable Cox proportional hazards models. Time-dependent area under the curve (tdAUC) and C-indices were also calculated for PFI. RESULTS: A total of 583 participants were included in the study, with 409 assigned to the validation cohort. The CHAI biomarker classified 229 participants (56%) as CHAI (+) and 180 (44%) as CHAI (-) in the validation set. CHAI (+) participants had worse PFI in a multivariable analysis adjusting for available clinicopathologic variables (hazard ratio (HR) = 2.65; 95% confidence interval (CI), 1.63-4.30). TdAUC for the CHAI biomarker was 0.60 (95% CI, 0.53-0.67) at 12 months, 0.62 (0.55-0.69) at 36 months, and 0.67 (0.55-0.79) at 60 months; the C-index was 0.62 (95% CI, 0.58-0.67). CONCLUSIONS: The CHAI platform was used to develop a prognostic digital pathology biomarker in CRC. This demonstrates the feasibility and potential to apply this artificial intelligence-based digital pathology biomarker platform for risk stratification in CRC and supports its further study.

Artificial intelligence

Neotelomeres and telomere-spanning chromosomal arm fusions in cancer genomes revealed by long-read sequencing.

Alterations in the structure and location of telomeres are pivotal in cancer genome evolution. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeats, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. These results provide a framework for the systematic study of telomeric repeats in cancer genomes, which could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Humans

Interpreting Mutation Co-Occurrence in Cancer Genomics Under Biological Context.

Somatic mutation patterns observed in cancer genomes are widely used to generate hypotheses about functional relationships among cancer genes and signaling pathways. However, mutation co-occurrence and mutual exclusivity are assessed at multiple levels, including cohorts, bulk specimens, lesions, regions, clones, and individual cells, although each observational level supports a different scope of inference. In this structured narrative review, we clarify these inferential boundaries and distinguish marginal from conditional association, as well as negative association from complete mutual exclusivity. A hypothetical numerical example of Simpson's reversal illustrates how marginal and conditional associations can differ and why negative association with non-zero overlap should be distinguished from complete mutual exclusivity. We then synthesize evidence from bulk, multi-region, phylogenetic, and single-cell analyses to examine spatial and clonal localization, interclonal cooperation, single-cell error and detection power, and genetic versus non-genetic resistance. We also provide a decision guide for method selection and a staged framework for functional validation. Overall, statistical association, physical localization, and functional interaction are related but distinct inferential targets that require different data, assumptions, and forms of validation.

clonal evolution

LCR-modules: a collection of workflows for cancer genome analysis.

MOTIVATION: The surge of genomic data from advanced sequencing technologies is outpacing current analytical pipelines. We introduce LCR-modules, an open-source suite of bioinformatics tools designed for flexible and automated cancer genome data analysis. LCR-modules enables reproducible analysis of diverse cancer genomics data at scale. The suite comprises 49 Snakemake-based workflows organized into three levels, facilitating tasks from low-level quality control to complex cohort-level analyses. LCR-modules supports various sequencing types and integrates pipelines such as mutation calling, expression quantification, and cohort-level aggregation, ensuring flexibility and reproducibility. LCR-modules represents a significant advancement in genomic data analysis, reducing barriers in reproducibility and scalability and has already been applied to a combination of exomes and genomes from over 10 800 samples. AVAILABILITY: No new data were generated in support of this research. The source code for the LCR-modules is openly available at https://github.com/LCR-BCCRC/lcr-modules.

Software

Neotelomeres and Telomere-Spanning Chromosomal Arm Fusions in Cancer Genomes Revealed by Long-Read Sequencing.

Alterations in the structure and location of telomeres are key events in cancer genome evolution. However, previous genomic approaches, unable to span long telomeric repeat arrays, could not characterize the nature of these alterations. Here, we applied both long-read and short-read genome sequencing to assess telomere repeat-containing structures in cancers and cancer cell lines. Using long-read genome sequences that span telomeric repeat arrays, we defined four types of telomere repeat variations in cancer cells: neotelomeres where telomere addition heals chromosome breaks, chromosomal arm fusions spanning telomere repeats, fusions of neotelomeres, and peri-centromeric fusions with adjoined telomere and centromere repeats. Analysis of lung adenocarcinoma genome sequences identified somatic neotelomere and telomere-spanning fusion alterations. These results provide a framework for systematic study of telomeric repeat arrays in cancer genomes, that could serve as a model for understanding the somatic evolution of other repetitive genomic elements.

Telomere

Mutual Information-based Prognostic Biomarker Discovery in Cancer Genomics: Conceptual Framework and Representative Applications of MI-POG.

Mutual information (MI)-based approaches have increasingly been applied to cancer genomics; however, their use for genome-wide prognostic biomarker discovery remains relatively underexplored. The present article summarizes the conceptual workflow of Mutual Information-based Prognostic Omics Gene (MI-POG) based on previously published applications in breast cancer, lower-grade glioma, and other cancer datasets. The framework consists of clinical endpoint discretization, genome-wide MI-based screening, candidate ranking, and downstream validation using conventional survival-analysis approaches. Previous MI-POG applications identified solute carrier family 20 member 1 (SLC20A1) as a prognostic biomarker in hormone receptor-positive breast cancer. Elevated SLC20A1 expression was associated with unfavorable survival outcomes and was independently validated in the Molecular Taxonomy of Breast Cancer International Consortium (METABRIC) cohort. Methodological analyses demonstrated how survival endpoints can be integrated into an information-theoretic framework through fixed-time outcome discretization, enabling model-independent assessment of molecular-clinical dependencies. Applications across multiple cancer datasets suggested the potential applicability of the framework across biologically distinct tumor types, although further validation will be required to establish its robustness and generalizability. In conclusion, MI-POG can be formalized as an information-theoretic framework for genome-wide identification of prognostic biomarkers by quantifying molecular-clinical dependencies using mutual information. Representative applications from previously published studies suggest that MI-POG may complement conventional survival-analysis approaches and provide a useful strategy for biomarker discovery, although additional benchmarking and prospective validation will be required.

Humans

Severus detects somatic structural variation and complex rearrangements in cancer genomes using long-read sequencing.

For the detection of somatic structural variation (SV) in cancer genomes, long-read sequencing is advantageous over short-read sequencing with respect to mappability and variant phasing. However, most current long-read SV detection methods are not developed for the analysis of tumor genomes characterized by complex rearrangements and heterogeneity. Here, we present Severus, a breakpoint graph-based algorithm for somatic SV calling from long-read cancer sequencing. Severus works with matching normal samples, supports unbalanced cancer karyotypes, can characterize complex multibreak SV patterns and produces haplotype-specific calls. On a comprehensive multitechnology cell line panel, Severus consistently outperforms other long-read and short-read methods in terms of SV detection F1 score (harmonic mean of the precision and recall). We also illustrate that compared to long-read methods, short-read sequencing systematically misses certain classes of somatic SVs, such as insertions or clustered rearrangements. We apply Severus to several clinical cases of pediatric leukemia/lymphoma, revealing clinically relevant cryptic rearrangements missed by standard genomic panels.

Humans

A multicenter survey on BRAF screening for the implementation of perioperative cancer genomic medicine for resectable colorectal oligometastases.

BACKGROUND: Genomic screening is an essential, but potentially time-consuming procedure, especially in neoadjuvant settings. We evaluated the preoperative screening of the BRAF V600E mutation for recruitment to a clinical trial among patients with resectable colorectal oligometastases (CRM). METHODS: In April 2022, an investigator-initiated trial was launched to investigate the efficacy and safety of perioperative use of the BEACON triplet regimen for BRAF V600E mutant resectable CRM. BRAF screening was retrospectively conducted in patients with resected colorectal liver metastases in 2019 for planning the trial and prospectively conducted in preoperative patients with resectable CRM from January 2022 to June 2025 for patient recruitment to the trial. RESULTS: BRAF V600E mutation was detected in 12 (3.2%) of 379 postoperative patients retrospectively and in 36 (1.7%) of 2140 preoperative patients prospectively, with 1840 patients (86.0%) carrying the wild-type and 264 patients (12.3%) classified as untested. The detection rate of the BRAF V600E mutation was significantly lower when the screening was performed prospectively in preoperative patients (P&#x2009;<&#x2009;0.001). The untested rates varied across metastatic organs, with 10.3% in the liver, 18.1% in the lungs, 12.0% in the lymph nodes, 16.7% in the peritoneum, and 7.8% in other organs. The untested rates decreased consistently across semiannual comparisons: 28.5% in the first evaluation, followed by 15.0%, 12.0%, 8.4%, 9.1%, 7.3%, and 7.1% (P&#x2009;<&#x2009;0.01 when compared with the first period). CONCLUSION: Raising physician awareness, as reflected by the untested rate, is a crucial factor in conducting clinical trials to implement perioperative cancer genomic medicine.

Humans

SeqQC-former: A sequence-quality fusion framework for QC-aware review prioritization of candidate somatic SNVs in cancer genomics.

The accurate prioritization of candidate somatic single-nucleotide variants (SNVs) remains a challenge due to the substantial variability in sequencing quality across genomic loci. SeqQC-Former is a sequence-quality fusion framework that integrates the local nucleotide context with read-level quality-control (QC) covariates derived from matched tumor-normal sequencing data. This integration generates QC-aware prioritization scores for the downstream review of candidate variants. Unlike conventional variant callers, SeqQC-Former is designed not to infer biological truth but to support post-calling review and prioritization under heterogeneous sequencing conditions. The framework was trained and evaluated on a SEQC2-derived dataset comprising 89,447 candidate loci, including 1378 positive and 88,069 negative loci. In chromosome-held-out validation, which aims to reduce potential genomic-position leakage, SeqQC-Former demonstrated strong discrimination (AUROC = 0.9479; AUPRC = 0.9448), indicating good generalization to previously unseen chromosomes. Given that the SEQC2-derived labels contain QC-associated information; these results should be interpreted as an evaluation of QC-aware prioritization capability rather than an independent validation of biological variant correctness. Ablation analyses revealed that structured QC covariates provided the dominant predictive signal under the current SEQC2-derived labeling regime. SeqQC-Former achieved a significantly higher AUROC than classical machine-learning baselines, as determined by DeLong's test (p&#x202f;<&#x202f;0.01). Application to 53,164 glioblastoma variants demonstrated that external predictions were sensitive to QC scaling and threshold selection, underscoring that model outputs should be interpreted as QC-dependent prioritization scores rather than calibrated probabilities or definitive biological classifications. Overall, SeqQC-Former offers a reproducible post-calling QC-aware prioritization framework for large-scale somatic SNV review and underscores the importance of explicitly modeling sequencing-quality information when interpreting structured cancer genomics datasets.

Humans

Amplification-Driven S100A11 Overexpression in Hepatocellular Carcinoma Is Associated with Metabolic Reprogramming, ECM Remodelling, and Immune Evasion: A Pan-Cancer Genomic Study.

BACKGROUND: S100A11, a calcium-binding S100 family protein, is increasingly implicated in carcinogenesis, yet its molecular regulation and clinical relevance across cancers remain unclear. Hepatocellular carcinoma (HCC) carries a dismal prognosis, in part due to a lack of reliable biomarkers for risk stratification of established disease. METHODS: We conducted a pan-cancer analysis of S100A11 genomic alterations across 31 studies (10,767 samples) obtained from TCGA, encompassing copy number alterations, somatic mutations, and DNA methylation. HCC-specific analyses evaluated S100A11 expression, its potential as a diagnostic/prognostic marker, co-expression networks, and pathway enrichment using TCGA-LIHC data, with univariate and multivariate Cox regression to assess survival associations. RESULTS: S100A11 alterations were predominantly driven by copy number amplification, with the highest frequencies in hepatobiliary cancers, lung and breast cancers. Copy number amplification showed a consistent inverse relationship with promoter methylation, indicating amplification-driven transcriptional activation. In HCC, S100A11 was markedly overexpressed compared with normal liver tissue, with strong diagnostic discriminatory capacity. High S100A11 expression was significantly associated with inferior overall survival (log-rank p = 0.032; HR = 1.46, 95% CI 1.03-2.06) and remained an independent predictor of overall survival after adjustment for age, sex, and AJCC pathologic stage (HR = 1.27, 95% CI 1.01-1.60, p = 0.038). Co-expression and pathway analyses demonstrated an association between S100A11 and metabolic reprogramming, extracellular matrix remodelling, and immune dysregulation. CONCLUSIONS: These findings identify S100A11 as a candidate diagnostic and prognostic biomarker in HCC whose overexpression is associated with metabolic reprogramming, ECM remodelling, and immune dysregulation, warranting experimental validation of a mechanistic role.

ECM

Protein arginine methyltransferase (PRMT8) in cancer: Genomic alterations, subcellular dynamics, and clinical implications.

PRMT8 encodes a protein arginine methyltransferase, which is primarily expressed in the brain and nervous system. Several studies have reported its alterations, which have been implicated in various cancers. However, the existing information remains unsystematic and fragmented due to inconsistency in methodology. This review aims to explore PRMT8 gene alterations in humans, their effects on cellular function and physiology, and their clinical implications. We conducted a narrative literature review covering all publications on PRMT8 alterations across different cancer types, their effect on tumour cell characteristics, and their impact on patient prognosis. Reported PRMT8 alterations include mutations, copy number amplifications, and single-nucleotide polymorphisms, which lead to overexpression or downregulation of PRMT8 protein in tumour cells. PRMT8 alterations compromise the efficacy of both chemotherapy and immune checkpoint inhibitor treatment. These alterations enable tumour cells to maintain pluripotency via activation of the PI3K/AKT/SOX2 signalling pathway, thereby promoting cellular proliferation, invasion, and colony formation. Clinically, these PRMT8 alterations drive disease progression and therapy resistance, resulting in poor prognosis and reduced patient survival. These findings underscore the need to incorporate PRMT8 alterations assessment in clinical practice to guide therapeutic decision-making and improve treatment outcomes in affected patient populations.

Humans

Aligning awareness, systems and policy to increase equitable access to genomically driven cancer care.

Genomic testing has the potential to transform cancer care across the patient pathway. However, its benefits remain unevenly realised across populations and health systems. Precision oncology is characterised by a strong promissory discourse, with expectations of improved outcomes and cost-effectiveness, yet real-world implementation remains variable and context dependent. This review examines how patient and public awareness interacts with, and is constrained by, structural, organisational and political-economic factors that shape equitable access to genomically driven cancer care across five themes: (1) the power of patient and public advocacy; (2) learning from the patient perspective; (3) culturally responsive communication; (4) structural and personal barriers and facilitators and (5) political economy of health. Examples are mapped across global regions to highlight how health system, structural and societal factors continue to limit the universal realisation of genomic medicine's benefits. We present four recommendations to strengthen the translation of awareness into equitable access and clinical impact: (1) expand and adequately power genomic studies in underserved populations; (2) improve risk communication and decision-making across the cancer pathway; (3) equitable validation and interpretation of emerging genomic technologies and (4) generate real-world evidence on access, uptake and outcomes of genome-matched therapies. Embedding awareness, trust, access and equity in future initiatives is imperative to realise the promise of precision medicine for all patients.

Biomarkers

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

Hierarchical modeling of tumor subtypes in cell lines using large-scale genomic datasets.

Cancer cell lines (CLs) are widely used to study tumor biology and drug response, yet their translational relevance is often limited by inaccurate subtype annotations. Existing CL-tumor matching approaches are frequently constrained by flat classification schemes, weak subtype definitions, and the exclusion of normal tissue references, leading to potential confounding of tumor-specific and tissue-of-origin signals. To address these limitations, a hierarchical classification (HC) framework is presented in which CLs are aligned with patient tumors across biological resolutions, from organ to molecular subtype. Gene expression profiles from 802 CLs, 5,612 tumors from The Cancer Genome Atlas (TCGA) , and 8,939 non-cancerous tissues were integrated to separate oncogenic signals from tissue-specific signals. Node-specific features were selected using maximum relevance minimum redundancy, and balanced accuracies of 89% in cross-validation and 75%, and 80% on external datasets were achieved. Through the framework, 43 CLs were reassigned, and clinically relevant underrepresented subtypes were identified.

cancer cell lines

Haplotype-resolved reconstruction and functional interrogation of cancer karyotypes.

Complex karyotype changes are widespread in cancer genomes. A major gap in cancer genome characterization is the resolution of rearranged chromosomes with chromosome-length continuity. Here, we describe a two-tiered approach to determine the segmental composition of rearranged chromosomes with haplotype resolution. First, we present refLinker, a bioinformatic method for robust determination of chromosomal haplotypes using cancer Hi-C data. By contrast with existing methods, refLinker is insensitive to the presence of large-scale DNA deletions, duplications, and high-level amplification in cancer genomes. Second, we demonstrate a computational strategy to determine the segmental structure of rearranged chromosomes using haplotype-specific Hi-C contacts. We apply these methods to breast cancer genomes and provide direct evidence for long-range transcriptional changes associated with rearrangements of the inactive X chromosome. Together, these results highlight refLinker's broad utility for studying the functional consequences of chromosomal rearrangements.

Humans

Identifying multigenic modules under selection in the tumor genome.

MOTIVATION: Genomic alterations in cancer arise from selective pressures acting on hallmark molecular modules, layered over a background of random mutagenic events. Methods to detect selection at the level of modules, as opposed to genes or nucleotides, are relatively underdeveloped. RESULTS: Here we present CanSRMaPP (Cancer Selection Recovery by Maximum Posterior Probability), a Bayesian model of the cancer genome that infers mutational selection on single genes and multi-genic modules while simultaneously modeling background events. Applying CanSRMaPP to lung adenocarcinoma genomes, we identify positive selection on 63 modules, yielding a model that parsimoniously explains the observed pattern of genetic alterations observed in new cancer cohorts. We further show that CanSRMaPP is adaptable to more tumor types and to alternative module definitions. We show that these modules serve as an effective scaffold for translating the cancer genome to molecular states, with prediction of cancer biomarker status as demonstration. AVAILABILITY: CanSRMaPP is freely available on GitHub. SUPPLEMENTARY INFORMATION: Supplementary Figs. S1-5, Supplementary Tables S1-5, and Supplementary Notes 1 and 2 are available at Bioinformatics online.

Journal Article