PubMed HealthSearch

SEARCH · PubMed Health

Results for “transcriptomic platform”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals

Transcriptomic Profiling of Canine Testicular Leydig Cell Tumors Uncovers Key Upregulated Gene Pathways.

Total RNA was isolated from sections of healthy testes and Leydig cell tumors of mixed-breed dogs using TMA Master II device. The RNA-seq libraries were sequenced on the Illumina platform. Following differential expression analysis, Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Gene Set Enrichment Analysis (GSEA) were applied with quality control obtained using FastQC and Trimmomatic. This analysis revealed 1500 transcripts, including 928 upregulated and 168 downregulated genes. The results demonstrated that a significant proportion of these differentially expressed genes are directly involved in the control of sex steroid production (CYP11A1, STAR, and 3β-HSD3B1) or tube formation, angiogenesis, and extracellular matrix remodeling in interstitial cells (ESM1, FGG, and VEGFA). Moreover, we identified the upregulation of transcripts responsible for neurotransmitter or neuroendocrine signaling (SLC6A4, GRIN2C, GABRB3) and cholesterol metabolism and its regulation (GPX3, MSMO1, DHCR24). These genes were strongly associated with the phosphatidylinositol-3-kinase (PI3K)-Protein Kinase B (Akt) cascade and extracellular matrix interactions, features shared with various malignancies. Alterations in estrogen and relaxin signaling appear to be distinctive, understudied mechanisms specific to canine Leydig cell tumors. Concurrently, downregulated genes (e.g., DMRTC2, SEMA3C, ALOX12) were linked with cell differentiation, signaling and immunoregulatory pathway suppression involved in tumorigenesis. A complex transcriptomic profile of canine Leydig cell tumors was developed, revealing a conserved oncogenic core shared in some aspects with human malignancies alongside unique species-specific alterations. Findings seem to be useful for identifying novel diagnostic biomarkers and targeted therapies in veterinary oncology, establishing canine reproductive tissues as a valuable comparative biomedical model for research in human.

Leydig cell tumor

MitoScribe single-cell molecular recorder logs graded signaling dynamics into mitochondrial DNA.

Genetically encoded DNA recorders convert transient biological events into stable genomic mutations, offering a means to reconstruct past cellular states. However, current approaches to log historical events by modifying genomic DNA have limited capacity to record the magnitude of biological signals within individual cells. Here, we introduce MitoScribe, a mitochondrial DNA (mtDNA)-based recording platform that uses mtDNA base editors (DdCBEs) to write graded biological signals into mtDNA as neutral, single-nucleotide substitutions at a defined site. Taking advantage of the hundreds to thousands of mitochondrial genome copies per cell, we demonstrate MitoScribe enables reproducible, highly sensitive, non-destructive, durable, and high-throughput measurements of molecular signals, including hypoxia, NF-κB activity, BMP and Wnt signaling. We show multiple modes of operation, including multiplexed recordings of two independent signals, and coincidence detection of temporally overlapping signals. Coupling MitoScribe with single-cell RNA sequencing and mitochondrial transcript enrichment, we further reconstruct signaling dynamics at the single-cell transcriptome level. Applying this approach during the directed differentiation of human induced pluripotent stem cells (iPSCs) toward mesoderm, we show that early heterogeneity in response to a differentiation cue predicts the later cell state. Together, MitoScribe provides a scalable platform for high-resolution molecular recording in complex cellular contexts.

Journal Article

A Practical Workflow for Spatial Transcriptomics Data Analysis: From Data Acquisition to Advanced Analyses.

Spatial transcriptomics (ST) profiles genome-wide gene expression while preserving the two-dimensional spatial context of mRNA molecules within tissue sections, enabling studies of tissue architecture and microenvironment-associated biology. However, ST analysis remains challenging because data import, quality control, integration, deconvolution, spatial statistics, and visualization often require multiple software environments and reproducible parameter choices. This protocol presents a practical computational workflow for public ST datasets in R, beginning with data acquisition and software setup and proceeding through Seurat-based data loading, quality control, normalization, multi-sample integration, clustering, and spatially variable gene analysis. The workflow then applies complementary deconvolution strategies, including reference-guided SPOTlight analysis and unsupervised STdeconvolve topic modeling, followed by Giotto-based spatial cell-cell communication analysis and interactive region-of-interest (ROI) selection using a custom Python Dash application. By emphasizing script-based execution, explicit parameter rationales, expected outputs, and troubleshooting checkpoints, the protocol provides an adaptable framework for standard array-based ST datasets and related platforms after dataset- and platform-specific parameter evaluation.

Spatial Transcriptomics

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article

Bridging Organ-on-a-Chip and Omics: A Multi-Dimensional Frontier in Biomedical Research.

Organ-on-a-Chip (OOC) technology offers a powerful platform for replicating human tissue-specific microenvironments, thereby narrowing the translational gap between conventional biomedical models and actual human physiology. Concurrently, omics technologies deliver comprehensive molecular-level insights into biological systems. This review highlights the transformative potential of integrating OOC platforms with high-throughput omics methodologies. We systematically examine the classification, structural configurations, and engineering principles underlying OOC systems, alongside the defining attributes of key omics domains-genomics, transcriptomics, proteomics, and metabolomics. The convergence of dynamic OOC models with advanced omics technologies enables high-resolution, multi-dimensional analyses across numerous biomedical applications, including drug metabolism, disease mechanisms, environmental toxicity assessments, and host-microbiome interactions. This interdisciplinary integration is driving a paradigm shift in precision and translational medicine. However, several challenges remain to be addressed, such as the development of whole-organ mimetics, adaptation of sample collection techniques, and real-time artificial intelligence-based integration of biosensor data with multi-omics datasets. Addressing these hurdles will be vital for unlocking the full potential of this technological synergy in biomedical science.

Multiomics

Establishment and Characterization of Patient-Derived Xenograft Organoids for Personalized Treatment of Castration-Resistant Prostate Cancer.

BACKGROUND: Basic research on castration-resistant prostate cancer (CRPC) is limited by the lack of clinically relevant models. This study aimed to establish patient-derived xenografts (PDX) and PDX-derived organoids from clinical CRPC specimens to develop a bidirectional experimental platform for in vivo xenografts and ex vivo organoids. METHODS: We established a new PDX library (KUCaP PDX series) using CRPC clinical specimens and derived prostate cancer organoids. Comprehensive biological characterization of clinical specimens, PDXs, PDX-derived organoids, and organoid-derived xenografts (ODXs) was performed to confirm the preservation of the original tumor features. Using our PDX library, we conducted genetic engineering and drug testing to explore novel therapeutic approaches. RESULTS: PDX-derived organoids were successfully established from all eight KUCaP PDX lines (100%). Four of the eight lines (50%) were maintained during the long-term culture experiments for over ten passages. Key features observed in the original clinical specimens, including genetic alterations and castration responsiveness, were maintained across the PDX, PDX-derived organoid, and ODX models. RNA sequencing revealed that transcriptomic profiles were consistently maintained across clinical specimens, PDXs, PDX-derived organoids, and ODXs. One PDX and organoid (KUCaP19) which exhibited a high homologous recombination deficiency (HRD) score, without any pathogenic homologous recombination repair (HRR) gene alterations, showed sensitivity to a poly ADP-ribose polymerase (PARP) inhibitor. In contrast, KUCaP12, which had no HRR alterations and a low HRD score, did not respond to PARP inhibition. CONCLUSIONS: We developed the KUCaP library as a novel experimental platform for CRPC research by integrating clinical specimens with PDX, organoid, and ODX models along with their genomic and transcriptomic data. These models largely retained the genetic profiles and responses to castration observed in the original tumors. The bidirectional use of personalized PDX and organoids will facilitate the elucidation of the molecular mechanisms of CRPC.

Male

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans

Paired analysis of primary adenoid cystic carcinoma and derived cell lines reveals a mesenchymal and stem-like shift associated with therapy resistance.

Adenoid cystic carcinoma (ACC) is a salivary gland malignancy characterized by slow but persistent growth, frequent local recurrence, and late metastatic progression. Patients with unresectable, recurrent, or metastatic disease have limited therapeutic options. Efforts to identify effective therapeutic targets have been hindered by the limited availability of well-characterized ACC models. In this study, we established 11 ACC cell lines and performed RNA sequencing of nine cell lines and their matched primary tumors to evaluate the preservation and evolution of molecular and lineage-associated characteristics during cell line establishment. Comparative transcriptomic analysis revealed reduced epithelial and luminal differentiation programs in the cell lines, accompanied by enrichment of myoepithelial, EMT-, and cancer stem cell-associated transcriptional programs. Digital deconvolution and single-sample gene set enrichment analysis supported enrichment of hybrid EMT/stem-like states during in vitro propagation, while comparison with publicly available primary-recurrent ACC data demonstrated partial preservation of recurrence-associated plasticity and invasion programs. Protein-level validation of representative epithelial, myoepithelial, EMT, and stemness markers supported the major transcriptomic changes. In addition, a cell line with a higher stemness signature showed reduced sensitivity to cisplatin. Together, these findings indicate that ACC cell line establishment is associated with transcriptional reprogramming and enrichment of plastic, EMT/stem-like states while retaining selected ACC lineage characteristics. These models provide experimentally tractable platforms for investigating ACC progression, therapeutic response, and mechanisms of treatment resistance.

Adenoid cystic carcinoma

Trustworthy Agentic AI in Bioinformatics: From Workflow Automation to Traceable and Validated Biological Inference.

Agentic artificial intelligence is extending bioinformatics beyond conversational assistance by enabling systems to select tools, execute code, revise analytical plans, and interpret biological data. These capabilities may accelerate research, but they also redistribute decisions that determine whether biological conclusions are valid. We conducted a targeted, structured PubMed search in July 2026 and identified 11 peer-reviewed agentic bioinformatics systems for descriptive review based on predefined eligibility criteria for analytical decision-making, tool or code execution, iterative evaluation, or coordinated agent activity. The evidence base covered single-cell transcriptomics, microbial genomics, cancer genomics, and omics applications, together with methodological literature on reproducibility and biological validation. We examined how current systems report delegated authority, provenance, validation, evidence, abstention, and human oversight. Existing platforms implement safeguards such as sandboxed execution, restricted commands, interaction logs, evidence identifiers, automated checks, critic agents, quality scores, and expert assessment. However, published reports rarely provide a connected account linking the original biological question to samples, reference resources, analytical decisions, computational actions, statistical results, supporting evidence, validation outcomes, and final claims. We distinguish inherited bioinformatics errors, errors amplified through autonomous action, and emergent failures arising from memory, retrieval, tool interaction, or agent coordination. We further propose a multidimensional decision-rights profile, consequence-sensitive validation gates, and a claim-to-evidence provenance architecture organized through the Traceable History of Research Evidence, Agent Actions, and Decisions in Bioinformatics (THREAD-Bio) framework. Illustrative cases show that technically successful execution may still support misleading inference. Trustworthy agentic bioinformatics therefore requires claims to remain reconstructible, challengeable, validated, and proportionate to the evidence.

accountable autonomy

Meta-Merging the Transcriptomes of Gastric Tumors Redefines the Connections among Molecular and Clinical Subtypes.

INTRODUCTION: The availability of a large number of cancer expression profiles presents an excellent opportunity to re-investigate various biological and clinical questions. While several expression profiles have been established for different cancers, merging them may provide a more powerful platform for extensively extrapolating molecular and clinical features across multiple cohorts. MATERIALS AND METHODS: In this study, five gastric tumor expression profiles from the Gene Expression Omnibus [GEO] and one in-house cohort comprising a total of 1,060 samples were merged. The batch effect was removed using non-parametric ComBat analysis, and the seamless merging of datasets was confirmed through various parameters. RESULTS: Extrapolation of ACRG [Asian Cancer Research Group] and TCGA [The Cancer Genome Atlas] molecular subtypes in the merged cohort of 1,060 gastric tumors revealed nine distinct clusters. Notably, the following patterns were observed: [i] mutual exclusivity between Epithelial to Mesenchymal Transition [EMT] and Microsatellite Instability [MSI] subtypes in 90% of tumors; [ii] overlapping occurrence of EMT and MSI subtypes in the remaining tumors; [iii] overlap between MSI and Epstein-Barr Virus [EBV] subtype tumors; [iv] both commonalities and differences between EMT and Genomically Stable [GS] subtypes; and [v] an association between EBV positivity and PI3K mutation. CONCLUSION: The current study demonstrates that compiling a larger expression profile is valuable for revisiting the molecular features and epidemiology associated with molecular subtypes, thereby aiding in the development of novel diagnostics and targeted therapeutics.

Humans

Transcriptome-based epigenetic screening identifies DNA hypermethylation signatures as prognostic biomarkers in oral squamous cell carcinoma.

Promoter DNA hypermethylation is a key epigenetic mechanism of gene silencing in cancer, yet the DNA hypermethylome of oral squamous cell carcinoma (OSCC) and its prognostic relevance remain poorly characterized. Here, we systematically identified and validated novel hypermethylated genes with prognostic significance in OSCC using a genome-wide discovery and multi-platform validation strategy. Candidate genes were first identified by pharmacologic demethylation combined with RNA sequencing across OSCC cell lines, then validated by quantitative RT-PCR, methylation-specific PCR, and bisulfite sequencing in OSCC cell lines, normal oral mucosa, and primary OSCC tumors, with independent confirmation in the TCGA-HNSC dataset. Immunohistochemistry confirmed protein-level silencing, and Kaplan-Meier survival analysis assessed prognostic significance across both cohorts. This pipeline identified five candidate genes, GPX3, ANG, CTGF, GPRC5B, and BAMBI, exhibiting cancer-specific promoter hypermethylation associated with transcriptional and protein silencing in OSCC. Validation in oral cavity tumor samples extracted from the TCGA-HNSC dataset confirmed tumor-specific hypermethylation and revealed significant inverse correlations between methylation and expression for GPX3, GPRC5B, and CTGF. Notably, CTGF hypermethylation was independently associated with poor overall survival in both cohorts (institutional cohort, p=0.03; oral tumor subset from TCGA-HNSC, p=0.01), and a combined ANG+CTGF methylation signature showed superior and reproducible prognostic performance across both platforms. Pathway analysis linked these genes to epithelial-mesenchymal transition and interferon response signaling. This study establishes the first validated DNA methylation biomarker panel for OSCC prognosis, identifying CTGF hypermethylation as a robust prognostic driver with translational potential for clinical risk stratification.

Humans

Molecular differences between young and mature stria vascularis from organotypic explants and transcriptomics.

The stria vascularis (SV) is an essential component of the inner ear that regulates the ionic environment required for hearing. SV degeneration disrupts cochlear homeostasis, leading to irreversible hearing loss, yet a comprehensive understanding of the SV, and consequently therapeutic availability for SV degeneration, is lacking. We developed a whole-tissue explant model from neonatal and mature mice to create a platform for advancing SV research. We validated our model by demonstrating that the proliferative behavior of the SV in vitro mimics SV in vivo. We also provided evidence for pharmacological experimentation by investigating the role of Wnt/β-catenin signaling in SV proliferation. Finally, we performed single-cell RNA sequencing from in vivo neonatal and mature mouse SV and surrounding tissue and revealed key genes and pathways that may play a role in SV proliferation and maintenance. Together, our results contribute new insights into investigating biological solutions for SV-associated hearing loss.

Biochemistry

Spatially guided in vivo single-cell functional genomics of postnatal heart.

Understanding how spatial organization and cell-cell interactions shape gene regulatory programs is central to decoding tissue development and function. The transition at birth, marked by increased circulatory demands and rapid tissue growth, requires precise spatiotemporal coordination of cardiac maturation. In this study, we generated a high-resolution spatial and temporal atlas of the postnatal mouse heart by integrating single-nucleus RNA sequencing with image-based spatial transcriptomics. This framework revealed dynamic cellular interactions, niche-specific signaling and transcriptional programs guiding cardiomyocyte maturation. To functionally test prioritized regulators in vivo and at scale, we developed PIP-seq (probe-based indel-detectable Perturb-seq), a high-throughput platform that detects single guide RNA identity, infers gene editing and profiles transcription from fixed nuclei. Applying PIP-seq to the developing postnatal heart, we identified 21 previously uncharacterized regulators of cardiomyocyte maturation, including genes essential for sarcomere assembly, metabolic reprogramming and electrophysiological transitions. Together, our findings define how microenvironmental signals and intrinsic gene programs cooperate to guide heart maturation and establish a broadly applicable framework for functional genomics in complex tissues.

Animals

LymphGen-Sig: Integrating Genetic and Transcriptional States to Predict Therapeutic Response in Diffuse Large B-Cell Lymphoma.

PURPOSE: Genetic classification may advance precision medicine in diffuse large B-cell lymphoma (DLBCL), but existing tools like LymphGen (LG) are limited by complexity and incomplete classification and do not incorporate nongenetic features that affect disease biology and therapeutic outcomes. To address these limitations, we developed LG-sig (LGsig), a gene expression-based platform that classifies all DLBCLs and harmonizes both genetic and nongenetic dimensions of the disease. METHODS: LGsig was built on the distinct subtype-specific gene expression signature of each LG class using paired genomic and transcriptomic data (National Cancer Institute/British Columbia Cancer Agency; N = 764). Model development was restricted to DLBCLs classified into MYD88L265P and CD79B mutations (MCD), BCL6 translocation and NOTCH2 mutations (BN2), EZH2 mutations and BCL2 translocation (EZB), or SGK1 and TET2 mutations (ST2). Gene features were selected by differential gene expression, with 294 genes being optimal for classification using a nearest shrunken centroid classifier. LGsig classifications were designated as MCDsig, BN2sig, ST2sig, and EZBsig. The final model was applied to RNAseq from archival samples from the POLARIX trial (N = 678) to assess outcomes after polatuzumab vedotin-R-CHP (pola-R-CHP) or rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP) for each LGsig subtype. RESULTS: LGsig accurately identified LG subtypes using transcriptional data alone and extended assignments to all previously LG-unclassified cases. Importantly, LG-unclassified DLBCLs reassigned by LGsig mirrored the transcriptional and clinical features of their corresponding LG counterparts, supporting their reclassification. In addition, LGsig reassigned LG A53 DLBCLs, characterized by aneuploidy and TP53 alterations, into more biologically and therapeutically relevant LGsig clusters. Finally, LGsig improved the performance of LG as a biomarker in the POLARIX study, by identifying distinct DLBCL subtypes exhibiting a survival benefit with pola-R-CHP over R-CHOP in both LG-classified and LG-unclassified cases. CONCLUSION: LGsig expands molecular classification beyond current genetic classifiers in DLBCL by integrating both genetic and transcriptional dimensions of the disease to better inform subtype-specific therapeutic strategies.

Journal Article

Advancing precision tacrolimus therapy: a systems genetics dissection in BXD platform.

BACKGROUND: Tacrolimus is a core immunosuppressant in organ transplantation, but its narrow therapeutic window and significant pharmacokinetic variability hinder precision dosing. Although CYP3A5-guided strategies have established clinical relevance for tacrolimus initial dose adjustment, they do not fully account for the marked interindividual variability in tacrolimus exposure, highlighting the need for complementary models to decode more complex genetic regulation. This study aimed to identify candidate genetic modulators of tacrolimus metabolism and develop an integrated predictive framework for individualized therapy. METHODS: Using 46 BXD recombinant inbred mouse strains, we characterized transcriptomics and machine learning, and validated key genes. We then constructed a clinical model using data from 168 renal transplant recipients. RESULTS: We identified 19 genomic loci associated with tacrolimus pharmacokinetic traits and supported DBP/CYP2A6 as candidate modulators associated with tacrolimus disposition. The clinical prediction model, incorporating these genes and clinical variables, achieved robust AUROC. CONCLUSIONS: These findings support a polygenic contribution to tacrolimus metabolism and provide an experimental and computational framework for identifying candidate modulators relevant to individualized dosing. The BXD mouse platform offers a systems-genetics approach for mechanistic discovery that may inform future translational studies on tacrolimus precision dosing.

Animals

The molecular similarity landscape of preclinical cancer models to patient tumors.

Selecting appropriate preclinical models is fundamental for translational oncology, yet a large-scale, multi-omic quantitative comparison of their similarity to primary human tumors is lacking. To address this, we integrated transcriptomic, proteomic, and genomic profiles from over 10,000 primary tumors from The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC), alongside 4,000 preclinical models. Using a robust computational framework, we revealed a clear hierarchy of transcriptomic and proteomic similarity to patient tumors: with patient-dervied xenografts (PDXs) having greater transcriptomic and proteomic similarity to patient tumors (>) compared with patient-derived organoids (PDOs), which are equal in hierarchy to that of PDX-dervied organoids (PDXOs) > cell lines. We also quantified high molecular conservation (Pearson correlation coefficient = 0.96) across paired in vitro to in vivo platform (organoids to PDX) transitions. Furthermore, genomic analysis demonstrated that whole-exome sequencing (WES) outperforms RNA-seq in detecting DNA variants, and it identified a clonal complexity hierarchy (cell lines > PDXOs > PDXs > PDOs) reflecting the effect of passaging history on intratumor heterogeneity. Ultimately, this study delivers a comprehensive quantitative benchmark, establishing a population-level hierarchy of molecular similarity between preclinical models and primary tumors and providing a data-driven reference for model selection. These findings offer a data-driven framework for selecting models that balance biological representativeness with experimental practicality.

Humans

OmicsPred as a centralised resource for genetic prediction of multi-omic traits.

Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.

Journal Article