PubMed HealthSearch

SEARCH · PubMed Health

Results for “Allelic Imbalance”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Molecular analysis of lung adenocarcinomas from the SAFIR02-Lung cohort reveals new metastasis-associated copy-number alterations including frequent mutant-specific KRAS-allelic imbalance and identifies CDKN2A homozygous deletions as an independent biomarker of poor prognosis.

BACKGROUND: Identifying molecular alterations specific to advanced lung adenocarcinomas could provide insights into tumour progression and dissemination mechanisms. METHOD: We analysed tumour samples, either from locoregional lesions or distant metastases, from patients with advanced lung adenocarcinoma from the SAFIR02-Lung trial by targeted sequencing of 45 cancer genes and comparative genomic hybridisation array and compared them to early tumours samples from The Cancer Genome Atlas. RESULTS: Differences in copy-number alterations frequencies suggest the involvement in tumour progression of LAMB3, TNN/KIAA0040/TNR, KRAS, DAB2, MYC, EPHA3 and VIPR2, and in metastatic dissemination of AREG, ZNF503, PAX8, MMP13, JAM3, and MTURN. Conversely, no meaningful difference was found in pathogenic single-nucleotide variant frequencies, reinforcing the notion that they are early events in tumorigenesis. CDKN2A homozygous deletion was linked to poor clinical outcome in patients with early tumours (overall survival hazard ratio 2.17, 95% CI: 1.43-3.28, corrected p-value = 0.01). Furthermore, we found that KRAS mutant allele specific imbalance, i.e. focal amplification of the mutant allele, is more prevalent in locoregional or distant samples of metastatic patients than in early lesions (8.4%, 13% and 2.8% respectively). This observation was replicated in three public cohorts. Tumours with KRAS mutant allele specific imbalance show specific patterns of co-occurrence and mutual exclusion with alterations in key cancer genes like CDKN2A, TP53, STK11 and NKX2-1, often in a tumour type dependent manner. CONCLUSION: Advanced LUAD tumours exhibit higher copy-number alteration burden, with distinct alterations associated with tumour progression and metastasis. CDKN2A homozygous deletions predict poor prognosis in early disease, while KRAS mutant allele-specific imbalance is enriched in advanced tumours.

Humans

Exploring genomic regions regulating the liver transcriptome and energy homeostasis in pigs.

In pigs, energy homeostasis has an impact on meat quality and health. In a Duroc pig population, 30 quantitative trait locus (QTL) regions associated with fatty acid (FA) composition in adipose tissue, plasma, liver and muscle were previously identified. Mapping of expression quantitative trait locus (eQTL) regions will provide a molecular hypothesis for genotype-phenotype interactions and may allow the identification of shared causal variants, key to increasing our understanding of the genetic regulation of FA composition and energy homeostasis. However, gene expression is impacted by environmental factors, while individual-level allelic imbalance (AI) can be more reliable and can be surveyed via allelic-specific expression (ASE) analysis. Furthermore, treatment of ASE as a quantitative trait allows the identification of allele-specific expression quantitative trait loci (aseQTLs), which are variants whose heterozygosity is linked to the AI of a nearby single-nucleotide polymorphism (SNP), pointing to regulatory elements. In this study, liver was selected as a key metabolic hub with an important role in the regulation of energy homeostasis, and 310 liver RNA sequencing samples were analysed using a combination of (1) eQTL mapping, (2) ASE analysis, and (3) aseQTL mapping methods. A total of 2 188 eQTL regions were identified, mostly cis-eQTL regions (73.17%). ASE analysis reported 1 964 ASE SNPs, associated with 633 genes. Finally, aseQTL mapping reported 64 172 aseQTL, associated with the AI of 31 genes. Colocalisation analysis combined with ASE analysis showed that the expression of FADS1 and FADS2 genes is associated with the polyunsaturated FA composition in several tissues, where microRNA regulation may be present. Finally, in the DGAT2 gene, annotated in a QTL region associated with multiple FAs in adipose tissue, ASE revealed allelic imbalance in the 3' untranslated region (UTR) of this gene. Allelic differential expression can be caused by a 13-bp insertion affecting messenger RNA stability, previously described, exemplifying how allelic imbalance is caused by a post-transcriptional regulatory mechanism undetectable by eQTL mapping. Furthermore, aseQTLs were associated with this gene, linked to a previously identified copy number variant not yet associated with DGAT2 expression. These results demonstrate how ASE analysis and aseQTL mapping can complement eQTL mapping, as they resolved a complex region affected by allelic heterogeneity, a main confounding effect of QTL mapping. In conclusion, the combination of eQTL mapping, ASE analysis and aseQTL mapping allowed the characterisation of the regulation of liver gene expression, improving our understanding of the genetic determinism of energy homeostasis.

Allele-specific expression

Polysomal Profiling Coupled to Allele-Specific Proteomics Reveals an EIF4H TranSNP Allele Possessing Higher mRNA Translation Potential.

To search for genetic sources of allele-specific mRNA translation, we leveraged heterozygous polymorphisms and variants present in the exome of HCT116 colorectal adenocarcinoma-derived cells, computing allelic fractions from both total and polysome-associated RNA from RNA-Seq data. Allelic imbalance in polysomal RNA led us to nominate 52 coding variants associated with allele-specific mRNA translation, of which 16 are nonsynonymous. To validate instances of allele-specific translation, a proteomics workflow was developed that combines label-free shotgun analysis, high-pH reversed-phase peptide fractionation, and targeted parallel reaction monitoring using isotope-labeled peptide standards. Using this approach, we provide proof-of-concept validation of the heterozygous G>A, R183H missense single-nucleotide variant rs1554710467 in the eukaryotic initiation factor 4H (EIF4H) gene. The variant is present in two EIF4H alternatively spliced variants, which showed equivalent translation efficiency in HCT116 cells but differ in abundance. The alternative peptide containing H183 was significantly more abundant than the corresponding reference peptide containing R183, consistent with the over-representation of the alternative allele in polysomal RNA in HCT116 cells. A dual-fluorescence ribosome-stalling assay confirmed the enhanced translation potential of the variant allele. The two EIF4H allelic proteins exhibited similar stability and subpolysomal localization. This study demonstrates the feasibility of using allele-specific proteomics at the endogenous protein levels by exploiting heterozygous coding variants. Overall, our approach extends the toolbox available to investigate allele-specific differences in mRNA translation potential, a relatively underexplored layer of gene expression regulation that could reveal interindividual differences in disease-relevant phenotypes.

Humans

Bayesian estimation of allele-specific expression in the presence of phasing uncertainty.

MOTIVATION: Allele-specific expression (ASE) analyses aim to detect imbalanced expression of maternal versus paternal copies of an autosomal gene. Such allelic imbalance can result from a variety of cis-acting causes, including disruptive mutations within one copy of a gene that impact the stability of transcripts, as well as regulatory variants outside the gene that impact transcription initiation. Current methods for ASE estimation suffer from a number of shortcomings, such as relying on only one variant within a gene, assuming perfect phasing information across multiple variants within a gene, or failing to account for alignment biases and possible genotyping errors. RESULTS: We developed BEASTIE, a Bayesian hierarchical model designed for precise ASE quantification at the gene level, based on given genotypes and RNA-Seq data. BEASTIE addresses the complexities of allelic mapping bias, genotyping error, and phasing errors by incorporating empirical phasing error rates derived from Genome-in-a-Bottle individual NA12878. BEASTIE surpasses existing methods in accuracy, especially in scenarios with high phasing errors. This improvement is critical for identifying rare genetic variants often obscured by such errors. Through rigorous validation on simulated data and application to real data from the 1000 Genomes Project, we establish the robustness of BEASTIE. These findings underscore the value of BEASTIE in revealing patterns of ASE across gene sets and pathways. AVAILABILITY AND IMPLEMENTATION: The software is freely available from Github (https://github.com/x811zou/BEASTIE); and Zendo (DOI: 10.5281/zenodo.15062124).

Bayes Theorem

Integrated clinicogenomic analysis reveals the evolution and metastatic tropisms of advanced colorectal cancer.

We performed an integrated clinical and genomic analysis of over 7,000 consecutively sequenced colorectal cancer (CRC) samples to comprehensively characterize genetic drivers and metastatic tropisms of CRC. We find that genomic evolutionary changes, such as clonal mutations and oncogenic mutant allelic imbalance, selectively enhance the impact of recurrent oncogenic alterations. We identify the relative timing of organ-specific metastasis, showing sequential metastatic progression in microsatellite stable CRC with brain and adrenal metastases as late events; metastatic sites that cluster together, such as lung, bone, and brain metastases; and genomic events that enhance or decrease risk for each metastatic site, with WNT pathway activation as overall protective while RAS pathway activation increased risk for spread to all metastatic sites. Our data suggest that despite the heterogeneity in CRC, genomic evolution increases the impact of recurrent alterations, and integrating information about tumor primary location and genomics can be used to predict organ-specific metastasis risk.

Humans

Prognostic and predictive value of HRD in early triple negative breast cancer (TNBC).

This review explores the emerging role of homologous recombination deficiency (HRD) as both a prognostic and predictive biomarker in early-stage triple-negative breast cancer (TNBC). HRD arises from the defective repair of DNA double-strand breaks through homologous recombination, resulting in genomic instability and increased sensitivity to DNA-damaging agents such as platinum compounds. The review outlines the biological basis of HRD, including genomic signatures such as loss of heterozygosity, telomeric allelic imbalance, and large-scale state transitions, and highlights its prevalence in TNBC compared with other breast cancer subtypes. Clinical trials have shown that HRD-positive patients often achieve higher pathological complete response rates and improved disease-free survival when treated with chemotherapy. However, conflicting evidence across trials underscores the need for more reliable and standardized methods for HRD assessment. The review also explores the therapeutic potential of poly(ADP-ribose) polymerase inhibitors in TNBC, particularly in BRCA-mutated or HRD-positive tumors. Agents such as olaparib, talazoparib, and niraparib have demonstrated promising efficacy in both neoadjuvant and adjuvant settings with some trials suggesting that selected patients may avoid chemotherapy. Furthermore, HRD-positive tumors are characterized by increased genomic instability and a higher neoantigen burden, promoting immune cell infiltration, particularly of tumor-infiltrating lymphocytes, which may enhance responsiveness to immune checkpoint inhibitors. Overall, current evidence supports the role of HRD as a promising biomarker in TNBC. However, further research is required to refine its clinical utility and to integrate HRD testing into personalized treatment strategies, especially in combination with emerging therapies such as immunotherapy.

Humans

A mouse organoid platform for modeling cerebral cortex development and cis-regulatory evolution in vitro.

Natural selection has shaped the gene regulatory networks that orchestrate cortical development, leading to structural and functional variation across mammals, but the molecular and cellular mechanisms underpinning these changes have only begun to be characterized. Here, we develop a reproducible protocol for cerebral cortex organoid generation from mouse epiblast stem cells (EpiSCs), which recapitulates the timing and cellular differentiation programs of the embryonic cortex. We generated cortical organoids from F1 hybrid EpiSCs derived from crosses between laboratory mice (C57BL/6J) and four wild-derived inbred strains spanning ∼1 M years of evolutionary divergence to comprehensively map cis-acting transcriptional regulatory variation across developing cortical cell types, using single-cell RNA sequencing (scRNA-seq). We identify hundreds of genes that exhibit dynamic allelic imbalances, providing the first insight into the developmental mechanisms underpinning changes in cortical structure and function between subspecies. These experimental methods and cellular resources represent a powerful platform for investigating gene regulation in the developing cerebral cortex.

Organoids

Structural diversity and evolutionary constraints of oxidative phosphorylation.

The oxidative phosphorylation (OxPhos) system is central to metabolism. The more than 90 structural subunits are encoded by different chromosome categories (autosomal, X, and mtDNA). The system is envisioned as an invariant structure between cells and individuals. However, a comprehensive analysis of the 1,000 Genomes Project data reveals unexpected genetic intra-individual variability resulting from the heterozygosity of diploid autosomal genes, while diversity at the population level is generated by variability in mtDNA. We characterized the different levels of structural constriction at evolutionary and population levels for all OxPhos protein residues. To support this analysis, we developed ConScore, a conservation-based predictor of variant impact within OxPhos proteins (area under the receiver operating characteristic curve [ROC-AUC] = 0.97; area under the precision-recall curve [PR-AUC] = 0.94). Notably, for the nuclear-encoded subunits, we found mechanisms limiting individual variability as allelic imbalance or homozygosity bias. Integrating structural, functional, and genetic data, we highlight the significance of each OxPhos protein position, expanding insights into its role in speciation and disease.

Oxidative Phosphorylation

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article

[Pathogenic large duplication in TP53 as a hereditary predisposing factor in breast cancer].

AIM: Germline pathogenic variants of TP53 are associated with Li-Fraumeni syndrome and represent a high risk for hereditary breast and ovarian cancer. We identified a germline, multi-exon heterozygous duplication variant in TP53, NM_000546.6:dup(ex2-5), in a young triple-negative breast cancer patient. We aimed to assign its pathogenicity. METHODS: DNA and RNA (cDNA) level specific amplification tests and sequencings were performed to identify the genomic duplication breakpoint and to detect the presence of defective transcripts. RESULTS: cDNA tests revealed aberrant transcript, which causes a shift in the reading frame. Allelic imbalance was also observed, indicating degradation of the defective RNA product. By locating the breakpoints at the DNA level, we determined that 6975 bp was repeated in tandem and in the same orientation. We also revealed the possible mechanism of the structural rearrangement. CONCLUSIONS: We established that the duplication identified at DNA level manifested in the mRNA and coded for a non-functional protein. Based on these data, we were able to classify this duplication variant as pathogenic, which affects the patient's therapeutic options and justifies genetic screening of family members.

Adult

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles

Serum amylase 2 polymorphism in pigs.

A genetic polymorphism of the pig amylase 2 system is described. It is controlled by two codominant genes, Am 2A and Am 2B. A comparison of some breeds from Byelorussian breeding farms gives the following frequencies for the alleles Am 2A and Am 2B, respectively: 0.35 and 0.865 in Large White (n = 682), 0.257 and 0.743 in Byelorussian Black and White (n = 400), 0.540 and 0.460 in Hampshire (n = 51). While studying the genotype distribution according to the Hardy-Weinberg law, a genetic imbalance in Byelorussian Black and White pigs was established (chi2 = 56.4). Genetic relationship between the Am 1 and Am 2 loci was not found.

Alleles

Imbalance in X-chromosome expression: evidence for a human X-linked gene affecting growth of hemopoietic cells.

In each of six family members who were heterozygous at the X-linked locus for glucose-6-phosphate dehydrogenase, only one or the other of the two alleles at that locus was almost exclusively expressed. The data are consistent with evidence that X-chromosome inactivation is a random process that may be followed by selection for one of the two resulting cell types on the basis of an unknown gene, which is located on the X chromosome and which can affect the rate of proliferation of hemopoietic cells in humans.

Alleles

A translocation X;Y system for detecting meiotic nondisjunction and chromosome breakage in males of Drosophila melanogaster.

A nondisjunction and chromosome breakage screening system devised by Craymer and modified in our laboratory, involves an X;Y translocation with the short arm of the Y (Ys), marked with the wild type allele of yellow, attached to the distal end of an X (break point 11D) carrying the recessive marker y; and the long arm of the Y chromosome (YL), marked with the dominant locus Bar of Stone (BS), attached to the proximal end of the X. A female tester strain carrying normal chromosomes homozygous for the yellow allele is employed in the mating scheme. Following normal disjunction in the male, all zygotes, which in this case receive aneuploid paternal sex-chromosomes and a normal euploid maternal complement, will die as a result of genetic imbalance. Thus all survivors from this corss can be classified as exceptions arising from: (1) nondisjunction in the female; (2) gross deletion of the paternal X;Y chromosome; (3) complete loss of the paternal X;Y chromosome; or (4) primary meiotic nondisjunction in the male. Results indicate the sensitivity of this scheme for the detection of events induced by x-rays and various chemicals. Positive results have been obtained with the known mutagens EMS and x-radiation.

Animals

Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.

MOTIVATION: Huntington's disease (HD) exhibits substantial variability in age of onset and disease progression that is not fully explained by CAG repeat length alone. Part of this residual variation is heritable, implicating additional genetic mechanisms. cis-regulatory variation, genetic variants that alter transcription and splicing of nearby genes, represents one such mechanism that can be quantified through allele-specific expression (ASE) analysis. However, methods for integrating ASE profiles into patient stratification frameworks remain underdeveloped, particularly for rare diseases with small cohorts and sparse data. RESULTS: We adapt a network-based stratification algorithm, originally developed for somatic tumour mutations, to ASE data. By propagating gene-level ASE imbalance profiles through a protein-protein interaction network, we stratified 20 HD patients into three distinct biological patient subgroups. Differential gene expression analysis highlights neuroinflammatory pathways, including microglial activation, immune cell activation, and cytokine regulation, as key sources of inter-patient heterogeneity, while differential ASE analysis implicates proteasomal and ubiquitin-dependent protein catabolic processes, immune activation, and central nervous system development. Intersection of differentially imbalanced and expressed genes identified FAM181B as a candidate gene with potential eQTL-mediated regulation, supported by independent cis-eQTL evidence for rs3780 in the caudate and putamen, the primary HD-affected striatal regions. FAM181B encodes a nuclear protein expressed in neural tissues acting as an interactor of the Hippo pathway TEAD transcription factors, implicating transcriptional regulatory variation as a potential contributor to molecular heterogeneity between patient subgroups. Differences in cortical and striatal neuropathological scores between clusters, even when adjusted for CAG repeat length, provide clinical support for the biological relevance of the identified subgroups. AVAILABILITY: All analysis code, Docker containers, and conda environments are available at https://github.com/macsbio/HD-ASE-NBS.

Huntington Disease

Paralogous evolution of the ITS2 region in Xiphophorus.

Ribosomal ITS2 is widely used in phylogenetic studies, yet its multigene organization and potential paralogy can obscure true species relationships. This proof-of-concept study investigates whether ITS2 sequences derived from long-read genomic data in multiple Xiphophorus species primarily reflect orthologous history or are shaped by ancient and local duplications. Phylogenetic analyses reveal two major, reciprocally mirroring ITS2 clades that represent long-standing paralogous rDNA lineages rather than simple allelic variants. The two paralogons show strong asymmetry in copy retention and loss for the majority of the species analyzed in this study. Exceptionally some other species are confined to one paralogon group and exhibit alternating ITS2 variants consistent with persistent ancestral polymorphism. A striking copy number imbalance in X. variatus, combined with its phylogenetic incongruence relative to the established species tree, is best explained by historical rDNA introgression followed by biased concerted evolution that nearly erased one paralogous copy. Despite incomplete homogenization, heterogeneous evolutionary rates, and occasional long-branch artifacts, the recovered paralog-specific topologies largely recapitulate the accepted Xiphophorus species phylogeny, indicating that ITS2 retains a robust organismal signal while also recording episodes of introgression and differential paralog evolution. These results demonstrate that explicit recognition of ITS2 paralogs can both improve phylogenetic interpretation and open avenues for future sequence-structure-based analyses of rDNA evolution and genus-level systematics in Xiphophorus.

Gene duplication

Genome-Wide Single-Nucleotide Polymorphism (SNP)-based Profiling of Loss of Heterozygosity Reveals Distinct Molecular Subgroup-Specific Patterns in Gastrointestinal Stromal Tumors (GIST).

PURPOSE: Gastrointestinal stromal tumors (GIST) are molecularly heterogeneous neoplasms defined by mutually exclusive driver alterations (KIT, PDGFRA, SDH, BRAF, RAS, and NF1). However, driver mutations alone do not fully explain their biological and clinical variability. Chromosomal imbalances and loss of heterozygosity (LOH) may represent an additional layer of tumor characterization. We developed a single-nucleotide polymorphism (SNP)-based next-generation sequencing panel enabling genome-wide LOH assessment from formalin-fixed paraffin-embedded tissue. MATERIALS AND METHODS: Forty-nine GIST cases molecularly classified using targeted next-generation sequencing (KIT n = 19, PDGFRA n = 9, SDH-deficient n = 8, NF1 n = 7, quadruple wild-type n = 6) were analyzed. LOH was inferred from variant allele frequency patterns across 1826 genome-wide SNPs. RESULTS: Chromosome 14 was the most commonly affected (63%), followed by chromosomes 22 (45%), 15 (41%), 21 (27%), and 13 (20%). Loss of chromosome arm 1p occurred in 43% of tumors. Distinct subgroup-specific patterns emerged: KIT-mutant GIST exhibited the highest degree of genomic instability, whereas both SDH-deficient tumors and PDGFRA-mutant GIST displayed minimal chromosomal instability. NF1-mutant tumors showed recurrent single-arm chromosome 17 LOH. Quadruple wild-type GISTs were heterogeneous, including 1 case with extensive chromosomal instability. CONCLUSIONS: Genome-wide SNP-based LOH profiling reveals distinct, subgroup-specific patterns of chromosomal imbalance in GIST and may serve as a feasible complementary approach to driver mutation analysis for refined molecular characterization and potential future clinical utility.

Humans

Bridging ancestry gaps in genomic risk prediction with tabular foundation models.

MOTIVATION: Models deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. RESULTS: Using large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. AVAILABILITY AND IMPLEMENTATION: All code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187.

Humans