PubMed HealthSearch

SEARCH · PubMed Health

Results for “variational inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Accounting for recombination rate variation improves inference of barrier loci and reveals the role of both natural and sexual selection in an incipient bird radiation.

Examining genomic patterns of differentiation across lineage pairs at different stages of the speciation continuum, in combination with recombination maps, can help disentangle the effects of linked and divergent selection and identify lineage-specific targets of selection that may act as barrier loci during speciation. Here, we apply this framework to genomic data from African and Indian Ocean bird species of the genus Zosterops (Zosteropidae) to identify candidate barrier loci between ecologically, phenotypically, and genetically distinct Reunion gray white-eye (Zosterops borbonicus) parapatric geographic forms. Using analyses that account for recombination rate variation, we show that putative targets of divergent selection are primarily located on the Z chromosome, except in comparisons between geographic forms that differ in their ecologies. Functional annotation revealed that candidate barrier loci between forms with similar environmental niches are associated with genes involved in song formation and immune function, whereas those between forms with different environmental niches are associated with adaptation to altitude, morphology, and song behavior. Our results highlight the combined roles of natural and sexual selection in the evolution of reproductive barriers in this incipient species radiation.

Animals

Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework.

Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA-protein (CITE-seq) and RNA-chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin-a nonhematopoietic tissue with continuous differentiation hierarchies-UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA-protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

Journal Article

Interactions between diverse proteinoids and microspheres in simulation of primordial evolution.

Experiments demonstrating an incorporation of different enzymelike activities into a single preparation of proteinoid microspheres provide a conceptual basis for the primitive lengthening of protometabolic pathways. An enhancement of one enzymelike activity by another proteinoid in the same microsphere has been found. This effect, plus the pathway-lengthening propensity of combinations of microspheres, indicates selective advantages contributing to adaptive protoselection. Data reported in this paper also bring into purview the concept of internally controlled variation. Inferences are derived for the origin of protosexuality in protocells. When allowance is made for a closer relationship to the environment than that needed in contemporary selection, the fundamental mechanistic requirements of protoevolution are regarded as met by the proteinoid microsphere.

Catalysis

Exploiting Omic Data to Advance Predictive Ecotoxicology.

Predicting species-specific chemical sensitivity using in silico approaches has the potential to transform environmental risk assessment, conservation, and biomonitoring, while reducing, and ultimately replacing, animal testing. Genomic and transcriptomic data capture extensive sensitivity-relevant variation, including differences in molecular targets, xenobiotic metabolism, and damage mitigation pathways. Large-scale sequencing initiatives therefore offer an unprecedented opportunity to address ecotoxicology's "too many species" problem. Although existing omic-based predictive tools provide proof of concept, they have so far been applied to a narrow set of relatively straightforward prediction scenarios. To achieve broader applicability, current and future tools must be firmly grounded in the diverse molecular mechanisms underlying differential chemical responses. Here, we critically evaluate the emerging field of predicting species sensitivity using molecular variation inferred from omic data. We analyze the strengths and limitations of current omic-based approaches and identify major sequence and ecotoxicological data gaps, as well as critical bioinformatic challenges. We then review the current knowledge of how molecular biology underlies differential chemical sensitivity, outlining research paths to allow the next generation of sensitivity prediction tools to exploit ever expanding omic data.

Ecotoxicology

NExON-Bayes: a Bayesian approach to network estimation informed by ordinal covariates.

MOTIVATION: In heterogeneous disease settings, accounting for intrinsic sample variability is crucial for obtaining reliable and interpretable omic network estimates. However, most graphical model analyses of biomedical data assume homogeneous conditional dependence structures, potentially leading to misleading conclusions. To address this, we propose a joint Gaussian graphical model that leverages sample-level ordinal covariates (e.g. disease stage) to account for heterogeneity and improve the estimation of partial correlation structures. RESULTS: Our modelling framework, called NExON-Bayes, extends the graphical spike-and-slab framework to account for ordinal covariates, jointly estimating their relevance to the graph structure and leveraging them to improve the accuracy of network estimation. To scale to high-dimensional omic settings, we develop an efficient variational inference algorithm tailored to our model. Through simulations, we demonstrate that our method outperforms the vanilla graphical spike-and-slab (with no covariate information), as well as other state-of-the-art network approaches which exploit covariate information. Applying our method to reverse phase protein array data from patients diagnosed with stage I, II or III breast carcinoma, we estimate the behaviour of proteomic networks as cancer progresses. Our model provides insights not only through inspection of the estimated proteomic networks, but also of the estimated ordinal covariate dependencies of key groups of proteins within those networks, offering a comprehensive understanding of how biological pathways shift across disease stages. AVAILABILITY AND IMPLEMENTATION: A user-friendly R package for NExON-Bayes with tutorials is available on Github at github.com/jf687/NExON, and archived at https://doi.org/10.5281/zenodo.20312938. The source of the dataset used is cited in the relevant section.

Bayes Theorem

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

TreeFlow: Probabilistic Modelling and Automatic Differentiation for Phylogenetics.

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

Bayesian inference

Gene-level complexity explains genome-wide variation in the distribution of fitness effects.

The distribution of fitness effects (DFE)-describing how harmful, neutral, or beneficial new mutations are-is central to understanding how populations evolve. Although the DFE varies across genomes and species, it remains unclear which aspects of genomic organization drive this variation. Here, we inferred gene-level selective constraints across the genomes of Mus musculus castaneus, Drosophila melanogaster and Saccharomyces cerevisiae using a combination of population genetics and machine learning trained on diverse gene features. Many gene features were predictive of selective constraint, with conservation, gene structure, and expression being the most informative. These selective constraints delineated gene classes with distinct DFEs. Genes with higher connectivity and expression-features reflecting how many traits a gene influences-experienced stronger and less dispersed deleterious effects with increasing selective constraint. Between species, the rate of adaptation decreased with increasing organismal complexity, whereas across the genome it did not decrease monotonically with selective constraint, but tended to be higher at intermediate levels. While between-species comparisons of DFE parameters were less consistent with predictions of Fisher's geometric model (FGM) based on organismal complexity, variation in DFE parameters across the genome aligned more closely with FGM when complexity was considered at the gene level. Our results suggest that gene-level complexity, captured by genomic feature proxies, provides a more informative definition of complexity for DFE variation than organism-level labels, and highlight the value of using gene features collectively to link genomic architecture, fitness landscapes, and patterns of molecular evolution.

Animals

Spatial proteomics reveals four-stage molecular evolution in cancer immunotherapy-related gastritis.

BACKGROUND: Immune checkpoint inhibitors (ICIs) have transformed cancer treatment, yet immune-related adverse events (irAEs) including immunotherapy-related gastritis (IRAEG) pose significant clinical challenges-often necessitating treatment interruption that may compromise antitumor efficacy. IRAEG presents with atypical symptoms, lacks specific biomarkers, and shows histopathological overlap with other forms of gastritis, complicating diagnosis and management. Despite increasing clinical recognition, a systematic understanding of spatial molecular alterations across the full disease course remains limited. Here, we used spatial proteomics to map the molecular landscape of IRAEG during disease progression and to define stage-specific patterns of molecular evolution relevant to cancer immunotherapy management. METHODS: We analyzed tissue samples from seven patients, including four non-immunotherapy-related gastritis controls and three cancer patients who developed IRAEG following ICI therapy for solid tumors, sampled longitudinally across four disease stages: baseline (G1), acute severe inflammation (G2), early recovery (G3), and complete recovery (G4). Using laser capture microdissection coupled with data-independent acquisition mass spectrometry, we profiled 177 spatially resolved gastric tissue regions. Multiplex immunohistochemistry and immunofluorescence characterized features of the immune microenvironment, while Gene Ontology, KEGG pathway analysis, Gene Set Variation Analysis, and xCell inference enabled functional, metabolic, and immune profiling. Key immune and NET-related findings were further validated by multiplex immunofluorescence in an independent, expanded cohort of IRAEG and non-immunotherapy-related gastritis samples. RESULTS: IRAEG was characterized by widespread HLA molecule activation and enhanced antigen processing, resembling the immune phenotype observed in organ transplant rejection. The acute G2 stage exhibited excessive neutrophil extracellular trap formation, profound metabolic suppression, and collapse of immune homeostasis-features that may inform early intervention strategies to preserve ICI treatment continuity. During early recovery (G3), inflammatory injury transitioned toward repair, marked by activation of fatty acid metabolism and PPAR signaling. Notably, even at complete clinical recovery (G4), more than 1,000 proteins remained differentially expressed, reflecting sustained enhancement of metabolic and immune functions and establishing a distinct molecular "memory" state with implications for ICI rechallenge decisions. CONCLUSIONS: These findings define four molecularly distinct stages of IRAEG progression and recovery. The stage-specific signatures identified here serve as candidate biomarkers for diagnosis, disease staging, and therapeutic response assessment, and may guide clinical decisions regarding irAE management, treatment modification, and safe ICI rechallenge to support continued antitumor therapy.

Humans

Advancing translational exposomics: bridging genome, exposome and personalized medicine.

Understanding the interplay between genetic predisposition and environmental and lifestyle exposures is essential for advancing precision medicine and public health. The exposome, defined as the sum of all environmental exposures an individual encounters throughout their lifetime, complements genomic data by elucidating how external and internal exposure factors influence health outcomes. This treatise highlights the emerging discipline of translational exposomics that integrates exposomics and genomics, offering a comprehensive approach to decipher the complex relationships between environmental and lifestyle exposures, genetic variability, and disease phenotypes. We highlight cutting-edge methodologies, including multi-omics technologies, exposome-wide association studies (EWAS), physiology-based biokinetic modeling, and advanced bioinformatics approaches. These tools enable precise characterization of both the external and the internal exposome, facilitating the identification of biomarkers, exposure-response relationships, and disease prediction and mechanisms. We also consider the importance of addressing socio-economic, demographic, and gender disparities in environmental health research. We emphasize how exposome data can contextualize genomic variation and enhance causal inference, especially in studies of vulnerable populations and complex diseases. By showcasing concrete examples and proposing integrative platforms for translational exposomics, this work underscores the critical need to bridge genomics and exposomics to enable precision prevention, risk stratification, and public health decision-making. This integrative approach offers a new paradigm for understanding health and disease beyond genetics alone.

Humans

Macrolide-resistant Mycoplasma pneumoniae resurgence in Chinese children in 2023: a longitudinal, cross-sectional, genomic epidemiology study.

BACKGROUND: After a prolonged period of low detection rates, Mycoplasma pneumoniae resurged in China, during September to November, 2023, raising global concern. This study aims to gain a better understanding of the genetic mechanisms underlying the 2023 increase in cases and the evolutionary dynamics of the epidemic populations, which has been previously hampered due to limited genomic data of this pathogen. METHODS: We sequenced 685 M pneumoniae isolates, including 248 isolates from 11 Chinese provinces and municipalities in 2023 and 437 isolates from Beijing (2013-22). By analysing these isolates and 436 publicly global sequences, we reconstructed the pathogen's evolutionary history using time-calibrated phylogenies and effective population size inference. We investigated potential genomic variations contributing to the 2023 resurgence through genome-wide association study and conducted phylogeographic analysis of the 2023 isolates across China. FINDINGS: Two macrolide-resistant epidemic clusters (T1-2-EC1 and T2-2-EC2) were responsible for the 2023 resurgence in China. Both clusters, having acquired the 23S ribosomal RNA A2063G mutation conferring macrolide resistance, emerged in approximately 1997 and 2014, respectively, and subsequently outcompeted their predecessor populations. This coincided with China's large-scale adoption of azithromycin for paediatric community-acquired pneumonia around the early 2000s. Aside from macrolide resistance, T1-2-EC1 independently acquired 17 clade-specific mutations and T2-2-EC2 four clade-specific mutations, which could further explain their increased competitiveness. Whole-genome analysis revealed no resurgence-specific mutations in the 2023 isolates. Phylogeographic analysis showed rapid mixing of T1-2-EC1 isolates between different sampled regions within China. INTERPRETATION: Our study provides evidence that the 2023 resurgence in China is a continuation of the pre-COVID epidemic, rather than emergence of novel variants. The high prevalence of macrolide resistance and rapid intranational spread emphasise the urgent need for enhanced global surveillance of this pathogen. FUNDING: National Key Research and Development Program of China, National Natural Science Foundation of China for Key Programs of China Grants, and Beijing High-Level Public Health Technical Talent Project.

Humans

Identification and masking of artifactual and misleading within-host variants in deep-sequencing SARS-CoV-2 data.

Deep-sequencing data are increasingly used to study within-host viral diversity and to inform evolutionary inference. For SARS-CoV-2, analyses based on intra-host single-nucleotide variants (iSNVs) have been widely applied to quantify within-host diversity and infer transmission dynamics. However, these applications critically depend on the reliable identification of low-frequency variants, which remain vulnerable to systematic and technical artifacts. In this study, we show that recurrent artifactual iSNVs are common in large-scale SARS-CoV-2 sequencing data and can persist even under conservative minor allele frequency thresholds. Using data from the UK's Office for National Statistics COVID-19 Infection Survey, we demonstrate that such artifacts are predominantly sequencing center-specific rather than primer-specific. Each center exhibits a modest, distinct set of recurrent artifactual variants showing little overlap with sites routinely masked at the consensus level. To address this, we developed a systematic, dataset-aware framework that uses recurrence within sequencing datasets to identify small, noise-adapted sets of artifactual iSNVs to mask. Applying this framework reduces spurious sharing of low-frequency variants between samples and qualitatively alters downstream inferences, including estimates of within-host diversity and transmission bottleneck sizes. Although this study focused on SARS-CoV-2, it is likely that recurrent artifactual iSNVs will be problematic for other viruses as mass-sequencing becomes increasingly routine. Together, these findings highlight the importance of explicit, dataset-aware artifact control for robust inference from within-host variation, particularly as genomic studies increasingly seek to exploit sub-consensus diversity in rapidly evolving pathogens.

Humans

SARS-CoV-2 intra-host variation shows evidence of transmission and convergent evolution in a university surveillance cohort.

Monitoring and understanding the transmission and evolution of SARS-CoV-2 remains a significant public health priority. Within-host genetic variation provides insight into viral evolution during infection and may help infer transmission events. In this study, we analysed intra-host variation in SARS-CoV-2 genome sequences from Boston University's testing mandate. Focusing on intra-host single nucleotide variants (iSNVs), we inferred transmission events and assessed the selective forces shaping within-host viral evolution. To minimize false-positive iSNVs resulting from systematic biases, we implemented stringent data filtering and developed a heuristic to exclude contamination-derived artefacts arising from batched sequencing. We find that intra-host variation is limited and infrequently transmitted during acute infections, suggesting that shared iSNVs serve as highly specific but insensitive markers of transmission. We also observed incomplete purifying selection shaping within-host diversity, with the loci most affected changing among variants of concern. Finally, we identified a highly recurrent iSNV (G11083T) which may represent a site of positive selection. Our results highlight that within-host variation provides insight into within-host pathogen evolution, in spite of its limited use in genomic epidemiology.

SARS-CoV-2

Population-scale detection of methylation outliers from long-read genome sequencing.

BACKGROUND: Aberrant DNA methylation can mediate the functional effects of rare genetic variation and contribute to imprinting disorders, repeat expansion diseases, and other pathogenic regulatory mechanisms. Long-read sequencing technologies now enable genome-wide detection of CpG methylation alongside genetic variation from a single assay. However, methods for systematic identification and interpretation of methylation outliers from long-read sequencing data remain limited. METHODS: We developed METAFORA, a computational workflow for detecting methylation outlier regions from PacBio and Oxford Nanopore long-read sequencing data. METAFORA constructs population-level methylation references, segments the genome into correlated CpG blocks, infers technical and biological sources of variation through hidden factor estimation, models uncertainty due to variable depth sequencing, and computes covariate-adjusted methylation outlier scores for individual samples. We applied METAFORA across large long-read sequencing cohorts and integrated methylation outliers with multi-omic data. METAFORA is implemented as a snakemake workflow available at https://github.com/tjense25/METAFORA. RESULTS: METAFORA identified methylation outlier regions associated with rare structural variants, tandem repeat expansions, and imprinting abnormalities. We found outlier regions were enriched for molecular outliers across transcriptomic and chromatin accessibility datasets, supporting their functional relevance in gene regulation. In a representative case, METAFORA identified an imprinting defect affecting the GNAS locus associated with an STX16 deletion. CONCLUSIONS: METAFORA enables scalable detection and interpretation of methylation outliers from long-read sequencing data and provides a framework for integrating epigenetic outliers with genomic and multi-omic analyses. These approaches may improve interpretation of rare regulatory variation and support discovery of clinically relevant epigenetic abnormalities in genomic medicine.

DNA methylation

Comparative Population Genomics of Relictual Caribbean Island Gossypium hirsutum.

Gossypium hirsutum is the world's most important source of cotton fibre, yet the diversity and population structure of its wild forms remain largely unexplored. The complex domestication history of G. hirsutum combined with reciprocal introgression with a second domesticated species, G. barbadense, has generated a wealth of morphological forms and feral derivatives of both species and their interspecies recombinants, which collectively are scattered across a large geographic range in arid regions of the Caribbean basin. Here we assessed genetic diversity within and among populations from two Caribbean islands, Puerto Rico (n = 43, five sites) and Guadeloupe (n = 25, one site), which contain putative wild or introgressed forms. Using whole-genome resequencing data and a phylogenomic framework derived from a broader genomic survey, we parsed individuals into feral derivatives and truly wild forms. Feral cottons display uneven levels of genetic and morphological resemblance to domesticated cottons, with diverse patterns of genetic variation and heterozygosity. These patterns are inferred to reflect a complex history of interspecific and intraspecific gene flow that is spatially highly variable in its effects. Wild cottons in both Caribbean islands appear to be relatively inbred, especially the Guadeloupe samples. Our results highlight the dynamics of population demographics in relictual wild cottons that experienced profound genetic bottlenecks associated with repeated habitat destruction superimposed on a natural ecogeographical distribution comprising widely scattered populations. These results have implications for conservation and utilisation of wild diversity in G. hirsutum.

Genetics, Population

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

Large-scale admixture mapping in the All of Us Research Program improves the characterization of cross-population phenotypic differences.

Admixed individuals have largely been understudied in medical research due to their complex genetic ancestries. However, the consideration of admixture can help identify ancestry-enriched genetic associations, delineating some of the genetic underpinnings of cross-population phenotypic variation. To this end, we performed local ancestry inference within the All of Us Research Program to identify individuals with recent admixture between African (AFR) and European (EUR) populations (N=48,921). We identified evidence of local AFR ancestry enrichment at the HLA locus, suggestive of putative selection since admixture. Furthermore, we performed the largest admixture mapping (ADM) efforts in AFR-EUR Admixed individuals for 22 traits, identifying 71 associations between inferred local AFR ancestries and a trait. Variants from published GWAS could only account for 18 (25%) of the ADM associations, highlighting novel loci where ancestral haplotypes explained some phenotypic variation. Previous studies likely have not identified these loci due to the low availability of high-powered GWAS in populations genetically similar to AFR. One such loci was 9q21.33, associated with 1.4-fold risk of end-stage kidney disease (ESKD) for carriers of inferred local AFR ancestries at the region. This locus contains the gene SLC28A3, which has previously been linked to kidney function but has never been associated with cross-population ESKD prevalence differences. Together, our results expand upon the existing literature on phenotypic differences between populations, highlighting loci where genetic ancestries play a critical role in the genetic architecture of disease.

Journal Article

Species-wide gene editing of a flowering regulator reveals hidden phenotypic variation.

Genes do not act in isolation, and the effects of a specific variant at one locus can often be greatly modified by polymorphic variants at other loci. A good example is FLOWERING LOCUS C (FLC), which has been inferred to explain much of the flowering time variation in Arabidopsis thaliana. We use a set of 62 flc species-wide mutants to document pleiotropic, genotype-dependent effects for FLC on flowering as well as several other traits. Time to flowering was greatly reduced in all mutants, with the remaining variation explained mainly by allelic variation at the FLC target FT. Analysis of FT sequence variation suggested that extremely early combinations of FLC and FT alleles should exist in the wild, which we confirmed by targeted collections. Our study provides a proof of concept on how pan-genetic analysis of hub genes can reveal the true extent of genetic networks in a species.

Gene Editing