PubMed HealthSearch

SEARCH · PubMed Health

Results for “Coupled model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SURROGATE SELECTION OVERSAMPLES EXPANDED T CELL CLONOTYPES.

Surrogate selection is an experimental design that without sequencing any DNA can restrict a sample of cells to those carrying certain genomic mutations. In immunological disease studies, this design may provide a relatively easy approach to enrich a lymphocyte sample with cells relevant to the disease response because the emergence of neutral mutations associates with the proliferation history of clonal subpopulations. A statistical analysis of clonotype sizes provides a structured, quantitative perspective on this useful property of surrogate selection. Our model specification couples within-clonotype birth-death processes with an exchangeable model across clonotypes. Beyond enrichment questions about the surrogate selection design, our framework enables a study of sampling properties of elementary sample diversity statistics; it also points to new statistics that may usefully measure the burden of somatic genomic alterations associated with clonal expansion. We examine statistical properties of immunological samples governed by the coupled model specification, and we illustrate calculations in surrogate selection studies of melanoma and in single-cell genomic studies of T cell repertoires.

Bayes’s rule

An integrated multiscale air quality modelling framework for industrial park pollution: Linking local emissions to regional transport.

Capturing the spatiotemporal distribution of pollutants in industrial parks remains challenging for regional air quality models because of their coarse resolution (3 km), resulting in uncertainties in local emission quantification. To address this, we developed the Integrated Multiscale Air Quality Modelling System for Industry (IAQMS-Industry), coupling the regional Nested Air Quality Prediction Modelling System (NAQPMS) with a city-scale chemical transport model. This framework integrates point-source locations and Gaussian plume dispersion to simulate particulate matter with a diameter smaller than 2.5 micrometres (PM2.5) at 100 m resolution. Applied to the Beijing Yi Zhuang and Tangshan industrial parks and evaluated against observations. The coupled model achieved a normalized mean bias (NMB) ranging from 3.1 % to 6.2 %, improving upon NAQPMS (-16.9 % to -7.7 %). Spatial analysis revealed that coarse regional grids underestimated the PM2.5​ concentrations at industrial sites by smoothing gradients, whereas IAQMS-Industry successfully resolved spatial patterns. Industrial point emissions accounted for 22.9 %-26.4 % of PM2.5 in the coupled model, which was significantly greater than the regional model estimates of 1.6 %-13.7 %. These findings indicate that regional models overestimate pollutant dispersion processes in industrial parks while underestimating local industrial impacts. By explicitly resolving point-source dynamics and linking them to regional transport, IAQMS-Industry provides a robust tool for designing targeted emission controls in industrial cities and balancing local air quality improvements with minimized regional pollution outflow. This study underscores the necessity of multiscale modelling for accurate source apportionment and informed environmental governance in industrial zones.

Air Pollution

Flux-sum coupling analysis of metabolic network models.

Metabolites acting as substrates and regulators of all biochemical reactions play an important role in maintaining the functionality of cellular metabolism. Despite advances in the constraint-based framework for genome-scale metabolic modeling, we lack reliable proxies for metabolite concentrations that can be efficiently determined and that allow us to investigate the relationship between metabolite concentrations in specific metabolic states in the absence of measurements. Here, we introduce a constraint-based approach, the flux-sum coupling analysis (FSCA), which facilitates the study of the interdependencies between metabolite concentrations by determining coupling relationships based on the flux-sum of metabolites. Application of FSCA on metabolic models of Escherichia coli, Saccharomyces cerevisiae, and Arabidopsis thaliana showed that the three coupling relationships are present in all models and pinpointed similarities in coupled metabolite pairs. Using the available concentration measurements of E. coli metabolites, we demonstrated that the coupling relationships identified by FSCA can capture the qualitative associations between metabolite concentrations and that flux-sum is a reliable proxy for metabolite concentration. Therefore, FSCA provides a novel tool for exploring and understanding the intricate interdependencies between the metabolite concentrations, advancing the understanding of metabolic regulation, and improving flux-centered systems biology approaches.

Escherichia coli

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

Coupling of spectroscopy and nitrogen-oxygen isotopes unveils the mechanisms of dissolved organic matter and nitrate pollution in lakes within the agro-pastoral transition zone.

Lakes in arid and semi-arid regions are subjected to severe ecological stress, such as organic pollution, eutrophication, and salinization, due to climate change and human activities. This study investigates Chagannur Lake, a typical arid-region lake that is representative and ecologically sensitive in Northern China's agro-pastoral ecotone, to uncover its pollution characteristics and mechanisms. We employed fluorescence spectroscopy and stable isotope analysis to trace dissolved organic matter (DOM) and nitrate sources. The DOM composition was dominated by microbial metabolic byproducts and protein-like substances, suggesting that microbial processes are key to organic matter transformation. Source apportionment revealed that pollutants primarily originated from livestock and poultry manure (37.6 %), agricultural fertilizers (35.6 %), and soil erosion (24.7 %), with agricultural fertilizers contributing most significantly in the Gogstai River (63.3 %). A structural equation model (SEM) coupling spectral and mass spectrometric data revealed that microbial transformation significantly impairs the lake's self-purification capacity, thereby promoting pollutant accumulation (path coefficient = 0.91,*p < 0.05). Moreover, microbial processes link endogenous and exogenous pollution, a mechanism effectively traced by isotopic and fluorescence indices (path coefficient = 0.55, &#x204e;&#x204e;p < 0.01). These findings enhance the understanding of pollution sources and transformation mechanisms in arid-region lakes and offer foundational theoretical support for policymakers engaged in pollution control strategies.

Lakes

Asymmetric integration of various cancer datasets for identifying risk-associated variants and genes.

MOTIVATION: Cancer genomic research provides an opportunity to identify cancer risk-associated genes, but often suffers from undesirable low statistical power due to a limited sample size. Integrated analysis with different cancers has the potential to enhance statistical power for identifying pan-cancer risk genes. However, substantial heterogeneity across various cancers makes this challenging. RESULTS: Recently, a novel asymmetric integration method was developed that can deal with data heterogeneity and exclude unhelpful datasets from the analysis. We adapted and applied this method to integrate genotype datasets with matched case and control individuals from the Michigan Genomics Initiative, using each cancer as the primary dataset of interest and the other cancers as auxiliary datasets, respectively. Conditional logistic regression models were coupled with the asymmetric integrated framework to handle the matched case-control study design and permutation tests were performed to control for false discovery rates (FDRs). At the same FDR level, the integrated analysis found more potential genetic variants and genes that are associated with the risks of various cancers, showcasing the promise of the proposed approach for integrated analysis of cancer datasets. AVAILABILITY AND IMPLEMENTATION: Our method is available as source code at https://github.com/rxxwang/integrate_cancer.

Journal Article

Multi-scale phylodynamic modelling of rapid punctuated pathogen evolution.

Computational multi-scale pandemic modelling remains a major and timely challenge. Here we identify specific requirements for a new class of models simulating pandemics across three scales: (1) pathogen evolution, often punctuated by the rapid emergence of new variants, (2) human interactions within a heterogeneous population, and (3) public health responses which constrain individual actions to control the disease transmission. We then present a pandemic modelling framework satisfying these requirements and capable of simulating feedback loops between dynamics unfolding at these different scales. The developed framework comprises a stochastic agent-based model of pandemic spread, coupled with a phylodynamic model that incorporates within-host pathogen evolution. It is validated with a case study, modelling the punctuated evolution of SARS-CoV-2, based on global and contemporary genomic surveillance data, which captures a large heterogeneous population. We demonstrate that the model replicates the essential features of the COVID-19 pandemic and virus evolution, while retaining computational tractability and scalability.

SARS-CoV-2

Steroid hormone biosynthesis and dietary related metabolites associated with excessive daytime sleepiness.

BACKGROUND: Excessive daytime sleepiness (EDS) is a complex sleep problem that affects approximately 33% of the United States population. Although EDS usually occurs in conjunction with insufficient sleep and other sleep and circadian disorders, recent studies have shown unique genetic markers and metabolic pathways underlying EDS. Here, we aimed to further elucidate the biological profile of EDS using large-scale single- and pathway-level metabolomics analyses. METHODS: Metabolomics data were available for 877 metabolites in 6071 individuals from the Hispanic Community Health Study/Study of Latinos (HCHS/SOL). EDS was assessed using the Epworth Sleepiness Scale (ESS) questionnaire. We performed linear regression for each metabolite on the continuous ESS score, adjusting for demographic, lifestyle, and physiological confounders, and in sex specific groups. Subsequently, gaussian graphical modelling was performed coupled with pathway and enrichment analyses to generate a holistic interactive network of the metabolomic profile of EDS associations. FINDINGS: We identified seven metabolites belonging to steroids, sphingomyelin, and long-chain fatty acids sub-pathways in the primary model associated with EDS, and an additional three metabolites in the male-specific analysis. INTERPRETATION: Our findings indicate that an EDS metabolomic profile is characterised by endogenous and dietary metabolites within the steroid hormone biosynthesis pathway, with some pathways that differ by sex. These pathways may be useful for understanding the causes or consequences of EDS and related sleep disorders. FUNDING: Details regarding funding supporting this work and all studies involved are provided in the acknowledgements section.

Humans

Impacts of climate-driven yield changes on the affordability of healthy diets: a modelling study.

BACKGROUND: Food security is central to global nutrition improvement and public health goals, and healthy diets represent a higher-level aspiration beyond merely avoiding hunger. Climate change poses an increasing threat to food systems by affecting crop yields and food prices. Although climate change-driven risks to hunger have been widely studied, the extent to which climate change undermines the affordability of healthy diets while accounting for socioeconomic responses and regional inequalities remains insufficiently understood. This study aimed to quantify the effects of climate change on the future affordability of healthy diets under alternative socioeconomic and climate scenarios. METHODS: We developed an integrated modelling framework that explicitly couples multimodel crop-yield projections with an integrated assessment model (Global Change Analysis Model [GCAM]). Yield responses from six global gridded crop models driven by four climate models were integrated into GCAM, allowing endogenous socioeconomic adjustments such as land-use shifts, production reallocation, and price responses to emerge under shared socioeconomic pathways (SSPs). Diet affordability was then assessed using the Food and Agriculture Organization of the UN's Cost and Affordability of a Healthy Diet framework across three socioeconomic-climate scenarios (SSP1-2.6, SSP2-4.5, and SSP3-6.0). FINDINGS: Under a high-emissions pathway (ie, SSP3-6.0), climate change was projected to render healthy diets unaffordable for a model-mean of 119 million people globally by 2100, even when CO2 fertilisation effects are included, with the upper end of the model ensemble reaching about 1&#xb7;6 billion people. In contrast, climate-induced affordability losses were found to be negligible under both a low-emissions pathway (ie, SSP1-2.6; -0&#xb7;3 million) and a medium-emission pathway (SSP2-4.5; +0&#xb7;2 million). Under a high-emission pathway, model-mean projections indicated that diet costs could increase by up to 12% in the most affected regions by the end of the century. Under medium emissions, cost increases were projected to remain below 4%, whereas under low emissions, affordability changes were projected to be minimum across regions (within approximately 0&#xb7;5%). Substantial regional disparities emerged, with the largest and most consistent affordability losses concentrated in low-income regions that contributed least to historical greenhouse gas emissions. Under SSP3-6.0, these disparities persisted particularly in regions of Africa and Asia despite projected three-to-five-fold increases in income over the century, with climate-induced disruptions to food systems increasing the number of people unable to afford a healthy diet through mid-century. INTERPRETATION: Climate change is likely to exacerbate global nutritional inequalities by disproportionately increasing the affordability risks of healthy diets in regions that have contributed least to historical greenhouse gas emissions. Under high-warming scenarios, socioeconomic development alone is insufficient to fully offset these risks, highlighting the structural vulnerability of low-income food systems to climate-driven price shocks. These findings suggest that in the absence of targeted interventions, climate change could continue to undermine progress towards equitable and health-oriented nutrition outcomes. FUNDING: Ministry of Science and Technology of the People's Republic of China; National Natural Science Foundation of China; National Aeronautics and Space Administration Goddard Institute for Space Studies Climate Impacts Group; Future of Life Institute; and Global Alliance for Improved Nutrition.

Journal Article

Twisting the End Game: How Telomere Chromatin Modifications Shape Telomere Maintenance.

Cell division inevitably shortens telomeric DNA owing to the end-replication problem. Eukaryotic chromosomes possess specialized telomere structures to maintain genomic stability. In most proliferative cells, telomerase adds telomeric repeats during S-phase. In differentiated cells where telomerase is silenced, telomeres shorten progressively, thereby compromising genomic integrity. Consequently, cancer cells universally activate alternative telomere maintenance mechanisms during malignant transformation: ~80% reactivate telomerase, while a portion of the rest rely on BIR (break-induced replication)-mediated homologous recombination-based ALT (alternative lengthening of telomeres). Although these mechanisms are stable once established, the initial determinants influencing a cancer cell's choice remain poorly understood. This review discusses recent molecular insights into how telomeric chromatin properties profoundly impact this choice. After briefly introducing telomere chromatin characteristics and key players in its maintenance and dynamics, we discuss the mechanisms by which cancer cells acquire distinct telomere replication capabilities. In particular, we present an in-depth analysis linking telomere heterochromatin status to ALT. Furthermore, based on recent advances, we propose a coupled feedforward loop model explaining how the ALT state becomes "locked in" once initiated. Finally, we offer novel perspectives on rational, telomere-centric therapeutic interventions for ALT-positive cancers, focusing on strategies designed to disrupt such feedforward loops by manipulating telomeric chromatin structure.

Humans

Characterization of a Ku-binding motif in the C-terminal region of RAG2.

We applied an unsupervised interactome analysis with the RAG2 C-terminal region (R2CT) in v-abl pro-B cells undergoing V(D)J recombination. Mass-spectrometry analyses showed that Ku70 and Ku80 were among the top 10&#x202f;hits. To further strengthen these observations, we performed Proximity Ligation Assay (PLA) and characterize the existence of a GFP-R2CT-Ku complex formation in cellulo. The interaction of several partners with Ku70/80 (Ku) through Ku-binding motifs (KBMs) in their sequences governs their enrolment in NHEJ repair complexes. Through sequence analysis, we identified a KBM within R2CT (R-KBM, amino acids 589-527). We confirmed by calorimetry a specific micromolar interaction between this RAG2 region and Ku70/80/DNA complex. The RAG2 motif KBM can be subdivided in two conserved parts that have no interaction individually. AlphaFold2 prediction coupled with molecular dynamic simulations indicate that the C-terminal part of the RAG2 motif interacts with Ku80 on the same site than the NHEJ factor XLF. These in silico analyses indicated that the N-terminal part of the RAG2 motif interacts with DNA adjacent to Ku with a major role of the K503 residue in agreement with disruption of the interaction observed with the K503E mutant. This study further extends the large ensemble of proteins recruited at DSBs by KBM motifs and substantiates the model of a tight coupling between DNA breakage and repair during V(D)J recombination, mediated by the Ku-RAG2 C-terminus interaction.

Ku Autoantigen

BMDx2: A Tool for Integrating Toxicogenomics-Based Dose-Dependency Analysis and AOP-Based Mechanistic Insights.

Despite the advent of mechanistic toxicology using omics data to link molecular perturbations with systemic outcomes, regulatory toxicology still lacks the application of mechanism-anchored metrics from such data. This is partially because traditional gene-centric analysis often falls short of linking molecular changes to adverse outcomes. To address this gap, BMDx2, an open-source tool that transforms multi-dose toxicogenomics datasets into quantitative, mechanistic evidence for human chemical safety assessment is developed. BMDx2 couples benchmark-dose modeling with Adverse Outcome Pathway (AOP) enrichment to derive transcriptomic-based points of departure, enabling potency ranking, chemical prioritization, and mechanistically anchored explanations of the effect of chemical exposures. BMDx2 can process a broad range of data, including DNA microarray and RNA sequencing studies. Here, case studies are used to illustrate the versatility of BMDx2 in characterizing the mechanism of action of chemicals. An initial case study on carbon nanotubes exposure applies integrative analysis of transcriptomics and genome-wide DNA methylation data, uncovering cellular reprogramming processes underlying fibrosis. A second case study on bleomycin exposure demonstrate how transcriptomic data alone can be mapped to fibrosis-related AOPs in a standardized, regulatory appropriate manner. Together, these examples show how BMDx2 supports the regulatory application of toxicogenomics and accelerates mechanism-based chemical safety evaluation.

Toxicogenetics

Prediction of bacterial protein-compound interactions with only positive samples.

MOTIVATION: Prediction of Compound-Protein Interactions (CPI) in bacteria is crucial to advance various pharmaceutical and chemical engineering fields, including biocatalysis, drug discovery, and industrial processing. However, current CPI models cannot be applied for bacterial CPI prediction due to the lack of curated negative interaction samples. RESULTS: We propose a novel Positive-Unlabeled (PU) learning framework, named BIN-PU, to address this limitation. BIN-PU generates pseudo positive and negative labels from known positive interaction data, enabling effective training of deep learning models for CPI prediction. We also propose a weighted positive loss function that weights to truly positive samples. We have validated BIN-PU coupled with multiple CPI backbone models, comparing the performance with the existing PU models using bacterial cytochrome P450 (CYP) data. Extensive experiments demonstrate the superiority of BIN-PU over the benchmark models in predicting CPIs with only truly positive samples. Furthermore, we have validated BIN-PU on additional bacterial proteins obtained from literature review, human CYP datasets, and uncurated data for its reproducibility. We have also validated the CPI prediction for the uncurated CYP data with biological and biophysical experiments. BIN-PU represents a significant advancement in CPI prediction for bacterial proteins, opening new possibilities for improving predictive models in related biological interaction tasks. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/datax-lab/CYP.

Bacterial Proteins

Multi-omics identification and functional validation of signal regulatory protein gamma as a prognostic biomarker and immune regulator in head and neck squamous cell carcinoma.

BACKGROUND: Head and neck squamous cell carcinoma (HNSCC) comprises biologically diverse tumors, and durable responses to immune-checkpoint blockade are achieved by only a subset of patients. There remains a need for markers that connect clinical outcome with malignant-cell phenotypes and tissue-level immune organization. METHODS: We integrated The Cancer Genome Atlas HNSCC cohort (TCGA-HNSC), five Gene Expression Omnibus (GEO) validation cohorts, single-cell RNA sequencing, Visium spatial transcriptomics, cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq)-informed protein-potential inference, pharmacogenomic screening, genetic-risk analysis and experimental validation. A reconstructed 296-pipeline survival modelling framework was used to prioritize prognostic hub genes across validation-cohort-specific analyses. RESULTS: SIRPG was repeatedly ranked among the top ten selected genes in all five validation cohorts. At single-cell resolution, SIRPG-high tumor cells showed stronger malignant-cell features, immune-inhibitory and metabolic programs, Scissor-positive risk association, CLCA2/P53-related perturbation signals and inferred SIRPG-CD47/signal regulatory protein (SIRP) communication. Spatial analyses placed this axis within an immune-checkpoint-coupled niche, supported by Maxspin/multiview intercellular spatial modelling (MISTy) spatial coupling, communication analysis by optimal transport (COMMOT)-inferred CD47-SIRPG communication and scProTrans-inferred CD47/SIRPG protein-potential overlap. Functionally, SIRPG knockdown reduced HNSCC cell viability and increased apoptosis, whereas re-expression of short hairpin RNA (shRNA)-resistant SIRPG restored the CLCA2-BAX/BCL2 protein response. CONCLUSION: Together, these findings identify SIRPG as an immune-related prognostic hub and context-dependent tumor-cell regulator associated with apoptosis, immune communication and spatial microenvironmental organization in HNSCC.

Humans

Histone modifications and Sp1 promote GPR160 expression in bone cancer pain within rodent models.

Bone cancer pain (BCP) affects ~70% of patients in advanced stages, primarily due to bone metastasis, presenting a substantial therapeutic challenge. Here, we profile orphan G protein-coupled receptors in the dorsal root ganglia (DRG) following tumor infiltration, and observe a notable increase in GPR160 expression. Elevated Gpr160 mRNA and protein levels persist from postoperative day 6 for over 18 days in the affected DRG, predominantly in small-diameter C-fiber type neurons specific to the tibia. Targeted interventions, including DRG microinjection of siRNA or AAV delivery, mitigate mechanical allodynia, cold, and heat hyperalgesia induced by the tumor. Tumor infiltration increases DRG neuron excitability in wild-type mice, but not in Gpr160 gene knockout mice. Tumor infiltration results in reduced H3K27me3 and increased H3K27ac modifications, enhanced binding of the transcription activator Sp1 to the Gpr160 gene promoter region, and induction of GPR160 expression. Modulating histone-modifying enzymes effectively alleviated pain behavior. Our study delineates a novel mechanism wherein elevated Sp1 levels facilitate Gpr160 gene transcription in nociceptive DRG neurons during BCP in rodents.

Animals

Optimizing gene panels for equitable reproductive carrier screening: The Goldilocks approach.

PURPOSE: Professional organizations recommend pan-ancestry carrier screening for autosomal recessive and X-linked conditions. Advances in DNA sequencing have allowed the analysis of hundreds of genes; however, the optimal number of genes for carrier screening remains unclear. The American College of Medical Genetics and Genomics (ACMG) has proposed a tiered approach recommending screening for 113 genes. METHODS: We analyzed ClinVar and gnomAD v4.1.0, for genes associated with serious autosomal recessive and X-linked conditions and modeled screening performance across panels of varying compositions and sizes in diverse genetic ancestries. We also reevaluated the ACMG gene list using the updated gnomAD data. RESULTS: We identified potential inconsistencies in the ACMG gene lists, particularly in the carrier test performance (defined as a positive yield) for underrepresented genetic ancestry groups. Modeling of the population data for 1310 genes revealed that the screening of 152, 248, 531, and 725 genes achieved 90%, 95%, 99%, and 99.7% positive yields, respectively, in couples. Real-world data from the screening of more than 60,000 couples were used to validate the model. CONCLUSION: Our methodology optimizes the gene content of carrier screening panels for diverse ancestry groups, provides a mechanism for continually updating guidelines, ensures consistency with genomic population data, and improves equity across populations.

Humans

A complex of MAST1 and 14-3-3&#x3b7; regulates Tau phosphorylation in the developing cortex.

The MAST family of serine/threonine kinases has been implicated in a spectrum of human neurodevelopmental disorders. However, little is known about their biological function or regulation. Seeking to fill these gaps in our knowledge, we have identified upstream and downstream partners of MAST1. 14-3-3&#x3b7;, a neuronal 14-3-3 paralog, specifically interacts with MAST1 at two regulatory serines, S90 and S161. p21-activated kinase (PAK), a neuronal regulator of the actin cytoskeleton, phosphorylates MAST1 to regulate its interaction with 14-3-3&#x3b7;. Exploiting mouse models of human Mega-Corpus-Callosum Syndrome (MCC) and whole brain phosphoproteomics, we identify the microtubule-associated protein Tau as a candidate substrate of MAST1. We show that pathogenic MAST1 mutations perturb protein function either through misfolding or attenuation of kinase activity. Our data are consistent with a model in which the MAST kinases couple PAK, a neuronal regulator of the actin cytoskeleton, to microtubule remodeling during the differentiation and specification of cortical neurons.

Animals

Active learning of enhancers and silencers in the developing neural retina.

Deep learning is a promising strategy for modeling cis-regulatory elements. However, models trained on genomic sequences often fail to explain why the same transcription factor can activate or repress transcription in different contexts. To address this limitation, we developed an active learning approach to train models that distinguish between enhancers and silencers composed of binding sites for the photoreceptor transcription factor cone-rod homeobox (CRX). After training the model on nearly all bound CRX sites from the genome, we coupled synthetic biology with uncertainty sampling to generate additional rounds of informative training data. This allowed us to iteratively train models on data from multiple rounds of massively parallel reporter assays. The ability of the resulting models to discriminate between CRX sites with identical sequence but opposite functions establishes active learning as an effective strategy to train models of regulatory DNA. A record of this paper's transparent peer review process is included in the supplemental information.

Retina