PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequence-to-function models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5 recordsLinked to original sources

DNA-aware evaluation and debiasing of sequence-to-function models.

MOTIVATION: Genome sequence-to-function (S2F) models are widely used to interpret base-resolution functional genomics assays. Most S2F models are trained and evaluated against observed counts and profile-shapes using statistical objectives and fidelity metrics. These choices are well motivated, but they are DNA-independent. At the same time, experimental measurements arise from DNA-dependent assays with distinct characteristics. This mismatch motivates a complementary DNA-aware evaluation of S2F-predicted and experimental functional genomic tracks. RESULTS: We study DNA-dependency of experimental and S2F-predicted tracks using track-conditional genome language models (cgLMs). cgLMs predict masked nucleotides from a conditioning track under controlled DNA visibility. Across ATAC-seq and TF ChIP-seq peaks from GM12878 and K562, cgLM-probing reveals a consistent masked DNA-decodability gap between many experimental and S2F-predicted tracks. In particular, single-task (e.g. BPNet) and multi-task (e.g. AlphaGenome) S2F-predicted tracks enabled cgLMs to recover masked nucleotides with significantly higher accuracy and confidence than matched experimental tracks. Analyses of nonpeak and dinucleotide-shuffled sequences show that this gap is not confined to peaks and is not captured by standard DNA-agnostic profile-shape fidelity metrics alone. ChromBPNet Tn5-denoised predictions were an exception and behaved closer to the experimental regime, suggesting that staged training may reduce the gap. We then convert this diagnostic into a critic-derived objective, DNA-dependency matching (DDM), using a frozen multi-headed cgLM critic. We introduce Critic-Guided Profile-Shape Editing (CGPSE), a preliminary post hoc debiasing framework for frozen S2F models. In GM12878 ATAC-seq, CGPSE partially reduces the masked DNA-decodability gap for AlphaGenome and BPNet predictions, while exposing a tradeoff with profile-shape fidelity. AVAILABILITY AND IMPLEMENTATION: https://github.com/li-lab-mcgill/dna-aware-s2f-eval.

DNA

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Refining sequence-to-activity models by increasing model resolution.

Decoding the cis-regulatory syntax that controls gene expression is essential for improving our understanding of cell differentiation and disease. To identify regulatory motifs and their regulatory syntax, deep learning based sequence-to-activity (S2A) models learn transcription factor binding motifs and their combinations from DNA sequence by modeling measured chromatin accessibility. Previously, we developed AI-TAC, a S2A model that predicts chromatin accessibility across various immune cell types in multi-task fashion, effectively decoding the regulatory syntax underlying immune cell differentiation. While ATAC-seq is commonly used to measure regional accessibility, it also provides high-resolution profiles, the distribution of Tn5 insertion sites, that offer additional insights into the precise location and strength of TF binding sites. Here we demonstrate that modeling ATAC-seq profiles alongside accessibility consistently improves predictions of differential chromatin accessibility across cell types. Moreover, we also find that multi-task learning across related immune cell types consistently outperforms single-task models. To understand what additional information bpAITAC learns from ATAC-seq profiles, we systematically compare sequence attributions from models trained with and without ATAC-seq profiles. We identify novel motifs with strong effect sizes that emerge only when profile data is included. Our findings suggest that modeling ATAC-seq at base-pair resolution enables the model to learn a more nuanced and sensitive representation of the cis-regulatory syntax driving immune cell-specific chromatin landscapes.

ATAC-seq

Engineered histones reshape chromatin in human cells.

Histone proteins and their variants have been found to play crucial and specialized roles in chromatin organization and the regulation of downstream gene expression; however, the relationship between histone sequence and its effect on chromatin organization remains poorly understood, limiting our functional understanding of sequence variation between distinct subtypes and across evolution and frustrating efforts to rationally design synthetic histones that can be used to engineer specified cell states. Here, we make the first advance towards engineered histone-driven chromatin organization. By expressing libraries of sequence variants of core histones in human cells, we identify variants that dominantly modulate chromatin structure. We further interrogate variants using a combination of imaging, proteomics, and genomics to reveal both cis and trans-acting mechanisms of effect. Functional screening with transcription factor libraries identifies transcriptional programs that are facilitated by engineered histone expression. Double mutation screens combined with protein language models allow us to learn sequence-to-function patterns and design synthetic histone proteins optimized to drive specific chromatin states. This work establishes a foundation for the high-throughput evaluation and engineering of chromatin-associated proteins and positions histones as tunable nodes for understanding and modulating mesoscale chromatin organization.

Journal Article

Comparative characterization of six teleost piscidins reveals distinct antimicrobial, antibiofilm and stability profiles.

Piscidins are cationic α-helical antimicrobial peptides (AMPs) that constitute a key component of the innate immune defense of teleost fish, yet the relationship between their genomic organization, structural properties, and functional specialization remains incompletely understood. In this study, six piscidin peptides from Epinephelus akaara, Seriola dumerili, Thunnus maccoyii, Argyrosomus regius, Dicentrarchus labrax, and Epinephelus coioides were characterized through an integrated sequence-to-function approach combining comparative genomics, structural modeling, physicochemical analysis, and in vitro validation, with the aim of identifying candidates with potential for biomedical and biotechnological applications. All genes studied exhibited the conserved four-exon, three-intron architecture characteristic of teleost piscidins. Structural modeling and circular dichroism confirmed α-helical conformations under membrane-mimetic conditions, despite measurable differences in hydrophobicity, charge distribution, and predicted membrane insertion parameters. Antimicrobial assays revealed distinct functional profiles: Sd_FI25 and Epinecidin_1 displayed broad antibacterial activity against Gram-positive and Gram-negative pathogens, whereas Dl_FI22 showed selective activity with reduced temporal persistence associated with lower peptide stability. Ea_FF25 exhibited comparatively weak antibacterial potency. Antibiofilm activity varied among peptides and did not uniformly parallel planktonic MIC values. Computational predictions further suggested antiviral and antitumoral potential for several sequences, extending their prospective relevance beyond classical antibacterial roles. Conserved genomic architecture and α-helical structure coexist with pronounced functional diversification among teleost piscidins. These findings demonstrate that integrating structural prediction with experimental validation is an effective strategy for identifying fish-derived innate immune peptides as candidates for biomedical applications.

Antimicrobial activity