PubMed HealthSearch

SEARCH · PubMed Health

Results for “Motif discovery”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Functional Motif Discovery in FOXO1 Through CRISPR/Cas9 Exon Tiling Scan.

The study of FOXO1, a pivotal transcription factor, has garnered significant attention due to its critical role in diverse cellular processes, including lineage differentiation, apoptosis, cell cycle regulation, and metabolism. To comprehensively understand the functional intricacies of FOXO1, an innovative approach is essential. This chapter highlights employing CRISPR exon scanning as a strategic tool to dissect the functional domains of FOXO1 and unravel its diverse regulatory functions. CRISPR exon scan allows for the identification of functionally important domains based on the levels of sgRNA depletion or enrichment within the FOXO1 gene, providing a unique opportunity to investigate the domain function under relevant biological contexts. This approach enables the systematic exploration of FOXO1's structural domains, shedding light on how distinct regions contribute to its overall function. The comprehensive exon scan analysis using CRISPR technology allows gaining a nuanced understanding of FOXO1's functional diversity and regulatory mechanisms.

Forkhead Box Protein O1

Immunopeptidomics-driven MHC class II peptide-binding motif discovery for 2 common canine DR alleles.

Despite the central role of major histocompatibility complex (MHC) class II in adaptive immunity, peptide-binding motifs have yet to be characterized for any canine MHC class II allele. Here, we report the first immunopeptidomics-derived binding motifs for DLA-DRB1*015:01 (DLA-DR15) and DLA-DRB1*012:01 (DLA-DR12), 2 alleles overrepresented in breeds predisposed to immune-mediated diseases. Because dogs co-express DLA-DR and DLA-DQ, the MHC class II Ab clone YKIX334.2 was validated to be DLA-DR-specific, enabling allele-selective immunoaffinity purification of DLA-DR molecules from homozygous DLA-DR15 and DLA-DR12 donor spleens. Mass spectrometry and GibbsCluster motif deconvolution of 838 DLA-DR15-associated and 644 DLA-DR12-associated peptides eluted from their respective peptide-binding grooves revealed distinct allele-specific binding motifs, with characterization of anchor residue preferences, peptide-length distributions, cross-species comparisons with human and murine MHC class II motifs, and source protein composition of the eluted self-peptidome. To evaluate the translational utility of these motifs, recombinant DLA-DR15 and DLA-DR12 molecules were used to screen rabies virus glycoprotein and nucleoprotein peptide libraries via fluorescence-based peptide competition assays, identifying high-affinity candidate binders for both alleles. Spearman rank correlation between immunopeptidomics-derived position-specific scoring matrix scores and peptide competition assay rankings demonstrated modest associations, consistent with these approaches capturing complementary dimensions of peptide-MHC class II interaction. Ultimately, these findings establish what we believe is the first allele-specific peptide-binding motif framework for canine MHC class II, providing a foundation for DLA-allele-informed CD4+ T-cell epitope discovery studies and Ag-specific immune response characterization in the dog.

Animals

GAMMA: gap-aware motif mining under incomplete labeling with applications to MHC motifs.

MOTIVATION: Sequence motif identification is crucial for understanding molecular recognition, particularly in immune responses involving peptide binding to major histocompatibility complex (MHC) Class I molecules for antigen presentation to T cells. Traditionally, MHC Class I binding motifs are assumed to be contiguous and span nine amino acids. However, structural evidence suggests that binding may involve nonadjacent residues, challenging the assumptions of existing methods. RESULTS: In this study, we propose Gap-Aware Motif Mining Algorithm (GAMMA), a probabilistic framework designed to identify noncontiguous motifs under conditions of incomplete labeling. GAMMA employs Bayesian inference with Markov chain Monte Carlo sampling to jointly estimate motif parameters, binding locations, and the relative spacing between binding positions. Through extensive simulations and real-world applications to MHC Class I peptide datasets, GAMMA outperforms existing motif discovery tools such as GLAM2 in accurately localizing binding residues and identifying the underlying motifs. Notably, our results suggest that the true number of binding residues may be eight, fewer than the commonly assumed nine. In addition, for longer peptides, the model captures increased flexibility in the central region, consistent with structural observations that peptides may bulge in the middle. AVAILABILITY AND IMPLEMENTATION: The raw data and the source codes are available on GitHub (https://github.com/RanLIUaca/GAMMAmotif).

Amino Acid Motifs

HPV as a Molecular Hacker: Computational Exploration of HPV-Driven Changes in Host Regulatory Networks.

Human Papillomavirus (HPV), particularly high-risk strains such as HPV16 and HPV18, is a leading cause of cervical cancer and a significant risk factor for several other epithelial malignancies. While the oncogenic mechanisms of viral proteins E6 and E7 are well characterized, the broader effects of HPV infection on host transcriptional regulation remain less clearly defined. This study explores the hypothesis that conserved genomic motifs within the HPV genome may act as molecular decoys, sequestering human transcription factors (TFs) and thereby disrupting normal gene regulation in host cells. Such interactions could contribute to oncogenesis by altering the transcriptional landscape and promoting malignant transformation.We conducted a computational analysis of the genomes of high-risk HPV types using MEME-ChIP for de novo motif discovery, followed by Tomtom for identifying matching human TFs. Protein-protein interactions among the predicted TFs were examined using STRING, and biological pathway enrichment was performed with Enrichr. The analysis identified conserved viral motifs with the potential to interact with host transcription factors (TFs), notably those from the FOX, HOX, and NFAT families, as well as various zinc finger proteins. Among these, SMARCA1, DUX4, and CDX1 were not previously associated with HPV-driven cell transformation. Pathway enrichment analysis revealed involvement in several key biological processes, including modulation of Wnt signaling pathways, transcriptional misregulation associated with cancer, and chromatin remodeling. These findings highlight the multifaceted strategies by which HPV may influence host cellular functions and contribute to pathogenesis. In this context, the study underscores the power of in silico approaches for elucidating viral-host interactions and reveals promising therapeutic targets in computationally predicted regulatory network changes.

Humans

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem

Motif-Cluster: Motif driven prioritization of transcription factor binding clusters.

Genome-wide analyses of transcription factor (TF) motif binding sites have largely emphasized individual high-affinity sites, while overlooking the regulatory importance of locally repetitive motif clusters. Such clusters, including combinations of weak and strong binding sites, can collectively enhance TF occupancy and regulatory activity. Here we present Motif-Cluster, an open-source framework for motif-driven prioritization and visualization of TF binding clusters using sequence information alone. Motif-Cluster integrates a density-based clustering strategy with flexible modeling of binding-site gaps and affinity signals, enabling the identification and ranking of candidate regulatory regions without requiring experimental binding data. Through simulations and multiple real-data analyses, we show that combining gap distributions with binding affinity effectively balances cluster size and signal strength while reducing noise from weak sites. Application to ZNF410 successfully recovers the previously characterized binding clusters in the CHD4 promoter, which are conserved between human and mouse. Additional case studies involving PHB1, TWIST1, and EGR1 further demonstrate the general applicability of the method across diverse transcription factors. Motif-Cluster also provides intuitive visualization and reproducible workflows to facilitate interpretation of spatially dense motif patterns. Overall, Motif-Cluster offers a robust and flexible approach for prioritizing transcription factor regulatory regions from genome-wide motif scans, enabling biological discovery and guiding experimental design, particularly in settings where direct genome-wide binding assays are unavailable.

Transcription Factors

Type and position of repeat interruptions as determinants of disease severity and expansion size in Friedreich ataxia.

PURPOSE: In Friedreich ataxia (FRDA) the size of the smaller GAA expansion is a major determinant of disease severity; interruption motifs were identified after the discovery of the pathogenic expansions; however, their impact is only recently investigated. METHODS: 164 patients with FRDA with biallelic expansions and 15 patients without FRDA were analyzed for interruption(s) number, position, and motif. Expansion size and age at onset of ataxia (AAO) were determined for patients with FRDA. RESULTS: Three groups of patients with FRDA were identified by the simultaneous analysis of the precise distance ("depth") between the interruptions (mostly nontriplet) and the 3' end of the expansion (P < .001), the smaller expansion size (P < .001), and AAO (P < .001). Classical FRDA corresponds to absence of interruption or interruption depth < 8 repeats, with AAO often <15 years (area under the curve [AUC] = 0.90; 95% CI, 0.84-0.96); LOFA to interruption depth of 8 to 18 repeats (AUC = 0.97; 95% CI, 0.94-1), with AAO 15 to 34 years (AUC = 1; 95% CI, 1-1); and vLOFA to interruption depth > 18 (AUC = 0.97; 95% CI, 0.92-1), with AAO > 34 years. Multiple (>5) triplet interruptions hamper further expansion. CONCLUSION: This study provides the molecular basis for a novel classification of FRDA that should be recommended for correct diagnosis.

Humans

A small cationic probe for accurate, punctate discovery of RNA tertiary structure.

RNA molecules fold into intricate three-dimensional tertiary structures that are central to their biological functions. Yet reliably discovering new motifs that form true tertiary interactions remains a major challenge. Here we show that RNA tertiary folding occasionally generates electronegative motifs that react selectively with the small, positively-charged probe trimethyloxonium (TMO). Sites with enhanced reactivity to TMO, compared with the neutral reagent dimethyl sulfate (DMS), are indicative of tertiary structure and define T-sites. These positions share a structural signature in which a reactive nucleobase is adjacent to non-bridging phosphate oxygens, creating localized regions of negative charge. T-sites consistently map to the cores of higher-order structural interactions and functional centers across diverse RNAs, including distinct states in conformational ensembles. In the 10,723-nt dengue virus genome, three strong T-sites were detected, each within a complex structure required for viral replication. Cation-based covalent chemistry enables high-confidence discovery and analysis of functional RNA tertiary motifs across long and complex RNAs, opening new opportunities for transcriptome-wide structural analysis.

RNA electrostatics

cfMethDB: A Comprehensive cfDNA Methylation Data Resource for Cancer Biomarkers.

Cancer is a major global health threat, and early detection is crucial for improving patient outcomes. DNA methylation in circulating cell-free DNA (cfDNA) has emerged as a promising biomarker for non-invasive cancer diagnosis. However, the integration and utilization of existing cfDNA methylation data have been limited, hindering comprehensive research efforts, particularly in the discovery of cfDNA methylation biomarkers. To address this challenge, we introduced cfMethDB, a comprehensive database dedicated to cfDNA methylation in cancer that encompasses 4828 publicly available datasets. Through standardized analysis, we identified 1,048,770 differentially methylated cytosines (DMCs) as candidate biomarkers across seven cancer types. With cfMethDB, we not only identified known cfDNA methylation biomarkers, but also discovered several genes, such as ZIC4, that could be novel biomarkers. Moreover, cfMethDB offers a suite of user-friendly tools, including biomarker evaluation, pan-cancer search, and end motif analysis. We hope that cfMethDB will serve as a valuable platform for the discovery of novel cancer cfDNA methylation biomarkers and facilitate cancer research and clinical applications. cfMethDB is publicly available at https://cfmethdb.hzau.edu.cn/home.

Humans

A noncontiguous code for RNA-guided DNA recognition at the origin of CRISPR-Cas.

CRISPR-Cas provides RNA-mediated adaptive immunity, but how its first RNA-guided effector arose is unclear. In this study, we report the discovery of Viral Interference Programmable Repeat (VIPR) systems consisting of a Vipr protein ancestral to the earliest CRISPR-Cas effectors and VIPR RNAs (vrRNAs) comprising alternating GGY/NN motifs. Unlike canonical guide RNAs that pair with target nucleic acids through contiguous complementarity, vrRNAs recognize double-stranded DNA through a noncontiguous code in which the variable NN dinucleotides collectively specify a gapped target sequence. Natural vrRNA targets suggest that VIPR systems act against competing phages, and we demonstrate programmable phage defense by redirecting the complex for transcriptional repression. These results suggest that adaptive immunity originated from ancient warfare between viruses, revealing a previously unidentified logic for encoding information in sequence.

CRISPR-Cas Systems

CircExor enables interpretable prediction of circRNA localization into extracellular vesicles.

Certain circular RNAs (circRNAs) are selectively enriched in extracellular vesicles (EVs), in which they contribute to intercellular communication and represent promising biomarkers, yet the sequence determinants of their sorting remain unclear. Existing computational predictors are optimized mainly for linear RNAs and rarely address circRNA localization into EVs. Here we introduce circExor, the first framework specifically designed for circRNA EV localization. We curate a dedicated benchmark data set of 2102 circRNAs and implement a variable-length end-to-end concatenation strategy together with k-mer frequency encoding to accommodate circular topology, long sequence length, and length heterogeneity. Using a tree-based classifier, circExor achieves superior performance compared with RNAlocate-v3 and ExoGRU, reaching an AUROC of 0.743 on the internal test set and an average AUROC of 0.680 on the held-out test set. SHAP-based analysis, sequence perturbation analysis, motif mapping, and cell-based experimental validation support the predicted EV tendency and identify YBX1, HNRNPK, HNRNPL, and NOVA2 as candidate RBPs potentially associated with circRNA sorting. CircExor therefore provides a predictive and interpretable framework that links in silico modeling to mechanistic hypotheses, and supports biomarker discovery and candidate prioritization for downstream studies of EV-associated circRNAs.

Journal Article

Circulating inflammatory proteins and osteomyelitis: A bidirectional Mendelian randomization and colocalization analysis.

Circulating inflammatory proteins (CIPs) have been implicated in the progression of osteomyelitis (OM); however, whether these proteins play a causal role or are merely a consequence remains unclear. This study aimed to assess the causal relationships between CIPs and OM using a bidirectional 2-sample Mendelian randomization (MR) approach. MR analyses were performed using genome-wide association study summary statistics for 91 inflammation-related proteins (n&#x2005;=&#x2005;14,824) and OM (1881 cases and 3,91,037 controls). The inverse variance weighted method was used as the primary analytical approach, supplemented by MR-Egger, weighted median, simple mode, and weighted mode methods. Sensitivity analyses were conducted to evaluate heterogeneity, horizontal pleiotropy, and robustness. Colocalization analysis was applied to identify shared causal variants, and pathway enrichment analysis was used to explore underlying biological mechanisms. Forward MR analysis revealed that elevated levels of tumor necrosis factor-beta (TNF-&#x3b2;) were significantly associated with increased OM risk (odds ratio [OR]&#x2005;=&#x2005;1.132; 95% confidence interval [CI]: 1.052-1.217; false discovery rate [FDR]&#x2005;=&#x2005;0.027). Conversely, decreased levels of osteoprotegerin (OR&#x2005;=&#x2005;0.772; 95% CI: 0.671-0.889; FDR&#x2005;=&#x2005;0.015) and adenosine deaminase (OR&#x2005;=&#x2005;0.811; 95% CI: 0.736-0.894; FDR&#x2005;<&#x2005;0.001) were associated with increased OM risk. Reverse MR analysis identified increased levels of interleukin-15 receptor alpha, C-X-C motif chemokine ligand 1, fms-related tyrosine kinase 3 ligand, interleukin-20, interleukin-10 (IL10), C-C motif chemokine ligand 19, and CXCL6 as being significantly associated with OM susceptibility (all FDR&#x2005;<&#x2005;0.05). Colocalization analysis provided strong evidence for a shared causal variant between TNF-&#x3b2; and OM (posterior probability for hypothesis 4&#x2005;=&#x2005;0.999). Enrichment analyses indicated involvement of implicated proteins in Toll-like receptor signaling and T-helper 17 cell differentiation pathways. This study identified several CIPs - including TNF-&#x3b2;, osteoprotegerin, and adenosine deaminase - as potentially causal in OM development. These findings highlight promising targets for future immunomodulatory therapies aimed at preventing or mitigating osteomyelitis.

Humans

Discovery of Isonitrile Lipopeptide Chalkophores from Pathogenic Mycobacteria.

The virulence-associated isonitrile lipopeptide (INLP) biosynthetic gene cluster is conserved across Mycobacterium tuberculosis and many nontuberculous mycobacteria (NTM) pathogens, yet the corresponding mycobacterial metabolites have not been fully characterized, and their biological functions are still debated. Here, we report a precursor neutral loss chromatography based mass spectrometry strategy that enables the targeted discovery of INLPs from Mycobacterium fortuitum, a fast-growing NTM pathogen. By monitoring a characteristic neutral loss of 27.1 Da corresponding to hydrogen cyanide, we identified a family of INLPs directly from bacterial culture extracts. Structural elucidation of a representative compound using NMR and high-resolution MS revealed a distinctive terminal methylated carboxyl group, contrasting with previously reported INLPs bearing linear alcohol, acetal, or cyclic motifs. Bioinformatic analysis and in vitro enzymatic assays identified a methyltransferase encoded within the INLP BGC responsible for methyl ester formation. Furthermore, metal-binding assays demonstrated selective chelation of Cu(I) and Cu(II) by the isolated INLP, but no detectable interaction with Zn(II), suggesting a role in copper homeostasis. These findings represent the first full structural characterization of an INLP from pathogenic mycobacteria, expand our understanding of the enzymes involved in INLP modification, and unequivocally support the copper-binding activity of INLPs from these pathogens.

Lipopeptides

Integrative proximal-ubiquitomics profiling for deubiquitinase substrate discovery applied to USP30.

The growing interest in deubiquitinases (DUBs) as drug targets for modulating critical molecular pathways in disease is fueled by the discovery of their specific cellular roles. A crucial aspect of this fact is the identification of DUB substrates. While mass spectrometry-based proteomic methods can be used to study global changes in cellular ubiquitination following DUB activity perturbation, these datasets often include indirect and downstream ubiquitination events. To enrich for the direct substrates of DUB enzymes, we have developed a proximal-ubiquitome workflow that combines proximity labeling methodology (ascorbate peroxidase-2 [APEX2]) with subsequent ubiquitination enrichment based on the K-&#x3b5;-GG motif. We applied this technology to identify altered ubiquitination events in the vicinity of the DUB ubiquitin-specific protease 30 (USP30) upon its inhibition. Our findings reveal ubiquitination events previously associated with USP30 on TOMM20 and FKBP8, as well as the candidate substrate LETM1, which is deubiquitinated in a USP30-dependent manner.

Humans

Shared CD4+ T cell receptor specificity groups in Crohn's disease and ulcerative colitis.

Inflammatory bowel disease (IBD), encompassing ulcerative colitis (UC) and Crohn's disease (CD), is marked by chronic intestinal inflammation and dysregulated immunity. Although UC and CD affect different areas of the gastrointestinal tract, both diseases share aberrant CD4+ memory T cell responses, with HLA-DRB1 as a major genetic risk factor. HLA-DRB1 encodes MHC class II molecules that influence the CD4+ T cell receptor (TCR) repertoire, yet how these genotypes shape TCR specificity in IBD remains unclear. Here, we genotyped HLA-DRB1 and profiled 3.13 million TCR&#x3b2; sequences from circulating memory CD4+ T cells in 33 IBD patients (20 UC, 13 CD) and 14 healthy controls. Using the GLIPH2 algorithm, we distilled 468,441 candidates based on CDR3 amino acid motifs into 440 high-confidence TCR specificity groups significantly enriched among individuals sharing HLA-DRB1 alleles. Notably, 5 specificity groups were IBD-enriched and were shared between UC and CD, suggesting common antigen targets in both diseases. We also observed increased frequencies of clonally expanded cytotoxic GZMB+PRF1+ memory CD4+ T cells and KIR+CD8+ T cells in a subset of risk-allele carriers with IBD. These findings elucidate distinct, HLA-linked TCR specificity groups in IBD and provide mechanistic insights that may advance antigen discovery and personalized medicine.

Humans

AVITI sequencing of a four-generation CEPH/Utah pedigree confirms low mutation rates at homopolymer loci despite their low sequence complexity.

BACKGROUND: Short tandem repeats (STRs) and homopolymers are among the most mutable loci in the human genome. Despite their presumed mutability owing to replication slippage, homopolymer loci exhibit lower mutation rates and minimal paternal age effects compared to other STRs. This paradox questions if technical limitations, rather than biological mechanisms, explain these observations. RESULTS: We used the Element Biosciences AVITI platform to sequence the genomes of a 48-member, four-generation CEPH/Utah pedigree. As the AVITI platform reduces error rates at repetitive sequences compared to Illumina, this design enabled accurate mutation discovery at 90% of assayed homopolymers and a 1.7-fold increase in discoverable mutations compared to Illumina. We identified a median of 35 de novo homopolymer mutations per trio and a mutation rate of 5.28 &#xd7; 10-5 DNMs per locus per generation, confirming a lower rate than dinucleotides (1.94 &#xd7; 10-4). Most DNMs were single base-pair expansions or contractions. Despite comprising <1% of homopolymer loci, G/C homopolymers showed 18-fold higher mutation rates than A/T homopolymers; in contrast, the high dinucleotide mutation rate is not driven by a particular motif class. Parent-of-origin analysis revealed 78% of homopolymer mutations are paternal in origin, but no significant paternal age effect was observed. CONCLUSIONS: This study confirms that homopolymers exhibit lower mutation rates and lack strong paternal age effects compared to other STRs, likely owing to the combination of a lower propensity to form slippage-causing secondary structures and more efficient mismatch repair. Our set of high-quality mutations suggest these phenomena are biological rather than technical in nature. Finally, we demonstrate that AVITI sequencing unlocks previously intractable regions of the genome and will be a powerful tool for continued investigation of repeat mutation.

AVITI

Proteomic Profiling of Pulmonary Function and Cardiovascular Disease Risk in the Atherosclerosis Risk in Communities Study.

BACKGROUND: Pulmonary function is linked to cardiovascular disease risk; however, the underlying mechanisms remain unclear. We aimed to identify protein biomarkers associated with pulmonary function and examine their impact on incident chronic obstructive pulmonary disease, coronary heart disease, heart failure, and all-cause mortality. METHODS: Data from White and Black Americans in the Atherosclerosis Risk in Communities study (visit 2: N=11&#x2009;354, mean age=57 years; visit 5: N=3517, mean age=75 years), a prospective cohort, were analyzed. Linear regression assessed associations between protein levels and pulmonary function measures, including forced expiratory volume in 1 second and forced vital capacity. The impact of the identified proteins on incident chronic obstructive pulmonary disease, coronary heart disease, heart failure, and mortality was estimated using logistic regression and Cox proportional hazards models. Pathway enrichment and Mendelian randomization explored underlying biological functions and causal effects. RESULTS: Of 4766 proteins analyzed, 364 were cross-sectionally associated with forced expiratory volume in 1 second (and forced vital capacity (false discovery rate<0.05). Ninety-four and 270 proteins had concordant positive and negative effects, respectively. Five pathways related to pulmonary and cardiac function were enriched. Of the 364 proteins, 112 were linked to all 4 outcomes, where 86 were associated with increased risk (odds ratio/hazard ratio [OR/HR], 1.05-1.42) and 26 with reduced risk (OR/HR, 0.69-0.96). Six proteins (STAT3 [signal transducer and activator of transcription 3], MIC-1 [growth differentiation factor 15], apoA-II [apolipoprotein A-II], TPST1 [protein-tyrosine sulfotransferase 1], integrin a1b1 [integrin alpha-I: beta-1 complex], and BLC [C-X-C motif chemokine 13]) showed potential inverse causal effects on with forced expiratory volume in 1 second and forced vital capacity, and integrin a1b1 demonstrated consistent inverse associations with chronic obstructive pulmonary disease, coronary heart disease, and heart failure risks. CONCLUSIONS: Proteins associated with pulmonary function may influence CVD risk. Six proteins, including integrin a1b1, represent promising targets for future interventions.

Aged

PepGen: conditional generation of peptides for MHC binding.

MOTIVATION: Peptide-MHC II binding drives adaptive immunity, yet discovery of novel binder peptides remains challenging due to open binding grooves of MHC-II that accommodate variable-length peptides. While discriminative models perform well, they are unfeasible for generation via enumeration due to vast peptide space (2013&#x2248;8&#xd7;1016 for peptides of length 13 amino acids). Generative AI approaches could accelerate binder design to enable vaccines targeted to particular MHC-II alleles or optimize other peptide chemical properties. RESULTS: We introduce PepGen, the first protein language model for MHC II peptide generation building on Generalized Language Modeling. PepGen conditions on alleles, arbitrary partial peptides including putative TCR-interacting motifs, and continuous binding affinity. Across multiple benchmarks including infilling and de novo generation, PepGen outperformed frequency sampling, Gibbs clustering, and autoregressive baselines. Adjusted log-probabilities enable good classification performance. Experimental validation confirmed that the SARS-CoV-2 peptide TEGALNTPKDHIGTR binding the HLA-DQA101:03-DQB106:03 allele can be redesigned to bind the HLA-DQA101:02-DQB105:02 allele. PepGen generated three putative TCR-motif-preserving binders gaining up to 70% of original MFI. Overall, PepGen provides scalable, motif-constrained MHC II peptide redesign and de novo generation, validated through thorough benchmarks and functional assays. AVAILABILITY AND IMPLEMENTATION: Code and Data are available at https://github.com/DaniTheOrange/PepGen.

Peptides