PubMed HealthSearch

SEARCH · PubMed Health

Results for “AlphaFold”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Methylome profiling of SetDB1-deficient ESCs reveals coordinated epigenetic cross-talk during pluripotency.

SetDB1 is best known for catalyzing H3K9me3, but it also influences H3K27me3 deposition, CTCF-binding, and DNA methylation (DNAme). Given the interplay between DNAme and the other epigenetic features, we profiled DNAme following Setdb1 knockout (KO) in ground-state and serum-grown mouse embryonic stem cells (ESCs) to illuminate DNAme-dependent and -independent functions of SetDB1. Time-course whole-genome bisulfite sequencing of serum-grown ESCs shows that nearly half of SetDB1 binding sites are enriched with DNAme and H3K9me3, primarily at retrotransposons. Upon Setdb1 KO, both H3K9me3 and DNAme are reduced, with DNAme rapidly removed at many sites by TET enzymes. Some retrotransposons, primarily IAPs, are TET-resistant and lose DNAme slowly via passive dilution. Notably, SetDB1-mediated regulation of H3K27me3, CTCF-binding, and SMAD3 are uncoupled from the DNAme-H3K9me3 axis, and from each other. AlphaFold modeling and co-immunoprecipitation mass spectrometry suggest this uncoupling involves competitive binding to distinct SetDB1 protein domains, highlighting the complex coordination underlying SetDB1 functions.

AlphaFold modeling

Identification, characterization and classification of prokaryotic nucleoid-associated proteins.

Common throughout life is the need to compact and organize the genome. Possible mechanisms involved in this process include supercoiling, phase separation, charge neutralization, macromolecular crowding, and nucleoid-associated proteins (NAPs). NAPs are special in that they can organize the genome at multiple length scales, and thus are often considered as the architects of the genome. NAPs shape the genome by either bending DNA, wrapping DNA, bridging DNA, or forming nucleoprotein filaments on the DNA. In this mini-review, we discuss recent advancements of unique NAPs with differing architectural properties across the tree of life, including NAPs from bacteria, archaea, and viruses. To help the characterization of NAPs from the ever-increasing number of metagenomes, we recommend a set of cheap and simple in vitro biochemical assays that give unambiguous insights into the architectural properties of NAPs. Finally, we highlight and showcase the usefulness of AlphaFold in the characterization of novel NAPs.

Archaea

Substrate-Dependent Crosslinking by the Cytochrome P450 From Aminopyruvatide Biosynthesis.

Cytochrome P450s catalyze an array of reactions including crosslinking of aromatic side chains in the biosynthesis of ribosomally synthesized and post-translationally modified peptides (RiPPs). ApyO is a cytochrome P450 that forms a C─C bond between two tyrosines in a YLY motif in the substrate ApyA, the precursor peptide of the RiPP aminopyruvatide. We utilized cell-free translation to generate ApyA variants and probe the substrate tolerance of ApyO. Through AlphaFold-based modelling and in vitro assays, we show that ApyO accepts the 10 C-terminal residues of ApyA and requires a conserved Arg/Lys in the substrate. Inspired by substrate sequences in orthologous biosynthetic gene clusters, we substituted one of the tyrosine residues with a tryptophan and observed that ApyO catalyzed formation of an N─C bond between the indole of Trp and Cε2 of Tyr. ApyO unexpectedly catalyzed formation of a C─O bond between the two tyrosine residues when we substituted the leucine residue in the YLY motif with tyrosine or tryptophan. A peptide containing a biaryl linkage and C-terminal aminopyruvate displayed sub-nanomolar inhibition of select proteases, with the aminopyruvate group critical for activity. Overall, this study demonstrates plasticity in the manner of macrocyclization catalyzed by the P450 ApyO.

biosynthesis

Deciphering the Function and Structure of PA1216 as an S-Adenosyl-l-Methionine Binding Protein Using Differential Scanning Fluorimetry and Circular Dichroism.

Microbes produce bioactive secondary metabolites as toxins, pigments, or virulence factors. These specialized compounds are produced by nonribosomal peptide synthetases (NRPS), polyketide synthases (PKS), or hybrid NRPS/PKS pathways. The genes encoding NRPS and PKS reside in biosynthetic gene clusters (BGCs), some of which have no identified metabolite associated with them. Characterization of these orphan BGCs could provide insights into potential bioactive compounds that have yet to be discovered. Here, we characterize PA1216, a putative methyltransferase embedded within an NRPS BGC in Pseudomonas aeruginosa strain PAO1. We cloned, expressed, and purified PA1216, and developed an optimized differential scanning fluorimetry assay to measure its thermal stability, demonstrating concentration-dependent stabilization in the presence of established methyltransferase cofactors and inhibitors. We then adapted this assay for high-throughput screening of potential PA1216 substrates, identifying destabilizing compounds, including glycyl-glycine dipeptides, amino esters with aromatic or basic side chains, and N-Boc-protected amino acids. In contrast, sodium salts of organic acids stabilized PA1216. Lastly, we employed AlphaFold to construct a predictive model, revealing that PA1216 contains a Rossmann-like fold and a glycine-rich loop, typical of class I methyltransferases, and we corroborated these secondary structural elements using circular dichroism spectroscopy. Overall, these studies illuminate PA1216 function and establish a platform for characterizing cryptic gene clusters within secondary metabolic pathways.

Circular Dichroism

Heterologous expression and optimization of the antimicrobial peptide acidocin 4356 in Komagataella phaffii to target Pseudomonas aeruginosa.

Multidrug-resistant (MDR) pathogens, particularly Pseudomonas aeruginosa, pose a serious global health threat due to their increasing prevalence and limited therapeutic options. Antimicrobial peptides (AMPs) offer promising alternatives to traditional antibiotics, yet their large-scale application remains constrained by high production costs and technical challenges. This research sought to develop a yeast-based system for the cost-efficient synthesis of acidocin 4356 (ACD), an antimicrobial peptide proven effective against P. aeruginosa. A codon-optimized ACD gene was cloned into the pPICZα-A expression vector and integrated into the Komagataella phaffii (formerly Pichia pastoris) GS115 genome. Colony PCR confirmed successful integration, and specific transformants demonstrated expression of the 6 × His-ECS-rACD fusion protein, as verified by SDS-PAGE and dot blot analysis. After Ni-NTA chromatography and enterokinase digestion, rACD was found at ~ 20 kDa instead of 8.3 kDa, suggesting oligomerization or post-translational modifications. Response surface methodology determined the optimal temperature, pH, and methanol concentration for peptide synthesis. Under optimal circumstances (21 °C, pH 6.24, and 1.089% methanol), rACD synthesis increased by 34.12% over baseline conditions (30 °C, pH 6, 1% methanol). AlphaFold structural modeling identified three α-helices in high-confidence regions, implicated in bacterial membrane disruption. Antimicrobial assays demonstrated potent rACD activity against P. aeruginosa, yielding a 58.29% reduction in growth at 150 µg/mL and MIC50 and MIC90 values of 143.04 and 320.64 µg/mL, respectively. These findings underscore K. phaffii as a robust platform for AMP production and highlight rACD's therapeutic potential as an effective agent against MDR P. aeruginosa, warranting further investigation into its clinical and industrial applications. KEY POINTS: • Developing a novel K. phaffii strain for heterologous expression supports efficient rACD peptide production. • Optimized conditions boosted expression yield by 34.12% above the reference fermentation settings. • Recombinant acidocin suppressed Pseudomonas aeruginosa growth by 58%, indicating anti-MDR activity.

Pseudomonas aeruginosa

The functional study of novel KLHL3 missense mutations associated with pseudohypoaldosteronism type II.

BACKGROUND: Pseudohypoaldosteronism type II (PHA II) is an inherited tubulopathy, clinically defined by three hallmark features, including secondary hypertension, hyperchloremic metabolic acidosis, and persistent hyperkalemia occurring despite maintained glomerular filtration function. Herein, we aim to investigate the association of kelch like family member 3 (KLHL3) gene mutations with PHA II. METHODS: Compound heterozygous KLHL3 mutations were identified through whole-exome sequencing and Sanger validation. AlphaFold-based structural modeling, site-directed mutagenesis of Flag-tagged plasmids, and co-immunoprecipitation (Co-IP)/immunoblotting in vivo were combined to analyze mutant protein interactions and ubiquitination effects. RESULTS: A Chinese patient was identified with two previously unreported KLHL3 variants (c.131G > A [p.R44Q] and c.744 C > G [p.Y248*]), exhibiting a biochemical triad of asymptomatic hyperkalemia, mild metabolic acidosis, and borderline hypertension. Administration of thiazide diuretics effectively normalized the patient’s hyperkalemia and hypertension. A p.R44Q missense mutation predicted as variants of uncertain significance (VOUS) by American College of Medical Genetics and Genomics (ACMG) guidelines, and a p.Y248* nonsense mutation predicted as variants of likely pathogenic. Functional study revealed that the two KLHL3 mutations impair its ubiquitination of with-no-lysine kinase 1 (WNK1) and with-no-lysine kinase 4 (WNK4), and further increase phosphorylation of both SPAK (sterile20/sporulation-specific protein-1 related proline/alanine-rich kinase)/OSR1 (oxidative stress response kinase-1) and Na-Cl-cotransporter (NCC). CONCLUSIONS: Our study characterized two previously unreported KLHL3 mutations, followed by comprehensive in vitro functional analyses to elucidate their pathophysiological contributions at the molecular level.

Humans

Application of emerging technologies in the antiviral field.

Viral diseases pose a serious threat to global public health, agriculture, and biosecurity. Conventional antiviral strategies are often limited by an incomplete understanding of disease mechanisms, poor targeting precision, and slow response times. Emerging technologies are now reshaping the landscape of antiviral research. This review examines the roles of four key frontiers, including organoid models, gene editing, AI-driven molecular design, and synthetic biology. Organoids provide physiologically relevant platforms that model virus-host interactions and disease progression. Viral infections remain a major challenge to human and animal health, agriculture, and biosecurity. Progress in antiviral research is constrained by the complexity of viral pathogenesis, the diversity and rapid evolution of viruses, and the limited translational relevance of some traditional model systems. Recent advances in organoid technology, gene editing, artificial intelligence, and synthetic biology are expanding the toolkit available for antiviral research and development. In this review, we discuss how these four technological frontiers contribute to disease modeling, target discovery, molecular design, and translational innovation. Organoids, in particular, provide physiologically relevant systems for investigating viral infection, tissue tropism, host responses, and pathogenesis. Gene editing tools, such as CRISPR, enable precise manipulation of host and viral genomes, facilitating the development of resistant organisms and next-generation vaccine platforms. AI technologies, including AlphaFold for structure prediction and platforms for de novo protein design, address long-standing bottlenecks in structural biology and offer powerful means to engineer antiviral proteins, antibodies, and vaccine antigens. Synthetic biology, guided by the Design-Build-Test-Learn cycle, integrates computational design, genetic assembly, and functional validation into a cohesive pipeline. Together, these technologies form a synergistic workflow that spans disease modeling, target discovery, molecular design, construction, testing, and iterative optimization. This integrated approach is shifting antiviral development from traditional empirical methods toward more precise, intelligent strategies. The review also highlights ongoing challenges in integration and scalability, stressing that high-quality biological datasets and stronger interdisciplinary collaboration are essential for realizing translational potential. By presenting a cohesive view of these converging methodologies, this review offers a framework to guide the intelligent evolution of antiviral strategies in both human and animal health.

Antiviral

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry

A novel PKHD1 missense variant disrupting splicing in a fetus with Caroli disease.

BACKGROUND: Caroli disease (CD) is a rare inherited disorder characterized by dilatation of intrahepatic bile ducts, and prenatal diagnosis of this disease is extremely rare. PKHD1 is the only known causative gene, yet the pathogenicity of most missense variants remains unclear. METHODS: Exome sequencing (ES) was performed on a fetus with clinical features of CD. Candidate variants were validated by Sanger sequencing in the family. The impact of the novel missense variant on pre-mRNA splicing was assessed using minigene assays, and structural modeling of the PKHD1 protein was conducted with AlphaFold 3. RESULTS: At 23 weeks of gestation, the fetus showed hepatic cysts on ultrasound and a "central dot" sign on MRI, suggesting a diagnosis of CD. The fetus also exhibited features of autosomal recessive polycystic kidney disease and oligohydramnios. ES identified and Sanger sequencing confirmed three PKHD1 variants: a paternal nonsense variant c.5323C>T; p.(Arg1775*), and two maternal missense variants c.6682G>C; p.(Glu2228Gln) and c.8012G>T; p.(Arg2671Leu). The variant c.6682G>C is novel and minigene assays demonstrated that it caused exon 40 skipping, leading to an in‑frame deletion (c.6491_6682del; p.(Gly2164_Arg2227del)). Structural modeling predicts that this deletion lies within a large β‑barrel domain and may compromise its structural stability. Conclusion We characterize a novel missense variant that causes aberrant splicing of PKHD1 in CD. This finding underscores the necessity of functional analysis for evaluating the pathogenicity of missense variants, especially those at the last nucleotide of an exon. Our study expands the mutation spectrum of PKHD1 and provides insights into genotype‑phenotype correlations.

Humans

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high β-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50 mM sodium acetate, 6.0 pH, and 60 °C temperature on the seventh day of production with a value of 155.77 ± 3.21 U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately ±34.8-49.1 kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862 Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting

Active Site Assembly by SMG5 as a Mechanism for SMG6 Endonuclease Licencing in Nonsense-mediated mRNA Decay.

Nonsense-mediated mRNA decay (NMD) is a conserved eukaryotic surveillance pathway that eliminates transcripts containing premature termination codons (PTCs). Substantial progress has been made in defining the transcript features that mark aberrant translation termination for NMD activation, yet key mechanistic steps remain incompletely understood - including how recruitment of the central NMD factor UPF1 is coupled to the downstream effector phase in which targeted mRNAs are nucleolytically degraded. In metazoans, NMD employs an endonucleolytic route mediated by SMG6, a PIN-domain nuclease, alongside SMG5 and SMG7, which act downstream of PTC recognition. SMG5 has recently been proposed to licence SMG6 activity, yet the molecular basis of this licencing has remained elusive. Here, we combine AlphaFold structural predictions with biochemical assays to investigate interactions among human SMG5, SMG6, and SMG7. Structural models predict a high-confidence interface between SMG5 and SMG6 PIN domains that forms a composite active site: a conserved SMG5 aspartate (D893) complements the SMG6 acidic triad to reinstate the canonical tetrad required for PIN-domain catalysis. In vitro, SMG6 alone exhibits weak endonucleolytic activity, which is enhanced ∼10-fold by the SMG5 PIN domain. Mutational analyses confirm that conserved residues from both proteins are essential for this composite configuration. Our findings reveal that the SMG5 PIN domain, previously considered catalytically inert, plays a critical role in activating SMG6 by completing its active site. This work provides mechanistic insight into the SMG5-dependent licencing step and uncovers a composite PIN nuclease architecture at the heart of the metazoan NMD effector phase.

Nonsense Mediated mRNA Decay

A conserved distal-tail helical extension defines a tailspike attachment architecture in Gram-negative siphophages.

Rapid growth of bacteriophage genome collections has outpaced functional annotation of tail-tip proteins, limiting comparative analysis of host-recognition structures. Starting from a shared distal-tail gene organization in the Salmonella phages 9NA and Jersey, I developed a morphogenetic bioinformatic framework integrating gene synteny, sequence comparison, profile hidden Markov model (HMM) screening, structural evidence, structure-aware searching, and AlphaFold modeling. Comparison with the experimentally characterized lambda and Sf11 tail assemblies identified a predominantly alpha-helical C-terminal extension of the distal-tail (DT) protein associated with tailspike attachment, termed the distal-tail helical extension (DT-helix). Screening 541,986 proteins from 5167 complete NCBI RefSeq tailed-phage genomes, followed by evidence-based evaluation of sequence, genomic context, and structural architecture, identified 165 curated DT-helical-extension-associated phages. Their DT proteins segregated into six sequence groups. In the four principal multi-member groups, cognate tailspikes showed group-specific conservation in proximal N-terminal regions but substantially greater downstream diversity, consistent with sequence constraint at the DT-tailspike attachment boundary. A complementary ProstT5/Foldseek search supported the established groups but revealed no convincing additional highly divergent family. Together with the experimentally characterized Sf11 attachment interface, these findings define a recurrent morphogenetic architecture linking conserved distal-tail scaffolds to more variable receptor-binding proteins across siphophages infecting Gram-negative bacteria. Although universal exchangeability is not established, the identified scaffold-receptor-binding boundaries provide a framework for molecular characterization and rational phage engineering. Accession-level information for the 165 curated phages is available through PhageTailDB.

Viral Tail Proteins

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

Vinigrol Tricyclic Scaffold Biosynthesis Employs an Atypical Terpene Cyclase and a Multipotent Cyclization Cascade.

Vinigrol (1) is a fungal diterpenoid consisting of a decahydro-1,5-butanonaphthalene ring system with no analogs in nature. Despite immense efforts in synthetic studies, the vinigrol biosynthesis pathway remains largely unknown. Herein, we identified a biosynthetic gene cluster for 1 and fully elucidated the biosynthetic pathway. By employing an AlphaFold-generated model structure, we identified the possible catalytic residues of the noncanonical terpene cyclase and analyzed their function by site-directed mutagenesis. We found that the G340A mutation opened a cryptic pathway for an unprecedented tetracyclic diterpene, defined here as virgarene. Retro-biosynthetic theoretical analysis provided a solid foundation for the complex cyclization pathway for the vinigrol scaffold, its chemical transformation to a structurally distinct bonnadiene, and redirection of the enzymatic cyclization cascade to virgarene. Close inspection of the terpene cyclization pathway via integrated experimental and theoretical approaches would allow efficient exploration of novel terpenoid chemistries.

Cyclization

Integrating multi-omics technologies to decipher microbiome functions.

Multi-omics approaches have revolutionized our understanding of microbial communities by enabling simultaneous interrogation of genomic, transcriptomic, proteomic, and metabolomic data. The systematic integration and analysis of these deep datasets help decipher the functional roles of microbiomes, providing critical insights into microbial activities, interactions, and dynamics across diverse environments. Biological complexity makes multi-omics analysis of a single, isolated organism demanding but highly informative, yet this complexity increases further when samples comprise hundreds to thousands of individual species. As microbiome research continues to expand into clinical, environmental, and engineered systems, standardized workflows, benchmarked datasets, and community-driven initiatives are essential to ensure reproducibility, standardization and interpretability. Establishing and disseminating best practices for experimental design, data processing, and integrative analyses will be critical for maximizing comparability and scientific rigor across studies. This perspective highlights recent advances in multi-omics microbiome research, outlines key obstacles in data integration and metadata harmonization, and proposes a collaborative roadmap for scalable, FAIR-compliant multi-omics investigations and potentially disruptive Artificial Intelligence (AI) advances comparable to those of AlphaFold in the field of microbiome science.

Multiomics

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core Cα RMSDs of 0.62 Å for MHC-I and 0.33 Å for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63 817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides

Rare variant analysis of whole genome sequenced juvenile idiopathic arthritis multiplex pedigrees identifies rare variants in NOD2 and ACVR1.

Juvenile idiopathic arthritis is a complex rheumatic disease that is influenced by environmental and genetic factors. Linkage and genome-wide association studies have identified genes that contribute to the risk of developing juvenile idiopathic arthritis but are limited in their ability to identify disease-risk variants of large effect. Penetrant, heritable risk variants can be detected in high-risk families, but such cases are uncommon due to the low prevalence of juvenile idiopathic arthritis. This study utilizes whole-genome sequencing of 23 multiplex families, the largest such cohort to date, to discover variants and genes relevant to JIA pathogenesis. Pathogenic variants in NOD2 associated with Blau syndrome, an ultra-rare Mendelian inflammatory disorder, are the most recurrent variants in the cohort, consistent with previous reports that milder presentations of Blau syndrome are oftentimes misdiagnosed as juvenile idiopathic arthritis. For the first time, however, rare variants in ACVR1 and SMAD6, integral components of the Bone Morphogenic Protein pathway, are found to be associated with juvenile idiopathic arthritis. Identified ACVR1 variants map to critical protein domains. AlphaFold modeling predicts that the ACVR1 interaction with its inhibitor OGT is disrupted by these variants, indicating that the patient-mutated protein has a gain-of-function phenotype. Drosophila melanogaster expressing either a wild-type or patient-mutated version of ACVR1 exhibit embryonic lethality, with the mutant exhibiting 1.4-fold greater lethality than wild-type. The combination of family-based cohorts for gene discovery, AI-based computational tools, and animal model studies for tests of variant function underscores shared disease pathogenesis between JIA and monogenic disorders of immunity and connective tissue.

Arthritis, Juvenile

Laboratory Evolution Reveals Transcriptional Mechanisms Underlying Thermal Adaptation of Escherichia coli.

Adaptive laboratory evolution is able to generate microbial strains, which exhibit extreme phenotypes, revealing fundamental biological adaptation mechanisms. Here, we use adaptive laboratory evolution to evolve Escherichia coli strains that grow at temperatures as high as 45.3 °C, a temperature lethal to wild-type cells. The strains adopted a hypermutator phenotype and employed multiple systems-level adaptations that made global analysis of the DNA mutations difficult. Given the challenge at the genomic level, we were motivated to uncover high-temperature tolerance adaptation mechanisms at the transcriptomic level. We employed independently modulated gene set (iModulon) analysis to reveal five transcriptional mechanisms underlying growth at high temperatures. These mechanisms were connected to acquired mutations, changes in transcriptome composition, sensory inputs, phenotypes, and protein structures. They are as follows: (i) downregulation of general stress responses while upregulating the specific heat stress responses, (ii) upregulation of flagellar basal bodies without upregulating motility and upregulation fimbriae, (iii) shift toward anaerobic metabolism, (iv) shift in regulation of iron uptake away from siderophore production, and (v) upregulation of yjfIJKL, a novel heat tolerance operon whose structures we predicted with AlphaFold. iModulons associated with these five mechanisms explain nearly half of all variance in the gene expression in the adapted strains. These thermotolerance strategies reveal that optimal coordination of known stress responses and metabolism can be achieved with a small number of regulatory mutations and may suggest a new role for large protein export systems. Adaptive laboratory evolution with transcriptomic characterization is a productive approach for elucidating and interpreting adaptation to otherwise lethal stresses.

Escherichia coli