PubMed HealthSearch

SEARCH · PubMed Health

Results for “AlphaFold predictions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity.

Accurately determining the binding affinity of a ligand with a protein is important for drug design, development, and screening. With the advent of accessible protein structure prediction methods such as AlphaFold, predicted protein 3D structures are readily available; however, methods for predicting binding affinity currently do not take full advantage of 3D protein information. Here, we present CASTER-DTA (Cross-Attention with Structural Target Equivariant Representations for Drug-Target Affinity), which uses an equivariant graph neural network to learn more robust protein representations alongside a standard graph neural network to learn molecular representations to predict drug-target affinity. We augment these representations by incorporating an attention-based mechanism between protein residues and drug atoms to improve interpretability. We show that CASTER-DTA represents a state-of-the-art improvement on multiple benchmarks for predicting drug-target affinity and that it generates novel insights for several related tasks. We then apply CASTER-DTA to create a large resource of the binding affinities of every FDA-approved drug against every protein in the human proteome and make these predictions freely available for download. We also make available a web server for researchers to apply a pretrained CASTER-DTA model for predicting binding affinities between arbitrary proteins and drugs.

deep learning

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Pseudomonas aeruginosa

A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy.

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

Journal Article

Targeting the Disease Response With NlpD and LytM for Effective Nonantibiotic Treatment of Urinary Tract Infections.

BACKGROUND: Finding new ways of treating bacterial infections is essential. The NlpD protein, which inhibits RNA polymerase II (Pol II), has shown therapeutic efficacy against urinary tract infection. This study investigated the mechanism of Pol II inhibition and protection by NlpD and its LytM peptide. METHODS: Recombinant NlpD and LytM were screened for interactions with constituents of the Pol II complex, using AlphaFold predictions and protein interaction technology. Treatment effects were quantified in infected tissues and regulated host response pathways identified by genome-wide transcriptomics analysis in models of acute pyelonephritis and acute cystitis in Irf3-/- and Asc-/- mice, respectively. RESULTS: LytM was shown to interact with constituents of the Pol II multiprotein complex, inhibiting the CDK12 kinase from phosphorylating the Pol II subunit RPB1 and disrupting Pol II complex formation by interfering with the interaction between PAF1C and RPB1. The protection by LytM against acute pyelonephritis was accompanied by a reduction in gene expression in infected kidneys from >1900 significantly regulated genes (fold change >6) in the placebo group to about 150 in LytM-treated mice. The inhibition of gene expression in infected kidneys particularly targeted the excessive innate immune response. A similar effect was observed in acute cystitis. Bacterial clearance was accelerated in both model by LytM treatment, with effects against antibiotic-sensitive and resistant Escherichia coli strains. CONCLUSIONS: The results suggest that inhibiting the disease response of the host, using NlpD or LytM, may offer an efficient alternative to antibiotics in these models.

Animals

Rapidly evolving aphid gall effector proteins exhibit saposin-like folds.

Many insects manipulate plants by injecting effector proteins. In one extreme example of this molecular "hijacking", Hormaphis cornu aphids inject bicycle proteins into Hamamelis virginiana (Witch Hazel), contributing to the development of novel organs called galls. Bicycle proteins share no amino acid sequence similarity with proteins of known function. Here, we report the crystal structures of two divergent bicycle proteins. Both proteins contain saposin-like folds: one with multiple disulfide bonds exhibits a helix swap; the other has no disulfide bonds and possesses two tandem domains. To explore the structural evolution of bicycle proteins, we predicted bicycle protein structures with Alphafold2 (AF2). While AF2 did not recover the two experimental structures using existing databases, it succeeded after we provided multiple sequence alignments (MSAs) containing protein sequences encoded in new genome sequences from closely related aphid species. Using this customized approach at scale, we generated 2400 high-confidence predictions for bicycle proteins from seven aphid species. This dataset revealed that bicycle proteins without cysteines are outliers in fold space and appear to have evolved from ancestral proteins with disulfide-bonded saposin-like folds. While all bicycle proteins contain predicted saposin-like folds, they display a vast diversity of structural and physicochemical properties. While this diversity thwarts prediction of conserved functions encoded in structure, it suggests that bicycle proteins have evolved to target diverse plant processes and/or to evade plant immune surveillance.

AlphaFold predictions

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high β-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50 mM sodium acetate, 6.0 pH, and 60 °C temperature on the seventh day of production with a value of 155.77 ± 3.21 U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately ±34.8-49.1 kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862 Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting

Active Site Assembly by SMG5 as a Mechanism for SMG6 Endonuclease Licencing in Nonsense-mediated mRNA Decay.

Nonsense-mediated mRNA decay (NMD) is a conserved eukaryotic surveillance pathway that eliminates transcripts containing premature termination codons (PTCs). Substantial progress has been made in defining the transcript features that mark aberrant translation termination for NMD activation, yet key mechanistic steps remain incompletely understood - including how recruitment of the central NMD factor UPF1 is coupled to the downstream effector phase in which targeted mRNAs are nucleolytically degraded. In metazoans, NMD employs an endonucleolytic route mediated by SMG6, a PIN-domain nuclease, alongside SMG5 and SMG7, which act downstream of PTC recognition. SMG5 has recently been proposed to licence SMG6 activity, yet the molecular basis of this licencing has remained elusive. Here, we combine AlphaFold structural predictions with biochemical assays to investigate interactions among human SMG5, SMG6, and SMG7. Structural models predict a high-confidence interface between SMG5 and SMG6 PIN domains that forms a composite active site: a conserved SMG5 aspartate (D893) complements the SMG6 acidic triad to reinstate the canonical tetrad required for PIN-domain catalysis. In vitro, SMG6 alone exhibits weak endonucleolytic activity, which is enhanced ∼10-fold by the SMG5 PIN domain. Mutational analyses confirm that conserved residues from both proteins are essential for this composite configuration. Our findings reveal that the SMG5 PIN domain, previously considered catalytically inert, plays a critical role in activating SMG6 by completing its active site. This work provides mechanistic insight into the SMG5-dependent licencing step and uncovers a composite PIN nuclease architecture at the heart of the metazoan NMD effector phase.

Nonsense Mediated mRNA Decay

Rare variant analysis of whole genome sequenced juvenile idiopathic arthritis multiplex pedigrees identifies rare variants in NOD2 and ACVR1.

Juvenile idiopathic arthritis is a complex rheumatic disease that is influenced by environmental and genetic factors. Linkage and genome-wide association studies have identified genes that contribute to the risk of developing juvenile idiopathic arthritis but are limited in their ability to identify disease-risk variants of large effect. Penetrant, heritable risk variants can be detected in high-risk families, but such cases are uncommon due to the low prevalence of juvenile idiopathic arthritis. This study utilizes whole-genome sequencing of 23 multiplex families, the largest such cohort to date, to discover variants and genes relevant to JIA pathogenesis. Pathogenic variants in NOD2 associated with Blau syndrome, an ultra-rare Mendelian inflammatory disorder, are the most recurrent variants in the cohort, consistent with previous reports that milder presentations of Blau syndrome are oftentimes misdiagnosed as juvenile idiopathic arthritis. For the first time, however, rare variants in ACVR1 and SMAD6, integral components of the Bone Morphogenic Protein pathway, are found to be associated with juvenile idiopathic arthritis. Identified ACVR1 variants map to critical protein domains. AlphaFold modeling predicts that the ACVR1 interaction with its inhibitor OGT is disrupted by these variants, indicating that the patient-mutated protein has a gain-of-function phenotype. Drosophila melanogaster expressing either a wild-type or patient-mutated version of ACVR1 exhibit embryonic lethality, with the mutant exhibiting 1.4-fold greater lethality than wild-type. The combination of family-based cohorts for gene discovery, AI-based computational tools, and animal model studies for tests of variant function underscores shared disease pathogenesis between JIA and monogenic disorders of immunity and connective tissue.

Arthritis, Juvenile

Laboratory Evolution Reveals Transcriptional Mechanisms Underlying Thermal Adaptation of Escherichia coli.

Adaptive laboratory evolution is able to generate microbial strains, which exhibit extreme phenotypes, revealing fundamental biological adaptation mechanisms. Here, we use adaptive laboratory evolution to evolve Escherichia coli strains that grow at temperatures as high as 45.3 °C, a temperature lethal to wild-type cells. The strains adopted a hypermutator phenotype and employed multiple systems-level adaptations that made global analysis of the DNA mutations difficult. Given the challenge at the genomic level, we were motivated to uncover high-temperature tolerance adaptation mechanisms at the transcriptomic level. We employed independently modulated gene set (iModulon) analysis to reveal five transcriptional mechanisms underlying growth at high temperatures. These mechanisms were connected to acquired mutations, changes in transcriptome composition, sensory inputs, phenotypes, and protein structures. They are as follows: (i) downregulation of general stress responses while upregulating the specific heat stress responses, (ii) upregulation of flagellar basal bodies without upregulating motility and upregulation fimbriae, (iii) shift toward anaerobic metabolism, (iv) shift in regulation of iron uptake away from siderophore production, and (v) upregulation of yjfIJKL, a novel heat tolerance operon whose structures we predicted with AlphaFold. iModulons associated with these five mechanisms explain nearly half of all variance in the gene expression in the adapted strains. These thermotolerance strategies reveal that optimal coordination of known stress responses and metabolism can be achieved with a small number of regulatory mutations and may suggest a new role for large protein export systems. Adaptive laboratory evolution with transcriptomic characterization is a productive approach for elucidating and interpreting adaptation to otherwise lethal stresses.

Escherichia coli

Application of emerging technologies in the antiviral field.

Viral diseases pose a serious threat to global public health, agriculture, and biosecurity. Conventional antiviral strategies are often limited by an incomplete understanding of disease mechanisms, poor targeting precision, and slow response times. Emerging technologies are now reshaping the landscape of antiviral research. This review examines the roles of four key frontiers, including organoid models, gene editing, AI-driven molecular design, and synthetic biology. Organoids provide physiologically relevant platforms that model virus-host interactions and disease progression. Viral infections remain a major challenge to human and animal health, agriculture, and biosecurity. Progress in antiviral research is constrained by the complexity of viral pathogenesis, the diversity and rapid evolution of viruses, and the limited translational relevance of some traditional model systems. Recent advances in organoid technology, gene editing, artificial intelligence, and synthetic biology are expanding the toolkit available for antiviral research and development. In this review, we discuss how these four technological frontiers contribute to disease modeling, target discovery, molecular design, and translational innovation. Organoids, in particular, provide physiologically relevant systems for investigating viral infection, tissue tropism, host responses, and pathogenesis. Gene editing tools, such as CRISPR, enable precise manipulation of host and viral genomes, facilitating the development of resistant organisms and next-generation vaccine platforms. AI technologies, including AlphaFold for structure prediction and platforms for de novo protein design, address long-standing bottlenecks in structural biology and offer powerful means to engineer antiviral proteins, antibodies, and vaccine antigens. Synthetic biology, guided by the Design-Build-Test-Learn cycle, integrates computational design, genetic assembly, and functional validation into a cohesive pipeline. Together, these technologies form a synergistic workflow that spans disease modeling, target discovery, molecular design, construction, testing, and iterative optimization. This integrated approach is shifting antiviral development from traditional empirical methods toward more precise, intelligent strategies. The review also highlights ongoing challenges in integration and scalability, stressing that high-quality biological datasets and stronger interdisciplinary collaboration are essential for realizing translational potential. By presenting a cohesive view of these converging methodologies, this review offers a framework to guide the intelligent evolution of antiviral strategies in both human and animal health.

Antiviral

[Analysis of a Chinese pedigree affected with Townes-Brocks syndrome due to a novel variant of SALL1 gene and a literature review].

OBJECTIVE: To analyze a novel exonic variant of the SALL1 gene and its impact on the binding site of SALL protein. METHODS: Clinical data of three children diagnosed with Townes-Brocks syndrome and their family members who had presented at the First Affiliated Hospital of Shandong First Medical University in April 2022 were retrospectively collected. The pathogenic variant was identified through whole-genome sequencing (WGS) and validated by Sanger sequencing. Protein structural prediction was performed using AlphaFold and PyMOL software to construct three-dimensional models of the wild-type and mutant proteins. Additionally, previously reported cases were systematically reviewed. This study was approved by the Medical Ethics Committee of the hospital (Ethics No.: 2023-386). RESULTS: The proband was one of triplet sisters born at 34+4 gestational weeks. All three cases had presented with anal atresia and rectovaginal fistula, and case 3 also had toe malformation of left foot. WGS revealed a novel heterozygous c.757C>T (p.Gln253*) variant in the SALL1 gene, which was predicted to be pathogenic. Sanger sequencing confirmed co-segregation of the variant with the disease within the family. Protein structural modeling demonstrated that the variant has introduced a premature stop codon at position 253, resulting in a truncated protein. CONCLUSION: Above finding has enriched the mutation spectrum of the SALL1 gene in association with Townes-Brocks syndrome, which also represented a rare case of anal atresia in triplets, and provided a basis for molecular diagnosis, genetic counseling, and further research.

Humans

Deciphering the Function and Structure of PA1216 as an S-Adenosyl-l-Methionine Binding Protein Using Differential Scanning Fluorimetry and Circular Dichroism.

Microbes produce bioactive secondary metabolites as toxins, pigments, or virulence factors. These specialized compounds are produced by nonribosomal peptide synthetases (NRPS), polyketide synthases (PKS), or hybrid NRPS/PKS pathways. The genes encoding NRPS and PKS reside in biosynthetic gene clusters (BGCs), some of which have no identified metabolite associated with them. Characterization of these orphan BGCs could provide insights into potential bioactive compounds that have yet to be discovered. Here, we characterize PA1216, a putative methyltransferase embedded within an NRPS BGC in Pseudomonas aeruginosa strain PAO1. We cloned, expressed, and purified PA1216, and developed an optimized differential scanning fluorimetry assay to measure its thermal stability, demonstrating concentration-dependent stabilization in the presence of established methyltransferase cofactors and inhibitors. We then adapted this assay for high-throughput screening of potential PA1216 substrates, identifying destabilizing compounds, including glycyl-glycine dipeptides, amino esters with aromatic or basic side chains, and N-Boc-protected amino acids. In contrast, sodium salts of organic acids stabilized PA1216. Lastly, we employed AlphaFold to construct a predictive model, revealing that PA1216 contains a Rossmann-like fold and a glycine-rich loop, typical of class I methyltransferases, and we corroborated these secondary structural elements using circular dichroism spectroscopy. Overall, these studies illuminate PA1216 function and establish a platform for characterizing cryptic gene clusters within secondary metabolic pathways.

Circular Dichroism

Non-syndromic premature ovarian insufficiency associated with monoallelic LIG4 mutation via haploinsufficiency.

BACKGROUND: Premature ovarian insufficiency (POI) is a heterogeneous reproductive disorder, with genetic factors, particularly defects in DNA damage response pathways, increasingly implicated in its pathogenesis. DNA ligase IV (LIG4) is a key enzyme in the non-homologous end joining (NHEJ) pathway responsible for repairing DNA double-strand breaks (DSBs). However, its role in non-syndromic POI remains unclear. This study aimed to investigate the potential contribution of LIG4 variants to non-syndromic POI. RESULTS: Whole-exome sequencing identified a heterozygous frameshift variant in LIG4 (c.1271_1275del) in a three-generation Han Chinese family with non-syndromic POI, which co-segregated with affected individuals. AlphaFold-based structural modeling predicted truncation of the C-terminal XRCC4 interaction region. Functional experiments demonstrated that the mutant LIG4 protein showed reduced stability and was predominantly mislocalized to the cytoplasm of cells. In ovarian KGN cells, LIG4 depletion reduced cell viability, induced stress-associated cellular senescence, and impaired DNA damage repair capacity. In LIG4 knockout 293T cells, co-transfection of wild-type and mutant constructs revealed dose-dependent functional impairment, resulting in increased apoptosis under basal conditions and after phleomycin induced DNA damage, together with delayed repair of DSBs. Reanalysis of public single-cell RNA sequencing data further showed stage specific upregulation of LIG4 during oocyte maturation. Co-expression network analysis revealed enrichment in the Fanconi anemia pathway, phosphatidylinositol 3-kinase signaling pathway, and glycan metabolism. CONCLUSIONS: Our findings suggest that monoallelic LIG4 mutations may represent a potential genetic etiology for non-syndromic POI with sex-limited penetrance. While further validation in more physiologically relevant models is warranted, our data indicate that LIG4 haploinsufficiency may impair DSB repair and disrupt molecular pathways crucial for oocyte maturation and survival, highlighting a potential role of the NHEJ pathway in maintaining human ovarian function.

Humans

PMGen: from peptide-MHC structure prediction to peptide generation.

MOTIVATION: Accurate structural modeling of peptide-major histocompatibility complex (pMHC) complexes is essential for structure-driven immunotherapy design, yet current prediction tools suffer from narrow class coverage, restricted peptide lengths, insufficient accuracy, and a lack of built-in structure-aware peptide sampling. Consequently, most mimotope and altered peptide ligand designs rely solely on sequence substitution, leaving spatial and biophysical insights from pMHC structures largely unexploited. RESULTS: We introduce peptide-MHC generator (PMGen), an integrated framework for structure prediction and structure-guided design of variable-length peptides across MHC Class I and II. PMGen enforces anchor constraints within AlphaFold2 through two complementary strategies, initial guess and template engineering, achieving state-of-the-art structural fidelity without model fine-tuning. On a comprehensive benchmark, PMGen outperforms all existing methods, yielding median peptide-core Cα RMSDs of 0.62 Å for MHC-I and 0.33 Å for MHC-II. We show that PMGen can recover incorrectly predicted anchor positions and that AlphaFold pLDDT scores enable sequence-independent binding-core identification. Applied to a published neoantigen/wild-type pair, PMGen accurately captures mutation-induced conformational changes. Beyond structure prediction, we show that ProteinMPNN sampling on PMGen-predicted backbones yields higher affinity peptides while preserving the parental 3D conformation. Using PMGen to generate 63 817 high-confidence pMHC structures as training data, we further improve ProteinMPNN's peptide sequence recovery from 0.14 to 0.64 on a test set of 85 unseen MHC-I alleles, highlighting the value of accurate predicted structures for downstream machine learning tasks. AVAILABILITY AND IMPLEMENTATION: PMGen is freely available at https://github.com/soedinglab/PMGen, with an interactive Colab notebook at https://colab.research.google.com/github/soedinglab/PMGen/blob/master/colab.ipynb.

Peptides

A novel hemizygous missense variant in the BEND2 gene is associated with nonobstructive azoospermia.

Nonobstructive azoospermia (NOA), the most severe form of male infertility, frequently arises from genetic defects that disrupt spermatogenesis. In this study, a novel hemizygous missense variant (NM_001184767.2 [c.G1069A; p.V357I]) is identified in the X-linked BEN domain-containing 2 ( BEND2 ) gene of a patient with NOA characterized by spermatocyte maturation arrest. Whole-exome sequencing and Sanger validation confirmed that this rare variant is absent in fertile controls and that no pathogenic variants were detected in established NOA genes. Computational analysis predicted potential structural alterations via AlphaFold modeling, leading to the hypothesis that the ability of BEND2 to recognize genomic targets may be compromised. The patient's phenotype phenocopies the meiotic arrest observed in Bend2 -knockout mice. Expression profiling confirmed predominant BEND2 transcription in human and mouse testes, peaking in early spermatocytes and coinciding with meiotic initiation, with reduced transcript levels detected in the proband's peripheral blood compared with those in an obstructive azoospermia control. This study reports a pathogenic BEND2 variant associated with NOA with spermatocyte arrest, highlighting its critical role in human meiosis and expanding the genetic etiology of male infertility.

Adult

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry

Inference of Cytochrome P450 Evolutionary History Using Structural and Physicochemical Metrics.

Cytochrome P450s are a superfamily of heme-binding monooxygenases involved with the detoxification of intrinsic and extrinsic toxins. They are near ubiquitous within biological domains and are found in all domains. Members of families within the superfamily are defined based on amino acid identity thresholds, with thresholds as low as 40% in some families. Relationships among Cytochrome P450 families have proven elusive due to sub-Twilight Zone interfamily identities (<30%) that result in poor multiple sequence alignment quality and thus low levels of support for downstream phylogenetic reconstructions. Despite the low identities, Cytochrome P450 structures are remarkably well conserved both within and among families. In such cases, structural phylogenetics has the potential to unveil elusive relationships because the selectively favored physicochemical properties giving rise to the structure and function of the proteins persist despite sequence-level divergence. Recently, in two separate publications, we demonstrated that by utilizing physicochemical vectors, dynamic time warping, and hierarchical clustering (PCDTW), large swaths of protein domain families and betacoronavirus receptor-binding domain clades were congruent with validated functional/structural relationships. These were important findings because anomalous sequence alignment-based maximum likelihood phylogenetic findings, which were not congruent with the known functional relationships, were resolved. That also validated the use of physicochemical vectors in making inferences about structural/functional homology. Additionally, it illuminated that the same methods might be applied to other protein families with relationships that are difficult to resolve from sequence data alone. Herein, we used Molecular Weight and Hydrophobicity Physicochemical Dynamic Time Warping (MWHP PCDTW) along with structural and sequence alignment-based phylogenetic methodologies to analyze all of the Cytochrome P450s found both in the high-fidelity Structural Classificaction of Proteins (SCOP) database and the reviewed sequences with both experimentally resolved and de novo predicted structures in the Protein Data Bank and the AlphaFold (AF) Protein Structure Database, respectively. We compared the resulting phylogenetic topologies and found that in some cases, structure-based methods may be less able to resolve random/convergent similarity than physicochemical and sequence-based methodologies. This finding agrees with previous findings that demonstrate the usefulness of physicochemical properties in resolving both random structural similarity and potentially convergent relationships.

Cytochrome P-450 Enzyme System

Characterization of putatively lytic bacteriophages able to infect Xanthomonas citri subsp. citri and identification of novel putative exopolysaccharide depolymerases.

Asiatic citrus canker (ACC), caused by the Gram-negative bacterium Xanthomonas citri subsp. citri (X. citri), leads to substantial economic losses in the global citrus industry, necessitating sustainable alternatives to conventional copper-based bactericides. In this study, we isolated and sequenced 72 putatively lytic bacteriophages (including 65 previously uncharacterized isolates from S&#xe3;o Paulo, Brazil) and characterized their host range, stability and biocontrol potential. Genomic analysis revealed highly successful but low-diversity phage genomic signatures; 70 isolates shared ~95%&#x2009;DNA similarity and were closely related to the Japanese phage CP2, mirroring the clonal nature of the endemic X. citri population. These phages primarily belong to the Autographiviridae family, with the exception of the Schitoviridae isolate XacP77. Using HHsearch and AlphaFold structural modelling, we identified conserved tail-fibre genes predicted to encode putative exopolysaccharide depolymerases with structural homology to carbohydrate-binding modules (CBMs), which may facilitate the degradation of the bacterial xanthan gum capsule during infection. While the phages exhibited robust stability across a wide pH range (4-11) and temperatures up to 55&#x2009;&#xb0;C, they were highly sensitive to UV exposure, reaching total inactivation after 160&#x2009;s. Greenhouse assays demonstrated that treatment with phage P27 reduced ACC lesion production by 60%, pointing to the potential of these viruses and their candidate CBM-containing proteins as components of a sustainable biocontrol development within integrated pest management strategies for X. citri.

Xanthomonas