PubMed Health⌕ Search

Biomedical subjects

Jesse D Bloom

Publications and source records attributed to Jesse D Bloom.

14 recordsLinked to original sources

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article↗

Near real-time data on the human neutralizing antibody landscape to influenza virus to inform vaccine-strain selection in September 2025.

The hemagglutinin of human influenza virus evolves rapidly to erode neutralizing antibody immunity. Twice per year, new vaccine strains are selected with the goal of providing maximum protection against the viruses that will be circulating when the vaccine is administered ~8-12 months in the future. To help inform this selection, here we quantify how the antibodies in recently collected human sera neutralize viruses with hemagglutinins from contemporary influenza strains. Specifically, we use a high-throughput sequencing-based neutralization assay to measure how 188 human sera collected from Oct 2024 to April 2025 neutralize 140 viruses representative of the H3N2 and H1N1 strains circulating in humans as of the summer of 2025. This data set, which encompasses 26,148 neutralization titer measurements, provides a detailed portrait of the current human neutralizing antibody landscape to influenza A virus. The full data set and accompanying visualizations are available for use in vaccine development and viral forecasting.

Journal Article↗

Pleiotropic mutational effects on function and stability constrain the antigenic evolution of influenza hemagglutinin.

The evolution of human influenza virus hemagglutinin (HA) involves simultaneous selection to acquire antigenic mutations that escape population immunity while preserving protein function and stability. Epistasis shapes this evolution, as an antigenic mutation that is deleterious in one genetic background may become tolerated in another. However, the extent to which epistasis can alleviate pleiotropic conflicts between immune escape and protein function/stability is unclear. Here, we measure how all amino acid mutations in the HA of a recent human H3N2 influenza strain affect its cell entry function, acid stability, and neutralization by human serum antibodies. We find that epistasis has entrenched certain mutations so that reverting to the ancestral amino acid identity in earlier strains is no longer tolerated. Epistasis has also enabled the emergence of antigenic mutations that were detrimental to HA's cell entry function in earlier strains. However, epistasis appears insufficient to overcome the pleiotropic costs of antigenic mutations that impair HA's stability, explaining why some mutations that strongly escape human antibodies never fix in nature. Our results refine our understanding of the mutational constraints that shape recent H3N2 influenza evolution: epistasis can enable antigenic change, but pleiotropic effects can restrict its trajectory.

Journal Article↗

Thermodynamics of neutral protein evolution.

Naturally evolving proteins gradually accumulate mutations while continuing to fold to stable structures. This process of neutral evolution is an important mode of genetic change and forms the basis for the molecular clock. We present a mathematical theory that predicts the number of accumulated mutations, the index of dispersion, and the distribution of stabilities in an evolving protein population from knowledge of the stability effects (delta deltaG values) for single mutations. Our theory quantitatively describes how neutral evolution leads to marginally stable proteins and provides formulas for calculating how fluctuations in stability can overdisperse the molecular clock. It also shows that the structural influences on the rate of sequence evolution observed in earlier simulations can be calculated using just the single-mutation delta deltaG values. We consider both the case when the product of the population size and mutation rate is small and the case when this product is large, and show that in the latter case the proteins evolve excess mutational robustness that is manifested by extra stability and an increase in the rate of sequence evolution. All our theoretical predictions are confirmed by simulations with lattice proteins. Our work provides a mathematical foundation for understanding how protein biophysics shapes the process of evolution.

Computer Simulation↗

Structural determinants of the rate of protein evolution in yeast.

We investigate how a protein's structure influences the rate at which its sequence evolves. Our basic hypothesis is that proteins with highly designable structures (structures that are encoded by many sequences) will evolve more rapidly. Recent theoretical advances argue that structures with a higher density of interresidue contacts are more designable, and we show that high contact density is correlated with an increased rate of sequence evolution in yeast. In addition, we investigate the correlations between the rate of sequence evolution and several other structural descriptors, carefully controlling for the strong effect of expression level on evolutionary rate. Overall, we find that the structural descriptors that we consider appear to explain roughly 10% of the variation in rates of protein evolution in yeast. We also show that despite the well-known trend for buried residues to be more conserved, proteins with a higher fraction of buried residues, nonetheless, tend to evolve their sequences more rapidly. We suggest that this effect is due to the increased designability of structures with more buried residues. Our results provide evidence that protein structure plays an important role in shaping the rate of sequence evolution and provide evidence to support recent theoretical advances linking structural designability to contact density.

Analysis of Variance↗

Structure-guided recombination creates an artificial family of cytochromes P450.

Creating artificial protein families affords new opportunities to explore the determinants of structure and biological function free from many of the constraints of natural selection. We have created an artificial family comprising 3,000 P450 heme proteins that correctly fold and incorporate a heme cofactor by recombining three cytochromes P450 at seven crossover locations chosen to minimize structural disruption. Members of this protein family differ from any known sequence at an average of 72 and by as many as 109 amino acids. Most (>73%) of the properly folded chimeric P450 heme proteins are catalytically active peroxygenases; some are more thermostable than the parent proteins. A multiple sequence alignment of 955 chimeras, including both folded and not, is a valuable resource for sequence-structure-function studies. Logistic regression analysis of the multiple sequence alignment identifies key structural contributions to cytochrome P450 heme incorporation and peroxygenase activity and suggests possible structural differences between parents CYP102A1 and CYP102A2.

Amino Acid Sequence↗

Protein stability promotes evolvability.

The biophysical properties that enable proteins to so readily evolve to perform diverse biochemical tasks are largely unknown. Here, we show that a protein's capacity to evolve is enhanced by the mutational robustness conferred by extra stability. We use simulations with model lattice proteins to demonstrate how extra stability increases evolvability by allowing a protein to accept a wider range of beneficial mutations while still folding to its native structure. We confirm this view experimentally by mutating marginally stable and thermostable variants of cytochrome P450 BM3. Mutants of the stabilized parent were more likely to exhibit new or improved functions. Only the stabilized P450 parent could tolerate the highly destabilizing mutations needed to confer novel activities such as hydroxylating the antiinflammatory drug naproxen. Our work establishes a crucial link between protein stability and evolution. We show that we can exploit this link to discover protein functions, and we suggest how natural evolution might do the same.

Base Pairing↗

Why highly expressed proteins evolve slowly.

Much recent work has explored molecular and population-genetic constraints on the rate of protein sequence evolution. The best predictor of evolutionary rate is expression level, for reasons that have remained unexplained. Here, we hypothesize that selection to reduce the burden of protein misfolding will favor protein sequences with increased robustness to translational missense errors. Pressure for translational robustness increases with expression level and constrains sequence evolution. Using several sequenced yeast genomes, global expression and protein abundance data, and sets of paralogs traceable to an ancient whole-genome duplication in yeast, we rule out several confounding effects and show that expression level explains roughly half the variation in Saccharomyces cerevisiae protein evolutionary rates. We examine causes for expression's dominant role and find that genome-wide tests favor the translational robustness explanation over existing hypotheses that invoke constraints on function or translational efficiency. Our results suggest that proteins evolve at rates largely unrelated to their functions and can explain why highly expressed proteins evolve slowly across the tree of life.

Computational Biology↗

Predicting the tolerance of proteins to random amino acid substitution.

We have recently proposed a thermodynamic model that predicts the tolerance of proteins to random amino acid substitutions. Here we test this model against extensive simulations with compact lattice proteins, and find that the overall performance of the model is very good. We also derive an approximate analytic expression for the fraction of mutant proteins that fold stably to the native structure, Pf(m), as a function of the number of amino acid substitutions m, and present several methods to estimate the asymptotic behavior of Pf(m) for large m. We test the accuracy of all approximations against our simulation results, and find good overall agreement between the approximations and the simulation measurements.

Amino Acid Sequence↗

Thermodynamic prediction of protein neutrality.

We present a simple theory that uses thermodynamic parameters to predict the probability that a protein retains the wild-type structure after one or more random amino acid substitutions. Our theory predicts that for large numbers of substitutions the probability that a protein retains its structure will decline exponentially with the number of substitutions, with the severity of this decline determined by properties of the structure. Our theory also predicts that a protein can gain extra robustness to the first few substitutions by increasing its thermodynamic stability. We validate our theory with simulations on lattice protein models and by showing that it quantitatively predicts previously published experimental measurements on subtilisin and our own measurements on variants of TEM1 beta-lactamase. Our work unifies observations about the clustering of functional proteins in sequence space, and provides a basis for interpreting the response of proteins to substitutions in protein engineering applications.

Amino Acid Substitution↗

Evolving strategies for enzyme engineering.

Directed evolution is a common technique to engineer enzymes for a diverse set of applications. Structural information and an understanding of how proteins respond to mutation and recombination are being used to develop improved directed evolution strategies by increasing the probability that mutant sequences have the desired properties. Strategies that target mutagenesis to particular regions of a protein or use recombination to introduce large sequence changes can complement full-gene random mutagenesis and pave the way to achieving ever more ambitious enzyme engineering goals.

Directed Molecular Evolution↗

Stability and the evolvability of function in a model protein.

Functional proteins must fold with some minimal stability to a structure that can perform a biochemical task. Here we use a simple model to investigate the relationship between the stability requirement and the capacity of a protein to evolve the function of binding to a ligand. Although our model contains no built-in tradeoff between stability and function, proteins evolved function more efficiently when the stability requirement was relaxed. Proteins with both high stability and high function evolved more efficiently when the stability requirement was gradually increased than when there was constant selection for high stability. These results show that in our model, the evolution of function is enhanced by allowing proteins to explore sequences corresponding to marginally stable structures, and that it is easier to improve stability while maintaining high function than to improve function while maintaining high stability. Our model also demonstrates that even in the absence of a fundamental biophysical tradeoff between stability and function, the speed with which function can evolve is limited by the stability requirement imposed on the protein.

Biophysics↗

Apparent dependence of protein evolutionary rate on number of interactions is linked to biases in protein-protein interactions data sets.

BACKGROUND: Several studies have suggested that proteins that interact with more partners evolve more slowly. The strength and validity of this association has been called into question. Here we investigate how biases in high-throughput protein-protein interaction studies could lead to a spurious correlation. RESULTS: We examined the correlation between evolutionary rate and the number of protein-protein interactions for sets of interactions determined by seven different high-throughput methods in Saccharomyces cerevisiae. Some methods have been shown to be biased towards counting more interactions for abundant proteins, a fact that could be important since abundant proteins are known to evolve more slowly. We show that the apparent tendency for interactive proteins to evolve more slowly varies directly with the bias towards counting more interactions for abundant proteins. Interactions studies with no bias show no correlation between evolutionary rate and the number of interactions, and the one study biased towards counting fewer interactions for abundant proteins actually suggests that interactive proteins evolve more rapidly. In all cases, controlling for protein abundance significantly decreases the observed correlation between interactions and evolutionary rate. Finally, we disprove the hypothesis that small data set size accounts for the failure of some interactions studies to show a correlation between evolutionary rate and the number of interactions. CONCLUSIONS: The only correlation supported by a careful analysis of the data is between evolutionary rate and protein abundance. The reported correlation between evolutionary rate and protein-protein interactions cannot be separated from the biases of some protein-protein interactions studies to count more interactions for abundant proteins.

Bacterial Proteins↗