PubMed Health⌕ Search

Biomedical subjects

Bruce R Donald

Publications and source records attributed to Bruce R Donald.

9 recordsLinked to original sources

A complete algorithm to resolve ambiguity for intersubunit NOE assignment in structure determination of symmetric homo-oligomers.

Assignment of nuclear Overhauser effect (NOE) data is a key bottleneck in structure determination by NMR. NOE assignment resolves the ambiguity as to which pair of protons generated the observed NOE peaks, and thus should be restrained in structure determination. In the case of intersubunit NOEs in symmetric homo-oligomers, the ambiguity includes both the identities of the protons within a subunit, and the identities of the subunits to which they belong. This paper develops an algorithm for simultaneous intersubunit NOE assignment and C(n) symmetric homo-oligomeric structure determinations, given the subunit structure. By using a configuration space framework, our algorithm guarantees completeness, in that it identifies structures representing, to within a user-defined similarity level, every structure consistent with the available data (ambiguous or not). However, while our approach is complete in considering all conformations and assignments, it avoids explicit enumeration of the exponential number of combinations of possible assignments. Our algorithm can draw two types of conclusions not possible under previous methods: (1) that different assignments for an NOE would lead to different structural classes, or (2) that it is not necessary to uniquely assign an NOE, since it would have little impact on structural precision. We demonstrate on two test proteins that our method reduces the average number of possible assignments per NOE by a factor of 2.6 for MinE and 4.2 for CCMP. It results in high structural precision, reducing the average variance in atomic positions by factors of 1.5 and 3.6, respectively.

Algorithms↗

Redesigning the PheA domain of gramicidin synthetase leads to a new understanding of the enzyme's mechanism and selectivity.

The PheA domain of gramicidin synthetase A, a non-ribosomal peptide synthetase, selectively binds phenylalanine along with ATP and Mg2+ and catalyzes the formation of an aminoacyl adenylate. In this study, we have used a novel protein redesign algorithm, K*, to predict mutations in PheA that should exhibit improved binding for tyrosine. Interestingly, the introduction of two predicted mutations to PheA did not significantly improve KD, as measured by equilibrium fluorescence quenching. However, the mutations improved the specificity of the enzyme for tyrosine (as measured by kcat/KM), primarily driven by a 56-fold improvement in KM, although the improvement did not make tyrosine the preferred substrate over phenylalanine. Using stopped-flow fluorometry, we examined binding of different amino acid substrates to the wild-type and mutant enzymes in the pre-steady state in order to understand the improvement in KM. Through these investigations, it became evident that substrate binding to the wild-type enzyme is more complex than previously described. These experiments show that the wild-type enzyme binds phenylalanine in a kinetically selective manner; no other amino acids tested appeared to bind the enzyme in the early time frame examined (500 ms). Furthermore, experiments with PheA, phenylalanine, and ATP reveal a two-step binding process, suggesting that the PheA-ATP-phenylalanine complex may undergo a conformational change toward a catalytically relevant intermediate on the pathway to adenylation; experiments with PheA, phenylalanine, and other nucleotides exhibit only a one-step binding process. The improvement in KM for the mutant enzyme toward tyrosine, as predicted by K*, may indicate that redesigning the side-chain binding pocket allows the substrate backbone to adopt productive conformations for catalysis but that further improvements may be afforded by modeling an enzyme:ATP:substrate complex, which is capable of undergoing conformational change.

Chorismate Mutase↗

Structure determination of symmetric homo-oligomers by a complete search of symmetry configuration space, using NMR restraints and van der Waals packing.

Structural studies of symmetric homo-oligomers provide mechanistic insights into their roles in essential biological processes, including cell signaling and cellular regulation. This paper presents a novel algorithm for homo-oligomeric structure determination, given the subunit structure, that is both complete, in that it evaluates all possible conformations, and data-driven, in that it evaluates conformations separately for consistency with experimental data and for quality of packing. Completeness ensures that the algorithm does not miss the native conformation, and being data-driven enables it to assess the structural precision possible from data alone. Our algorithm performs a branch-and-bound search in the symmetry configuration space, the space of symmetry axis parameters (positions and orientations) defining all possible C(n) homo-oligomeric complexes for a given subunit structure. It eliminates those symmetry axes inconsistent with intersubunit nuclear Overhauser effect (NOE) distance restraints and then identifies conformations representing any consistent, well-packed structure to within a user-defined similarity level. For the human phospholamban pentamer in dodecylphosphocholine micelles, using the structure of one subunit determined from a subset of the experimental NMR data, our algorithm identifies a diverse set of complex structures consistent with the nine intersubunit NOE restraints. The distribution of determined structures provides an objective characterization of structural uncertainty: backbone RMSD to the previously determined structure ranges from 1.07 to 8.85 A, and variance in backbone atomic coordinates is an average of 12.32 A(2). Incorporating vdW packing reduces structural diversity to a maximum backbone RMSD of 6.24 A and an average backbone variance of 6.80 A(2). By comparing data consistency and packing quality under different assumptions of oligomeric number, our algorithm identifies the pentamer as the most likely oligomeric state of phospholamban, demonstrating that it is possible to determine the oligomeric number directly from NMR data. Additional tests on a number of homo-oligomers, from dimer to heptamer, similarly demonstrate the power of our method to provide unbiased determination and evaluation of homo-oligomeric complex structures.

Algorithms↗

Improved Pruning algorithms and Divide-and-Conquer strategies for Dead-End Elimination, with application to protein design.

MOTIVATION: Structure-based protein redesign can help engineer proteins with desired novel function. Improving computational efficiency while still maintaining the accuracy of the design predictions has been a major goal for protein design algorithms. The combinatorial nature of protein design results both from allowing residue mutations and from the incorporation of protein side-chain flexibility. Under the assumption that a single conformation can model protein folding and binding, the goal of many algorithms is the identification of the Global Minimum Energy Conformation (GMEC). A dominant theorem for the identification of the GMEC is Dead-End Elimination (DEE). DEE-based algorithms have proven capable of eliminating the majority of candidate conformations, while guaranteeing that only rotamers not belonging to the GMEC are pruned. However, when the protein design process incorporates rotameric energy minimization, DEE is no longer provably-accurate. Hence, with energy minimization, the minimized-DEE (MinDEE) criterion must be used instead. RESULTS: In this paper, we present provably-accurate improvements to both the DEE and MinDEE criteria. We show that our novel enhancements result in a speedup of up to a factor of more than 1000 when applied in redesign for three different proteins: Gramicidin Synthetase A, plastocyanin, and protein G. AVAILABILITY: Contact authors for source code.

Algorithms↗

A subgroup algorithm to identify cross-rotation peaks consistent with non-crystallographic symmetry.

Molecular replacement (MR) often plays a prominent role in determining initial phase angles for structure determination by X-ray crystallography. In this paper, an efficient quaternion-based algorithm is presented for analyzing peaks from a cross-rotation function in order to identify model orientations consistent with proper non-crystallographic symmetry (NCS) and to generate proper NCS-consistent orientations missing from the list of cross-rotation peaks. The algorithm, CRANS, analyzes the rotation differences between each pair of cross-rotation peaks to identify finite subgroups. Sets of rotation differences satisfying the subgroup axioms correspond to orientations compatible with the correct proper NCS. The CRANS algorithm was first tested using cross-rotation peaks computed from structure-factor data for three test systems and was then used to assist in the de novo structure determination of dihydrofolate reductase-thymidylate synthase (DHFR-TS) from Cryptosporidium hominis. In every case, the CRANS algorithm runs in seconds to identify orientations consistent with the observed proper NCS and to generate missing orientations not present in the cross-rotation peak list. The CRANS algorithm has application in every molecular-replacement phasing effort with proper NCS.

Algorithms↗

Phylogenetic classification of protozoa based on the structure of the linker domain in the bifunctional enzyme, dihydrofolate reductase-thymidylate synthase.

We have determined the crystal structure of dihydrofolate reductase-thymidylate synthase (DHFR-TS) from Cryptosporidium hominis, revealing a unique linker domain containing an 11-residue alpha-helix that has extensive interactions with the opposite DHFR-TS monomer of the homodimeric enzyme. Analysis of the structure of DHFR-TS from C. hominis and of previously solved structures of DHFR-TS from Plasmodium falciparum and Leishmania major reveals that the linker domain primarily controls the relative orientation of the DHFR and TS domains. Using the tertiary structure of the linker domains, we have been able to place a number of protozoa in two distinct and dissimilar structural families corresponding to two evolutionary families and provide the first structural evidence validating the use of DHFR-TS as a tool of phylogenetic classification. Furthermore, the structure of C. hominis DHFR-TS calls into question surface electrostatic channeling as the universal means of dihydrofolate transport between TS and DHFR in the bifunctional enzyme.

Amino Acid Sequence↗

Probabilistic disease classification of expression-dependent proteomic data from mass spectrometry of human serum.

We have developed an algorithm called Q5 for probabilistic classification of healthy versus disease whole serum samples using mass spectrometry. The algorithm employs principal components analysis (PCA) followed by linear discriminant analysis (LDA) on whole spectrum surface-enhanced laser desorption/ionization time of flight (SELDI-TOF) mass spectrometry (MS) data and is demonstrated on four real datasets from complete, complex SELDI spectra of human blood serum. Q5 is a closed-form, exact solution to the problem of classification of complete mass spectra of a complex protein mixture. Q5 employs a probabilistic classification algorithm built upon a dimension-reduced linear discriminant analysis. Our solution is computationally efficient; it is noniterative and computes the optimal linear discriminant using closed-form equations. The optimal discriminant is computed and verified for datasets of complete, complex SELDI spectra of human blood serum. Replicate experiments of different training/testing splits of each dataset are employed to verify robustness of the algorithm. The probabilistic classification method achieves excellent performance. We achieve sensitivity, specificity, and positive predictive values above 97% on three ovarian cancer datasets and one prostate cancer dataset. The Q5 method outperforms previous full-spectrum complex sample spectral classification techniques and can provide clues as to the molecular identities of differentially expressed proteins and peptides.

Algorithms↗

A novel ensemble-based scoring and search algorithm for protein redesign and its application to modify the substrate specificity of the gramicidin synthetase a phenylalanine adenylation enzyme.

Realization of novel molecular function requires the ability to alter molecular complex formation. Enzymatic function can be altered by changing enzyme-substrate interactions via modification of an enzyme's active site. A redesigned enzyme may either perform a novel reaction on its native substrates or its native reaction on novel substrates. A number of computational approaches have been developed to address the combinatorial nature of the protein redesign problem. These approaches typically search for the global minimum energy conformation among an exponential number of protein conformations. We present a novel algorithm for protein redesign, which combines a statistical mechanics-derived ensemble-based approach to computing the binding constant with the speed and completeness of a branch-and-bound pruning algorithm. In addition, we developed an efficient deterministic approximation algorithm, capable of approximating our scoring function to arbitrary precision. In practice, the approximation algorithm decreases the execution time of the mutation search by a factor of ten. To test our method, we examined the Phe-specific adenylation domain of the nonribosomal peptide synthetase gramicidin synthetase A (GrsA-PheA). Ensemble scoring, using a rotameric approximation to the partition functions of the bound and unbound states for GrsA-PheA, is first used to predict binding of the wildtype protein and a previously described mutant (selective for leucine), and second, to switch the enzyme specificity toward leucine, using two novel active site sequences computationally predicted by searching through the space of possible active site mutations. The top scoring in silico mutants were created in the wetlab and dissociation/binding constants were determined by fluorescence quenching. These tested mutations exhibit the desired change in specificity from Phe to Leu. Our ensemble-based algorithm, which flexibly models both protein and ligand using rotamer-based partition functions, has application in enzyme redesign, the prediction of protein-ligand binding, and computer-aided drug design.

Adenosine Triphosphate↗