PubMed Health⌕ Search

Biomedical subjects

Alexandre V Morozov

Publications and source records attributed to Alexandre V Morozov.

11 recordsLinked to original sources

Statistical mechanical modeling of genome-wide transcription factor occupancy data by MatrixREDUCE.

MOTIVATION: Regulation of gene expression by a transcription factor requires physical interaction between the factor and the DNA, which can be described by a statistical mechanical model. Based on this model, we developed the MatrixREDUCE algorithm, which uses genome-wide occupancy data for a transcription factor (e.g. ChIP-chip) and associated nucleotide sequences to discover the sequence-specific binding affinity of the transcription factor. Advantages of our approach are that the information for all probes on the microarray is efficiently utilized because there is no need to delineate "bound" and "unbound" sequences, and that, unlike information content-based methods, it does not require a background sequence model. RESULTS: We validated the performance of MatrixREDUCE by inferring the sequence-specific binding affinities for several transcription factors in S. cerevisiae and comparing the results with three other independent sources of transcription factor sequence-specific affinity information: (i) experimental measurement of transcription factor binding affinities for specific oligonucleotides, (ii) reporter gene assays for promoters with systematically mutated binding sites, and (iii) relative binding affinities obtained by modeling transcription factor-DNA interactions based on co-crystal structures of transcription factors bound to DNA substrates. We show that transcription factor binding affinities inferred by MatrixREDUCE are in good agreement with all three validating methods. AVAILABILITY: MatrixREDUCE source code is freely available for non-commercial use at http://www.bussemakerlab.org/. The software runs on Linux, Unix, and Mac OS X.

Algorithms↗

Electron density redistribution accounts for half the cooperativity of alpha helix formation.

The energy of alpha helix formation is well known to be highly cooperative, but the origin and relative importance of the contributions to helical cooperativity have been unclear. Here we separate the energy of helix formation into short range and long range components by using two series of helical dimers of variable length. In one dimer series two monomeric helices interact by forming hydrogen bonds, while in the other they are coupled only through long range, primarily electrostatic interactions. Using Density Functional Theory, we find that approximately half of the cooperativity of helix formation is due to electrostatic interactions between residues, while the other half is due to nonadditive many-body effects brought about by redistribution of electron density with helix length.

Algorithms↗

Protein-DNA binding specificity predictions with structural models.

Protein-DNA interactions play a central role in transcriptional regulation and other biological processes. Investigating the mechanism of binding affinity and specificity in protein-DNA complexes is thus an important goal. Here we develop a simple physical energy function, which uses electrostatics, solvation, hydrogen bonds and atom-packing terms to model direct readout and sequence-specific DNA conformational energy to model indirect readout of DNA sequence by the bound protein. The predictive capability of the model is tested against another model based only on the knowledge of the consensus sequence and the number of contacts between amino acids and DNA bases. Both models are used to carry out predictions of protein-DNA binding affinities which are then compared with experimental measurements. The nearly additive nature of protein-DNA interaction energies in our model allows us to construct position-specific weight matrices by computing base pair probabilities independently for each position in the binding site. Our approach is less data intensive than knowledge-based models of protein-DNA interactions, and is not limited to any specific family of transcription factors. However, native structures of protein-DNA complexes or their close homologs are required as input to the model. Use of homology modeling can significantly increase the extent of our approach, making it a useful tool for studying regulatory pathways in many organisms and cell types.

Base Sequence↗

Potential functions for hydrogen bonds in protein structure prediction and design.

Hydrogen bonds are an important contributor to free energies of biological macromolecules and macromolecular complexes, and hence an accurate description of these interactions is important for progress in biomolecular modeling. A simple description of the hydrogen bond is based on an electrostatic dipole-dipole interaction involving hydrogen-donor and acceptor-acceptor base dipoles, but the physical nature of hydrogen bond formation is more complex. At the most fundamental level, hydrogen bonding is a quantum mechanical phenomenon with contributions from covalent effects, polarization, and charge transfer. Recent experiments and theoretical calculations suggest that both electrostatic and covalent components determine the properties of hydrogen bonds. Likely, the level of rigor required to describe hydrogen bonding will depend on the problem posed. Current approaches to modeling hydrogen bonds include knowledge-based descriptions based on surveys of hydrogen bond geometries in structural databases of proteins and small molecules, empirical molecular mechanics models, and quantum mechanics-based electronic structure calculations. Ab initio calculations of hydrogen bonding energies and geometries accurately reproduce energy landscapes obtained from the distributions of hydrogen bond geometries observed in protein structures. Orientation-dependent hydrogen bonding potentials were found to improve the quality of protein structure prediction and refinement, protein-protein docking, and protein design.

Hydrogen↗

Analysis of anisotropic side-chain packing in proteins and application to high-resolution structure prediction.

pi-pi, Cation-pi, and hydrophobic packing interactions contribute specificity to protein folding and stability to the native state. As a step towards developing improved models of these interactions in proteins, we compare the side-chain packing arrangements in native proteins to those found in compact decoys produced by the Rosetta de novo structure prediction method. We find enrichments in the native distributions for T-shaped and parallel offset arrangements of aromatic residue pairs, in parallel stacked arrangements of cation-aromatic pairs, in parallel stacked pairs involving proline residues, and in parallel offset arrangements for aliphatic residue pairs. We then investigate the extent to which the distinctive features of native packing can be explained using Lennard-Jones and electrostatics models. Finally, we derive orientation-dependent pi-pi, cation-pi and hydrophobic interaction potentials based on the differences between the native and compact decoy distributions and investigate their efficacy for high-resolution protein structure prediction. Surprisingly, the orientation-dependent potential derived from the packing arrangements of aliphatic side-chain pairs distinguishes the native structure from compact decoys better than the orientation-dependent potentials describing pi-pi and cation-pi interactions.

Amino Acids↗

Close agreement between the orientation dependence of hydrogen bonds observed in protein structures and quantum mechanical calculations.

Hydrogen bonding is a key contributor to the exquisite specificity of the interactions within and between biological macromolecules, and hence accurate modeling of such interactions requires an accurate description of hydrogen bonding energetics. Here we investigate the orientation and distance dependence of hydrogen bonding energetics by combining two quite disparate but complementary approaches: quantum mechanical electronic structure calculations and protein structural analysis. We find a remarkable agreement between the energy landscapes obtained from the electronic structure calculations and the distributions of hydrogen bond geometries observed in protein structures. In contrast, molecular mechanics force fields commonly used for biomolecular simulations do not consistently exhibit close correspondence to either quantum mechanical calculations or experimentally observed hydrogen bonding geometries. These results suggest a route to improved energy functions for biological macromolecules that combines the generality of quantum mechanical electronic structure calculations with the accurate context dependence implicit in protein structural analysis.

Amino Acids↗

An improved protein decoy set for testing energy functions for protein structure prediction.

We have improved the original Rosetta centroid/backbone decoy set by increasing the number of proteins and frequency of near native models and by building on sidechains and minimizing clashes. The new set consists of 1,400 model structures for 78 different and diverse protein targets and provides a challenging set for the testing and evaluation of scoring functions. We evaluated the extent to which a variety of all-atom energy functions could identify the native and close-to-native structures in the new decoy sets. Of various implicit solvent models, we found that a solvent-accessible surface area-based solvation provided the best enrichment and discrimination of close-to-native decoys. The combination of this solvation treatment with Lennard Jones terms and the original Rosetta energy provided better enrichment and discrimination than any of the individual terms. The results also highlight the differences in accuracy of NMR and X-ray crystal structures: a large energy gap was observed between native and non-native conformations for X-ray structures but not for NMR structures.

Algorithms↗

Protein-protein docking predictions for the CAPRI experiment.

We predicted structures for all seven targets in the CAPRI experiment using a new method in development at the time of the challenge. The technique includes a low-resolution rigid body Monte Carlo search followed by high-resolution refinement with side-chain conformational changes and rigid body minimization. Decoys (approximately 10(6) per target) were discriminated using a scoring function including van der Waals and solvation interactions, hydrogen bonding, residue-residue pair statistics, and rotamer probabilities. Decoys were ranked, clustered, manually inspected, and selected. The top ranked model for target 6 predicted the experimental structure to 1.5 A RMSD and included 48 of 65 correct residue-residue contacts. Target 7 was predicted at 5.3 A RMSD with 22 of 37 correct residue-residue contacts using a homology model from a known complex structure. Using a preliminary version of the protocol in round 1, target 1 was predicted within 8.8 A although few contacts were correct. For targets 2 and 3, the interface locations and a small fraction of the contacts were correctly identified.

Algorithms↗

An orientation-dependent hydrogen bonding potential improves prediction of specificity and structure for proteins and protein-protein complexes.

Hydrogen bonding is a key contributor to the specificity of intramolecular and intermolecular interactions in biological systems. Here, we develop an orientation-dependent hydrogen bonding potential based on the geometric characteristics of hydrogen bonds in high-resolution protein crystal structures, and evaluate it using four tests related to the prediction and design of protein structures and protein-protein complexes. The new potential is superior to the widely used Coulomb model of hydrogen bonding in prediction of the sequences of proteins and protein-protein interfaces from their structures, and improves discrimination of correctly docked protein-protein complexes from large sets of alternative structures.

Amino Acid Sequence↗

Simple physical models connect theory and experiment in protein folding kinetics.

Our understanding of the principles underlying the protein-folding problem can be tested by developing and characterizing simple models that make predictions which can be compared to experimental data. Here we extend our earlier model of folding free energy landscapes, in which each residue is considered to be either folded as in the native state or completely disordered, by investigating the role of additional factors representing hydrogen bonding and backbone torsion strain, and by using a hybrid between the master equation approach and the simple transition state theory to evaluate kinetics near the free energy barrier in greater detail. Model calculations of folding phi-values are compared to experimental data for 19 proteins, and for more than half of these, experimental data are reproduced with correlation coefficients between r=0.41 and 0.88; calculations of transition state free energy barriers correlate with rates measured for 37 single domain proteins (r=0.69). The model provides insight into the contribution of alternative-folding pathways, the validity of quasi-equilibrium treatments of the folding landscape, and the magnitude of the Arrhenius prefactor for protein folding. Finally, we discuss the limitations of simple native-state-based models, and as a more general test of such models, provide predictions of folding rates and mechanisms for a comprehensive set of over 400 small protein domains of known structure.

Bacterial Proteins↗

Two-frequency mutual coherence function of electromagnetic waves in random media: a path-integral variational solution.

By use of path-integral methods, a general expression is obtained for the two-frequency, two-position mutual coherence function of an electromagnetic pulse propagating through turbulent atmosphere. This expression is valid for arbitrary models of refractive-index fluctuations, wide band pulses, and turbulence of arbitrary strength. The approach presented in this paper was examined in the cases of plane-wave, spherical wave, and Gaussian beam propagation in power-law turbulence and compared with existing numerical and exact results. A number of new results were obtained for the Gaussian beam pulse. Expressions derived here should be applicable to a wide range of practical pulse propagation problems.

Journal Article↗