PubMed HealthSearch

Biomedical subjects

R A Abagyan

Publications and source records attributed to R A Abagyan.

At least 19 recordsLinked to original sources

The 1.8 A crystal structure of the dimeric peroxisomal 3-ketoacyl-CoA thiolase of Saccharomyces cerevisiae: implications for substrate binding and reaction mechanism.

The dimeric, peroxisomal 3-ketoacyl-CoA thiolase catalyses the conversion of 3-ketoacyl-CoA into acyl-CoA, which is shorter by two carbon atoms. This reaction is the last step of the beta-oxidation pathway. The crystal structure of unliganded peroxisomal thiolase of the yeast Saccharomyces cerevisiae has been refined at 1.8 A resolution. An unusual feature of this structure is the presence of two helices, completely buried in the dimer and sandwiched between two beta-sheets. The analysis of the structure shows that the sequences of these helices are not hydrophobic, but generate two amphipathic helices. The helix in the N-terminal domain exposes the polar side-chains to a cavity at the dimer interface, filled with structured water molecules. The central helix in the C-terminal domain exposes its polar residues to an interior polar pocket. The refined structure has also been used to predict the mode of binding of the substrate molecule acetoacetyl-CoA, as well as the reaction mechanism. From previous studies it is known that Cys125, His375 and Cys403 are important catalytic residues. In the proposed model the acetoacetyl group fits near the two catalytic cysteine residues, such that the oxygen atoms point towards the protein interior. The distance between SG(Cys125) and C3(acetoacetyl-CoA) is 3.7 A. The O2 atom of the docked acetoacetyl group makes a hydrogen bond to N(Gly405), which would favour the formation of the covalent bond between SG(Cys125) and C3(acetoacetyl-CoA) of the intermediate complex of the two-step reaction. The CoA moiety is proposed to bind in a groove on the surface of the protein molecule. Most of the interactions of the CoA molecule are with atoms of the loop domain. The three phosphate groups of the CoA moiety are predicted to interact with side-chains of lysine and arginine residues, which are conserved in the dimeric thiolases.

Acetyl-CoA C-Acyltransferase

Do aligned sequences share the same fold?

Sequence comparison remains a powerful tool to assess the structural relatedness of two proteins. To develop a sensitive sequence-based procedure for fold recognition, we performed an exhaustive global alignment (with zero end gap penalties) between sequences of protein domains with known three-dimensional folds. The subset of 1.3 million alignments between sequences of structurally unrelated domains was used to derive a set of analytical functions that represent the probability of structural significance for any sequence alignment at a given sequence identity, sequence similarity and alignment score. Analysis of overlap between structurally significant and insignificant alignments shows that sequence identity and sequence similarity measures are poor indicators of structural relatedness in the "twilight zone", while the alignment score allows much better discrimination between alignments of structurally related and unrelated sequences for a wide variety of alignment settings. A fold recognition benchmark was used to compare eight different substitution matrices with eight sets of gap penalties. The best performing matrices were Gonnet and Blosum50 with normalized gap penalties of 2.4/0.15 and 2.0/0.15, respectively, while the positive matrices were the worst performers. The derived functions and parameters can be used for fold recognition via a multilink chain of probability weighted pairwise sequence alignments.

Databases as Topic

Contact area difference (CAD): a robust measure to evaluate accuracy of protein models.

A simple unified measure to evaluate the accuracy of three-dimensional atomic protein models is proposed. This measure is a normalized sum of absolute differences of residue-residue contact surface areas calculated for a reference structure and a model. It employs more rigorous quantitative evaluation of a contact than previously used contact measures. We argue that the contact area difference (CAD) number is a robust single measure to evaluate protein structure predictions in a wide range of model accuracies, from ab initio and threading models to models by homology, since it reflects both backbone topology and side-chain packing, is smooth, continuous and threshold-free, is not sensitive to typical crystallographic errors and ambiguities, adequately penalizes domain and/or secondary structure rearrangements and protein plasticity, and has consistent linear and matrix representations for more detailed analysis. The CAD quality of crystallographic structures, NMR structures, models by homology, and unfolded and misfolded structures is evaluated. It is shown that the CAD number discriminates between models better than Cartesian root-mean-square deviation (cRMSD). Structural variability of the NMR structures was found to be three times larger than deformations of crystallographic structures in different packing environments.

Amino Acids

Protein engineering with monomeric triosephosphate isomerase (monoTIM): the modelling and structure verification of a seven-residue loop.

Protein engineering experiments have been carried out with loop-1 of monomeric triosephosphate isomerase (monoTIM). Loop-1 of monoTIM is disordered in every crystal structure of liganded monoTIM, but in the wild-type TIM it is a very rigid dimer interface loop. This loop connects the first beta-strand with the first alpha-helix of the TIM-barrel scaffold. The first residue of this loop, Lys13, is a conserved catalytic residue. The protein design studies with loop-1 were aimed at rigidifying this loop such that the Lys13 side chain points in the same direction as seen in wild type. The modelling suggested that the loop should be made one residue shorter. With the modelling package ICM the optimal sequence of a new seven-residue loop-1 was determined and its structure was predicted. The new variant could be expressed and purified and has been characterized. The catalytic activity and stability are very similar to those of monoTIM. The crystal structure (at 2.6 A resolution) shows that the experimental loop-1 structure agrees well with the modelled loop-1 structure. The direct superposition of the seven loop residues of the modelled and experimental structures results in an r.m.s. difference of 0.5 A for the 28 main chain atoms. The good agreement between the predicted structure and the crystal structure shows that the described modelling protocol can be used successfully for the reliable prediction of loop structures.

Amino Acid Sequence

Towards protein folding by global energy optimization.

Different components of the theoretical protein folding problem are evaluated critically. It is argued that: (i) as a rule, small- and medium-sized proteins are in the free energy minimum; (ii) long-living metastable states may either appear occasionally with growing protein size, or be selected by evolution for a specific function; (iii) functions discriminating against incorrect folds would fail if they were used directly in the global optimization, unless they approximate the true free energy accurately; (iv) surface and electrostatic free energies should be treated separately; (v) conformational entropy (of side chains in particular) should be taken into account; (vi) Monte Carlo procedures considering all free energy terms and combining global knowledge-based random moves with local optimization have the largest potential for success.

Electrochemistry

A technique for identifying atoms from a screen image.

Improving the interfaces in molecular graphics applications, making them more natural and easy to use, is an important task, given the current complexity of the displayed objects and of modeling operations. Clicking near an atom center is the usual method of atom selection. However, this method has certain disadvantages when working with images composed of different atomic representations such as sticks, CPK, or dotted surfaces. We propose another technique allowing the user to obtain the correct answer when he or she clicks on any element of the atom image.

Algorithms

Proposed structure for the DNA-binding domain of the helix-loop-helix family of eukaryotic gene regulatory proteins.

A modelled tertiary structure for the dimeric HLH domain of the E47 protein is presented. Structural information was obtained from the aligned sequences of > 40 members of the HLH family. The information was used to model each monomer as an alpha-helical hairpin, with knobs-into-holes packing of side-chains as found in antiparallel coiled-coil. The dimer forms a four-helix bundle with additional knobs-into-holes packing at the dimer interface. The size and electrostatic properties of core-forming residues are all accounted for in the model. The model does not violate any known properties of protein structure. The monomers are related by two-fold rotational symmetry, in agreement with the observed DNA-binding sites which are imperfect inverted repeats. The N-terminal basic region, in which DNA binding and base specificity reside, forms the first part of helix 1. A prediction based on the model structure is that the HLH domains do not bind to DNA in its B form but require a partially unwound conformation in order to enter the major groove.

Amino Acid Sequence

Electrophoretic behavior of d(GGAAAAAAGG)n, d(CCAAAAAACC)n, and (CCAAAAAAGG)n and implications for a DNA bending model.

Double stranded multimers (C2A6C2)n, (C2A6G2)n and (G2A6G2)n were prepared from chemically synthesized oligonucleotides to study the influence of sequences flanking the An tract on the curvature of DNA. All these duplexes, including polypurine.polypyrimidine one, exhibit strong retardation in polyacrylamide gel which is indicative of pronounced DNA curvature. It has been proposed previously that among the bends at the boundary with the oligo(A) tract two types should be distinguished: 5'-bends and 3'-bends (Koo et al., 1986) This distinction was deduced from different relative mobilities of two specially designed sequences having phased 5'-bends and 3'-bends, respectively. Our data indicate that the substitutions of nucleotides at both 5' and 3' boundaries of A6 tract result in comparable changes in relative mobility. Therefore, for B-B' bends it is important to take into account not only whether they are at the 5' or 3' end of an oligo(dA) tract, but also the particular sequences at the boundaries of this tract.

Base Sequence

Sequence dependent modulating effect of camptothecin on the DNA-cleaving activity of the calf thymus type I topoisomerase.

High-resolution mapping of topol cleavages in the regions of human DNA including the oncogene c-Ha-ras and p53, has revealed three kinds of topol cleavage sites: cleavage sites not affected by camptothecin; cleavage sites reinforced only in the presence of camptothecin, and cleavage sites which weaken in the presence of camptothecin. Statistical analysis of sequences revealed certain nucleotide or dinucleotide preferences for three groups studied. The preferences in camptothecin-reduced sites predominate upstream from the cleavage point, whereas in camptothecin-induced sites the situation is reversed. The influence of camptothecin on cleavage sites induced by two molecular forms of topol has been also studied.

Animals

Structure of the hydration shells of oligo(dA-dT).oligo(dA-dT) and oligo(dA).oligo(dT) tracts in B-type conformation on the basis of Monte Carlo calculations.

Monte Carlo simulations [(N, V, T)-ensemble] were performed for the hydration shell of poly(dA-dT).poly(dA-dT) in canonical B form and for the hydration shell of poly(dA).poly(dT) in canonical B conformation and in a conformation with narrow minor groove, highly inclined bases, but with a nearly zero-inclined base pair plane (B' conformation). We introduced helical periodic boundary conditions with a rather small unit cell and a limited number of water molecules to reduce the dimensionality of the configuration space. The coordinates of local maxima of water density and the properties of one- and two-membered water bridges between polar groups of the DNA were obtained. The AT-alternating duplex hydration mirrors the dyad symmetry of polar group distribution. At the dApdT step, a water bridge between the two carbonyl oxygens O2 of thymines is formed as in the central base-pair step of Dickerson's dodecamer. In the major groove, 5-membered water chains along the tetranucleotide pattern d(TATA).d(TATA) are observed. The hydration geometry of poly(dA).poly(dT) in canonical B conformation is distinguished by autonomous primary hydration of the base-pair edges in both grooves. When this polymer adopts a conformation with highly inclined bases and narrow minor groove, the water density distribution in the minor groove is in excellent agreement with Dickerson's spine model. One local maximum per base pair of the first layer is located near the dyad axis between adjacent base pairs, and one local maximum per base pair in the second shell lies near the dyad axis of the base pair itself. The water bridge between the two strands formed within the first layer was observed with high probability. But the water molecules of the second layer do not have a statistically favored orientation necessary for bridging first layer waters. In the major groove, the hydration geometry of the (A.T) base-pair edge resembles the main features of the AT-pair hydration derived from other sequences for the canonical B form. The preference of the B' conformation for oligo(dA).oligo(dT) tracts may express the tendency to common hydration of base-pair edges of successive base pairs in the grooves of B-type DNA. The mean potential energy of hydration of canonical B-DNA was estimated to be -60 to -80 kJ/mole nucleotides in dependence on the (G.C) contents. Because of the small system size, this estimation is preliminary.

Base Sequence

An automatic search for similar spatial arrangements of alpha-helices and beta-strands in globular proteins.

A fast search algorithm to reveal similar polypeptide backbone structural motifs in proteins is proposed. It is based on the vector representation of a polypeptide chain fold in which the elements of regular secondary structures are approximated by linear segments (Abagyan and Maiorov, J. Biomol. Struct. Dyn. 5, 1267-1279 (1988)). The algorithm permits insertions and deletions in the polypeptide chain fragments to be compared. The fast search algorithm implemented in FASEAR program is used for collecting beta alpha beta supersecondary structure units in a number of alpha/beta proteins of Brookhaven Data Bank. Variation of geometrical parameters specifying backbone chain fold is estimated. It appears that the conformation of the majority of the fragments, although almost all of them are right-handed, is quite different from that of standard beta alpha beta units. Apart from searching for specific type of secondary structure motif, the algorithm allows automatically to identify new recurrent folding patterns in proteins. It may be of particular interest for the development of tertiary template approach for prediction of protein three-dimensional structure as well for constructing artificial polypeptides with goal-oriented conformation.

Algorithms

New methodology for computer-aided modelling of biomolecular structure and dynamics. 1. Non-cyclic structures.

A general methodology is proposed for the conformational modelling of biomolecular systems. The approach allows one: (i) to describe the system under investigation by an arbitrary set of internal variables, i.e., torsion angles, bond angles, and bond lengths; it offers a possibility to pass from the free structure to a completely fixed one with the number of variables from 3N to zero, respectively, where N is the number of atoms; (ii) to consider both, a single molecule and a complex of many molecules, (e.g., proteins, water, ligands, etc.) in terms of one universal model; (iii) to study the dynamics of the system using explicit analytical Lagrangian equations of motion, thus opening up possibilities for investigations of slow concerted motions such as domain oscillations in proteins etc.; (iv) to calculate the partial derivatives of various functions of conformation, e.g., the conformational energy or external constraints imposed, using a standard efficient procedure regardless of the variables and the structure of the system. The approach is meant to be used in various investigations concerning the conformations and dynamics of biomacromolecules.

Computer Simulation

New methodology for computer-aided modelling of biomolecular structure and dynamics. 2. Local deformations and cycles.

A new methodology for the conformational modelling of biomolecular systems (1) is extended to local deformations of chain molecules and to flexible molecular rings. It is shown that these two cases may be reduced to considering an equivalent molecular model with a regular tree-like topology. A simple procedure is developed to analyze any flexible rings (the five- and six-membered sugar rings of carbohydrates and nucleic acids, in particular) and local deformation regions by energy minimization. Dynamic equations are also derived for such molecular systems. As a result, a unified approach is proposed for the efficient energy minimization and simulation of dynamic behavior of multimolecular systems having any set of variable internal coordinates, local deformation regions and cycles. Advantages and domains of applicability of the approach are discussed.

Computer Simulation

A simple qualitative representation of polypeptide chain folds: comparison of protein tertiary structures.

A new simple quantitative representation of three-dimensional structure of globular proteins is proposed which is useful for comparison of distantly related problems, computer sorting of large sets of conformations, and search of structurally similar domains in protein data base. The folding course of the polypeptide backbone is approximated by a set of successive vectors corresponding to the elements of regular secondary structure (e.g. alpha-helices, strands of beta-sheets) and non-regular segments. The parameters specifying the spatial organization of segments in this vector model are internal coordinates, namely, lengths of the vectors, planar and dihedral angles. Quantitative representation proposed allows to circumvent the problem of insertions/deletions and to avoid the stage of best superposition during protein comparison. An application was made to the comparison of three-dimensional structures of scorpion toxins Centruroides sculpturatus Ewing v-3, Buthus eupeus M9 and I5A, which have different chain lengths and low sequence similarity.

Models, Molecular

Structural basis of stable bending in DNA containing An tracts. Different types of bending.

Structural determinants of DNA bending of different types have been studied by theoretical conformational analysis of duplexes. Their terminal parts were fixed either in an ordinary low-energy B-like conformation or in "anomalous" conformations with a narrowed minor groove typical of An tracts. The anomalous conformations had different negative tilt angles (up to about zero), different propeller twists and minor groove widths. Calculations have been performed for DNA fragments AnTm, TnAm, AnGCTm, AnCGTm, TmGCAn, TmCGAn which are the models of the junction of two anomalous structures on An and Tm tracts. On the AT step of the AnTm fragment the minor groove can be easily narrowed so that a whole unbent fragment of anomalous structure is formed on AnTm. According to our energy estimates, there should not be any reliable bending on AnTm. In contrast, in all other cases there was a pronounced roll-like bending into the major groove in the chemical symmetry region. Calculations of the junction between the anomalous and ordinary B-like structure for GnTm and CnAm have shown that there is an equilibrium bending with a tilt component towards the chain having the anomalous structure at the 5'-end. From our calculations it is impossible to determine precisely the direction of bending, though it can be suggested that the roll component of bending might be directed towards the major groove. The anomalous structure is the main reason of bending; alternations of pyrimidines and purines can modulate the value and the direction of equilibrium bending (only the value in the case of self-complementary fragments).(ABSTRACT TRUNCATED AT 250 WORDS)

Base Composition

[Two types of tripeptide conformation in collagen. Calculation of the structure of (Gly-Pro-Ser)n and (Gly-Val-Hyp)n polytripeptides].

Conformational analysis of polypeptides (Gly-Pro-Ser)n and (Gly-Val-Hyp)n was carried out for collagen-like triple helical complexes (coiled coils with screw symmetry). The lowest energy structure of the first polymer (helical parameters t 52,8, h 0,282 nm) is very close to that of (Gly-Pro-Hyp)n. The hydroxyl group of a serine residue does not form any intramolecular hydrogen bonds in this structure. (Gly-Val-Hyp)n triple complex is shown to unwind to t 7,7, h 0,297 nm as a result of optimization procedure. These findings confirm the assumption, made earlier on the basis of conformational analysis of (Gly-Pro-Hyp)n, (Gly-Pro-Ala)n, (Gly-Ala-Hyp)n, (Gly-Ala-Ala)n, that the collagen triple helix contains stable wound triplets with proline in the second position, while the absence of imino acid in the 2nd position facilitates the unwinding of the triple helix. Thus, a collagen helix appears to have different parameters for the sites differing in the amino acid sequence. The values measured in the X-ray experiments (h 0,29 nm, t' 36) should be considered as a result of averaging. The model allows to reconcile the X-ray data for collagen and crystalline (Gly-Pro-Pro)10 oligomer.

Collagen