PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein structure”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

How inaccuracies in protein structure models affect estimates of protein-ligand interactions: computational analysis of HIV-I protease inhibitor binding.

The influence of possible inaccuracies that can arise during homology modeling of protein structures used for ligand binding studies were investigated with the molecular mechanics generalized Born surface area (MM-GBSA) method. For this, a family of well-characterized HIV-I protease-inhibitor complexes was used. Validation of MM-GBSA led to a correlation coefficient ranging from 0.72 to 0.93 between calculated and experimental binding free energies DeltaG. All calculated DeltaG values were based on molecular dynamics simulations with explicit solvent. Errors introduced into the protein structure through misplacement of side-chains during rotamer modeling led to a correlation coefficient between DeltaG(calc) and DeltaG(exp) of 0.75 compared with 0.90 for the correctly placed side chains. This is in contrast to homology models for members of the retroviral protease family with template structures ranging in sequence identity between 32% and 51%. For these protein models, the correlation coefficients vary between 0.84 and 0.87, which is considerably closer to the original protein (0.90). It is concluded that HIV-I low sequence identity with the template structure still allows creating sufficiently reliable homology models to be used for ligand-binding studies, although placement of the rotamers is a critical step during the modeling.

Amino Acid Sequence↗

Protein structures sustain evolutionary drift.

A protein sequence folds into a unique three-dimensional protein structure. Different sequences, though, can fold into similar structures. How stable is a protein structure with respect to sequence changes? What percentage of the sequence is 'anchor' residues, that is, residues crucial for protein structure and function? Here, answers to these questions are pursued by analyzing large numbers of structurally homologous protein pairs. Most pairs of similar structures have sequence identity as low as expected from randomly related sequences (8-9%). On average, only 3-4% of all residues are 'anchor' residues. The symmetric shape of the distribution at low sequence identity suggests that for most structures, four billion years of evolution was sufficient to reach an equilibrium. The mean identities for convergent (different ancestor) and divergent (same ancestor) evolution of proteins to similar structures are quite close and hence, in most cases, it is difficult to distinguish between the two effects. In particular, low levels of sequence identity appear not to be indicative of convergent evolution.

Bias↗

Constraint Logic Programming approach to protein structure prediction.

BACKGROUND: The protein structure prediction problem is one of the most challenging problems in biological sciences. Many approaches have been proposed using database information and/or simplified protein models. The protein structure prediction problem can be cast in the form of an optimization problem. Notwithstanding its importance, the problem has very seldom been tackled by Constraint Logic Programming, a declarative programming paradigm suitable for solving combinatorial optimization problems. RESULTS: Constraint Logic Programming techniques have been applied to the protein structure prediction problem on the face-centered cube lattice model. Molecular dynamics techniques, endowed with the notion of constraint, have been also exploited. Even using a very simplified model, Constraint Logic Programming on the face-centered cube lattice model allowed us to obtain acceptable results for a few small proteins. As a test implementation their (known) secondary structure and the presence of disulfide bridges are used as constraints. Simplified structures obtained in this way have been converted to all atom models with plausible structure. Results have been compared with a similar approach using a well-established technique as molecular dynamics. CONCLUSIONS: The results obtained on small proteins show that Constraint Logic Programming techniques can be employed for studying protein simplified models, which can be converted into realistic all atom models. The advantage of Constraint Logic Programming over other, much more explored, methodologies, resides in the rapid software prototyping, in the easy way of encoding heuristics, and in exploiting all the advances made in this research area, e.g. in constraint propagation and its use for pruning the huge search space.

Computational Biology↗

Use of glycerol, polyols and other protein structure stabilizing agents in protein crystallization.

A protein preparation to be used for crystallization should be homogeneous and should remain so throughout the course of a prolonged crystallization experiment. General methods for preparation of pure proteins and for prevention of their covalent modification (through proteolysis, sulfhydryl oxidation, etc.) during prolonged incubation are well known. Crystallographers are less aware of general methods for stabilization of proteins against non-covalent modifications (partial denaturation, heterogeneous aggregation) which can also introduce structural heterogeneity into a protein preparation. Related to this issue are methods to suppress protein conformational flexibility which can be a source of dynamic structural heterogeneity and which presents an entropic barrier to crystallization. However, for many years agents which stabilize protein structure have been described in the biochemical literature. Recently the most widely used of these structure-stabilizing agents, glycerol, was used to crystallize T7 RNA polymerase. The observation that this compound has general structure-stabilizing effects and that it was essential for crystallization of at least this one protein led to the suggestion that it might be generally useful in crystallizing flexible proteins and in inducing order in disordered segments of crystalline proteins. Subsequently, glycerol was used with good effect in the crystallization of a number of proteins. Other recent results suggest that soaking crystals in solutions containing glycerol can have 'structure-ordering' effects on the crystalline protein. These observations support the utility of glycerol in protein crystallization and suggest that the information in the biochemical literature on protein structure-stabilizing agents will find useful application in the field of protein crystal growth.

Journal Article↗

Evaluation of structural similarity based on reduced dimensionality representations of protein structure.

Protein similarity estimations can be achieved using reduced dimensional representations and we describe a new application for the generation of two-dimensional maps from the three-dimensional structure. The code for the dimensionality reduction is based on the concept of pseudo-random generation of two-dimensional coordinates and Monte Carlo-like acceptance criteria for the generated coordinates. A new method for calculating protein similarity is developed by introducing a distance-dependent similarity field. Similarity of two proteins is derived from similarity field indices between amino acids based on various criteria such as hydrophobicity, residue replacement factors and conformational similarity, each showing a one factor Gaussian dependence. Results on comparisons of misfolded protein models with data sets of correctly folded structures show that discrimination between correctly folded and misfolded structures is possible. Tests were carried out on five different proteins, comparing a misfolded protein structure with members of the same topology, architecture, family and domain according to the CATH classification.

Computational Biology↗

Bioinformatics methods to predict protein structure and function. A practical approach.

Protein structure prediction by using bioinformatics can involve sequence similarity searches, multiple sequence alignments, identification and characterization of domains, secondary structure prediction, solvent accessibility prediction, automatic protein fold recognition, constructing three-dimensional models to atomic detail, and model validation. Not all protein structure prediction projects involve the use of all these techniques. A central part of a typical protein structure prediction is the identification of a suitable structural target from which to extrapolate three-dimensional information for a query sequence. The way in which this is done defines three types of projects. The first involves the use of standard and well-understood techniques. If a structural template remains elusive, a second approach using nontrivial methods is required. If a target fold cannot be reliably identified because inconsistent results have been obtained from nontrivial data analyses, the project falls into the third type of project and will be virtually impossible to complete with any degree of reliability. In this article, a set of protocols to predict protein structure from sequence is presented and distinctions among the three types of project are given. These methods, if used appropriately, can provide valuable indicators of protein structure and function.

Algorithms↗

Automated construction of structural motifs for predicting functional sites on protein structures.

Structural genomics initiatives are beginning to rapidly generate vast numbers of protein structures. For many of the structures, functions are not yet determined and high-throughput methods for determining function are necessary. Although there has been extensive work in function prediction at the sequence level, predicting function at the structure level may provide better sensitivity and predictive value. We describe a method to predict functional sites by automatically creating three dimensional structural motifs from amino acid sequence motifs. These structural motifs perform comparably well with manually generated structural motifs and perform better than sequence motifs. Automatically generated structural motifs can be used for structural-genomic scale function prediction on protein structures.

Amino Acid Motifs↗

A simple method for protein structural classification.

Since the concept of structural classes of proteins was proposed, the problem of protein classification has been tackled by many groups. Most of their classification criteria are based only on the helix/strand contents of proteins. In this paper, we proposed a method for protein structural classification based on their secondary structure sequences. It is a classification scheme that can confirm existing classifications. Here a mathematical model is constructed to describe protein secondary structure sequences, in which each protein secondary structure sequence corresponds to a transition probability matrix that characterizes and differentiates protein structure numerically. Its application to a set of real data has indicated that our method can classify protein structures correctly. The final classification result is shown schematically. So it is visual to observe the structural classifications, which is different from traditional methods.

Algorithms↗

The correlation of protein structure and evolution of a protein-coding gene: phylogenetic inference using cytochrome oxidase III.

Discriminating phylogenetic signal from noise in DNA sequence data is a difficult problem in phylogenetic inference at higher systematic levels. For protein-coding genes, noise at synonymous (silent) positions can be filtered by deleting entire codon positions or types of change at a codon position. This method is not appropriate for replacement sites, because changes at each site within a codon may not be independent. This research presents a method using information from protein structure to evaluate variation in replacement sites. Analysis of the correlation of amino acid variation with protein structure identified rapidly evolving codons in the COIII gene. In a series of phylogenetic analyses attempting to recover a known set of vertebrate relationships, downweighting these labile codons produced the most accurate results. Structural correlates of variable and invariant residues identified in this study can be used to increase the accuracy of models used for phylogenetic inference. Viewing amino acid variation within a phylogenetic framework provided insight into residue changes important in the secondary and tertiary structures of the molecule, changes that were correlated between pairs of neighboring residues or between residues in neighboring helices.

Amino Acid Substitution↗

Progress in protein structure prediction.

If protein structure prediction methods are to make any impact on the impending onerous task of analyzing the large numbers of unknown protein sequences generated by the ongoing genome-sequencing projects, it is vital that they make the difficult transition from computational 'gedankenexperiments' to practical software tools. This has already happened in the field of comparative modelling and is currently happening in the threading field. Unfortunately, there is little evidence of this transition happening in the field of ab initio tertiary-structure prediction.

Algorithms↗

The expert system approach to predicting protein structure.

Prediction of protein structure is an open-ended problem. Since an approach from first principles cannot be taken in reasonable computer time, short-cuts using further data are necessary. Such data include information about the specific protein in question, and information in databases which are about proteins in general. Is it possible to write a general, flexible 'superalgorithm' which would suit most circumstances? If so, it would seem likely to overcome one of the most understated but nonetheless greatest difficulties associated with molecular modelling and computer-aided drug design--reproducibility. To this end, a 'polymorphic programming environment' has been developed which represents both an expert system and a high-level language for theoretical chemists and molecular biologists. This language is GLOBAL (Ball et al., 1990). In a series of earlier studies, and more recently by means of GLOBAL itself, the nature of reproducibility and its rather surprising limits have been explored, and in general the current status and future potential of protein modelling have been examined.

Expert Systems↗

Molecular confinement influences protein structure and enhances thermal protein stability.

The sol-gel method of encapsulating proteins in a silica matrix was investigated as a potential experimental system for testing the effects of molecular confinement on the structure and stability of proteins. We demonstrate that silica entrapment (1) is fully compatible with structure analysis by circular dichroism, (2) allows conformational studies in contact with solvents that would otherwise promote aggregation in solution, and (3) generally enhances thermal protein stability. Lysozyme, alpha-lactalbumin, and metmyoglobin retained native-like solution structures following sol-gel encapsulation, but apomyoglobin was found to be largely unfolded within the silica matrix under control buffer conditions. The secondary structure of encapsulated apomyoglobin was unaltered by changes in pH and ionic strength of KCl. Intriguingly, the addition of other neutral salts resulted in an increase in the alpha-helical content of encapsulated apomyoglobin in accordance with the Hofmeister ion series. We hypothesize that protein conformation is influenced directly by the properties of confined water in the pores of the silica. Further work is needed to differentiate the steric effects of the silica matrix from the solvent effects of confined water on protein structure and to determine the extent to which this experimental system mimics the effects of crowding and confinement on the function of macromolecules in vivo.

Animals↗

Prediction of protein structure by evaluation of sequence-structure fitness. Aligning sequences to contact profiles derived from three-dimensional structures.

The problem of protein structure prediction is formulated here as that of evaluating how well an amino acid sequence fits a hypothetical structure. The simplest and most complicated approaches, secondary structure prediction and all-atom free energy calculations, can be viewed as sequence-structure fitness problems. Here, an approach of intermediate complexity is described, which involves; (1) description of a protein structure in terms of contact interface vectors, with both intra-protein and protein-solvent contacts counted, (2) derivation of sequence preferences for 2 up to 29 contact interface types, (3) generation of numerous hypothetical model structures by placing the input sequence into a large set of known three-dimensional structures in all possible alignments, (4) evaluation of these models by summing the sequence preferences over all structural positions and (5) choice of predicted three-dimensional structure as that with the best sequence-structure fitness. Evolutionary information is incorporated by using position-dependent core weights derived from multiple sequence alignments. A number of tests of the method are performed: (1) evaluation of cyclic shifts of a sequence in its native structure; (2) alignment of a sequence in its native structure, allowing gaps; (3) alignment search with a sequence or sequence fragment in a database of structures; and (4) alignment search with a structure in a database of sequences. The main results are: (1) a native sequence can very well find its native structure among a large number of alternatives, in correct alignment; (2) substructures, such as (beta alpha)n units, can be detected in spite of very low sequence similarity; (3) remote homologous can be detected, with some dependence on the set of parameters used; (4) contact interface parameters are clearly superior to classical secondary structure parameters; (5) a simple interface description in terms of just two states, protein-protein and protein-water contacts, performs surprisingly well; (6) the use of core weights considerably improves accuracy in detection of remote homologues; (7) based on a sequence database search with a myoglobin contact profile, the C-terminal domain of a viral origin of replication binding protein is predicted to have an all-helical fold. The sequence-structure fitness concept is sufficiently general to accommodate a large variety of protein structure prediction methods, including new models of intermediate complexity currently being developed.

Amino Acid Sequence↗

Correlation of serotype specificity and protein structure of the five U.S. serotypes of bluetongue virus.

The relationship between serotype specificity and protein structure was studied by polyacrylamide gel electrophoresis, peptide mapping and radioimmune precipitation (RIP) of structural and non-structural proteins of the five U.S. serotypes of bluetongue virus (BTV). The surface proteins, VP2 and VP5, showed the most variation in size among the serotypes. Peptide mapping of the proteins showed that VP2 is unique for each of the U.S. serotypes. The nucleocapsid and non-structural proteins showed a high degree of conservation, whereas the other surface protein, VP5, showed intermediate conservation among the serotypes. Monospecific neutralizing antiserum produced in rabbits against each serotype was used in cross-RIP against cytoplasmic extracts prepared from cells infected with each BTV serotype. There were extensive cross-reactions among those proteins which showed a high degree of structural conservation, whereas VP2 was immunoprecipitated best in the homologous RIP system. Thus, a correlation between serotype specificity and protein structure was shown among the five U.S. serotypes of BTV.

Antigens, Viral↗

Link protein cDNA sequence reveals a tandemly repeated protein structure.

Link protein stabilizes the cartilage proteoglycan/hyaluronic acid aggregate by binding to both components. We screened a cDNA library prepared from rat chondrosarcoma mRNA in the lambda gt11 expression vector with monoclonal antibodies and polyclonal antisera to link protein. We obtained a clone for two-thirds of the link protein cDNA and identified it based on its deduced amino acid sequence. There are four RNA transcripts for link protein, ranging from 1.5 to 5.5 kilobases in size. The deduced amino acid sequence for link protein shows two domains of 100 residues each, which share 44% homology; within each domain is a 19-residue stretch 74% homologous with its counterpart. This structure indicates that link protein may have been formed by gene duplication.

Amino Acid Sequence↗

Simulations of apo and holo-fatty acid binding protein: structure and dynamics of protein, ligand and internal water.

Two molecular dynamics simulations of 5 ns each have been carried out for rat intestinal fatty acid binding protein, in apo-form and with bound palmitate. The fatty acid and a number of water molecules are encapsulated in a large interior cavity of the barrel-shaped protein. The simulations are compared to experimental data and analyzed in terms of root mean square deviations, atomic B-factors, secondary structure elements, hydrogen bond patterns, and distance constraints derived from nuclear Overhauser experiments. Excellent agreement is found between simulated and experimental solution structures of holo-FABP, but a number of differences are observed for the apo-form. The ligand in holo-FABP shows considerable displacement after about 1.5 ns and displays significant configurational entropy. A novel computational approach has been employed to identify internal water and to capture exchange pathways. Orifices in the portal and gap regions of the protein, discussed in the experimental literature, have been confirmed as major openings for solvent exchange between the internal cavity and bulk water. A third opening on the opposite side of the barrel experiences significant exchange but it does not provide a pathway for further passage to the central cavity. Internal water is characterized in terms of density distributions, interaction energies, mobility, protein contact times, and water molecule coordination. A number of differences are observed between the apo and holo-forms and related to differences in the protein structure. Solvent inside apo-FABP, for example, shows characteristics of a water droplet, while solvent in holo-FABP benefits from interactions with the ligand headgroup and slightly stronger interactions with protein residues.

Animals↗