PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein structure”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Comparison of crystal structures of two homologous proteins: structural origin of altered domain interactions in immunoglobulin light-chain dimers.

The sequence and structure of a second human kappa 1 immunoglobulin light-chain variable domain, Wat, has been determined. The R-factor is 15.7% for 1.9-A data. One hundred and ninety-five water molecules were identified; 30 water molecules were located in identical positions in each of the monomers. Some of the water molecules are integral parts of the domains. This light chain is encoded by the same variable domain gene that encoded the previously characterized kappa I variable domain, Rei. Due to limited somatic mutation, the two highly homologous proteins differ in only 20 of the 108 residues. Wat crystallized in space group P6(4) while Rei crystallized in space group P6(1); in both crystals, the asymmetric unit was the noncovalent dimer. Although the basic domain structure is the same for both proteins, the relative positions of the domains within the two dimers differ. This difference is most likely accounted for by the replacement of Tyr36 in Rei by Phe in the Wat protein. Residue Tyr36 is part of the hydrogen-bonding network in the interface between the domains in Rei. Losing the hydrogen-bonding capability of residue 36 by replacement of Tyr by Phe alters the network of hydrogen bonds between the domains, resulting in a different domain-domain contact. The details of lattice contacts in the two crystals were compared. One type of contact that extends the beta-sheet of the individual domains was conserved, but because it involved different symmetry elements within the crystal, different crystal packing resulted. In the Wat crystal, one of the contacts shows an example of how a symmetrical binding site can "bind" an asymmetrical object. Further, the examination of the Wat crystal also illustrates how the different crystalline environments of the domains of the dimer results in different distributions of temperature factors for the residues within the domains.

Amino Acid Sequence↗

Datamining protein structure databanks for crystallization patterns of proteins.

A study of 345 protein structures selected among 1,500 structures determined by nuclear magnetic resonance (NMR) methods, revealed useful correlations between crystallization properties and several parameters for the studied proteins. NMR methods of structure determination do not require the growth of protein crystals, and hence allow comparison of properties of proteins that have or have not been the subject of crystallographic approaches. One- and two-dimensional statistical analyses of the data confirmed a hypothesized relation between the size of the molecule and its crystallization potential. Furthermore, two-dimensional Bayesian analysis revealed a significant relationship between relative ratio of different secondary structures and the likelihood of success for crystallization trials. The most immediate result is an apparent correlation of crystallization potential with protein size. Further analysis of the data revealed a relationship between the unstructured fraction of proteins and the success of its crystallization. Utilization of Bayesian analysis on the latter correlation resulted in a prediction performance of about 64%, whereas a two-dimensional Bayesian analysis succeeded with a performance of about 75%.

Bayes Theorem↗

Protein structure modeling in the proteomics era.

Structural proteomics aims to understand the structural basis of protein interactions and functions. A prerequisite for this is the availability of 3D protein structures that mediate the biochemical interactions. The explosion in the number of available gene sequences set the stage for the next step in genome-scale projects -- to obtain 3D structures for each protein. To achieve this ambitious goal, the slow and costly structure determination experiments are supplemented with theoretical approaches. The current state and recent advances in structure modeling approaches are reviewed here, with special emphasis on comparative protein structure modeling techniques.

Animals↗

Accurate classification of protein structural families using coherent subgraph analysis.

Protein structural annotation and classification is an important problem in bioinformatics. We report on the development of an efficient subgraph mining technique and its application to finding characteristic substructural patterns within protein structural families. In our method, protein structures are represented by graphs where the nodes are residues and the edges connect residues found within certain distance from each other. Application of subgraph mining to proteins is challenging for a number reasons: (1) protein graphs are large and complex, (2) current protein databases are large and continue to grow rapidly, and (3) only a small fraction of the frequent subgraphs among the huge pool of all possible subgraphs could be significant in the context of protein classification. To address these challenges, we have developed an information theoretic model called coherent subgraph mining. From information theory, the entropy of a random variable X measures the information content carried by X and the Mutual Information (MI) between two random variables X and Y measures the correlation between X and Y. We define a subgraph X as coherent if it is strongly correlated with every sufficiently large sub-subgraph Y embedded in it. Based on the MI metric, we have designed a search scheme that only reports coherent subgraphs. To determine the significance of coherent protein subgraphs, we have conducted an experimental study in which all coherent subgraphs were identified in several protein structural families annotated in the SCOP database (Murzin et al, 1995). The Support Vector Machine algorithm was used to classify proteins from different families under the binary classification scheme. We find that this approach identifies spatial motifs unique to individual SCOP families and affords excellent discrimination between families.

Algorithms↗

Recoverable one-dimensional encoding of three-dimensional protein structures.

One-dimensional (1D) structures of proteins such as secondary structure and contact number provide intuitive pictures to understand how the native three-dimensional (3D) structure of a protein is encoded in the amino acid sequence. However, it is still not clear whether a given set of 1D structures contains sufficient information for recovering the underlying 3D structure. Here we show that the 3D structure of a protein can be recovered from a set of three types of 1D structures, namely, secondary structure, contact number and residue-wise contact order which is introduced here for the first time. Using simulated annealing molecular dynamics simulations, the structures satisfying the given native 1D structural restraints were sought for 16 proteins of various structural classes and of sizes ranging from 56 to 146 residues. By selecting the structures best satisfying the restraints, all the proteins showed a coordinate RMS deviation of <4 A from the native structure, and, for most of them, the deviation was even <2 A. The present result opens a new possibility to protein structure prediction and our understanding of the sequence-structure relationship.

Amino Acid Sequence↗

Protein structure and function at low temperatures.

Proteins represent the major components in the living cell that provide the whole repertoire of constituents of cellular organization and metabolism. In the process of evolution, adaptation to extreme conditions mainly referred to temperature, pH and low water activity. With respect to life at low temperatures, effects on protein structure, protein stability and protein folding need consideration. The sequences and topologies of proteins from psychrophilic, mesophilic and thermophilic organisms are found to be highly homologous. Commonly, adaptive changes refer to multiple alterations of the amino acid sequence, which presently cannot be correlated with specific changes of structure and stability; so far it has not been possible to attribute specific increments in the free energy of stabilization to well-defined amino-acid exchanges in an unambiguous way. The stability of proteins is limited at high and low temperatures. Their expression and self-organization may be accomplished under conditions strongly deviating from optimum growth conditions. Molecular adaptation to extremes of temperature seems to be accompanied by a flattening of the temperature profile of the free energy of stabilization. In principle, the free energy of stabilization of proteins is small compared to the total molecular energy. As a consequence, molecular adaptation to extremes of physical conditions only requires marginal alterations of the intermolecular interactions and packing density. Careful statistical and structural analyses indicate that altering the number of ion pairs and hydrophobic interactions allows the flexibility of proteins to be adjusted so that full catalytic function is maintained at varying temperatures.

Drug Stability↗

Small libraries of protein fragments model native protein structures accurately.

Prediction of protein structure depends on the accuracy and complexity of the models used. Here, we represent the polypeptide chain by a sequence of rigid fragments that are concatenated without any degrees of freedom. Fragments chosen from a library of representative fragments are fit to the native structure using a greedy build-up method. This gives a one-dimensional representation of native protein three-dimensional structure whose quality depends on the nature of the library. We use a novel clustering method to construct libraries that differ in the fragment length (four to seven residues) and number of representative fragments they contain (25-300). Each library is characterized by the quality of fit (accuracy) and the number of allowed states per residue (complexity). We find that the accuracy depends on the complexity and varies from 2.9A for a 2.7-state model on the basis of fragments of length 7-0.76A for a 15-state model on the basis of fragments of length 5. Our goal is to find representations that are both accurate and economical (low complexity). The models defined here are substantially better in this regard: with ten states per residue we approximate native protein structure to 1A compared to over 20 states per residue needed previously. For the same complexity, we find that longer fragments provide better fits. Unfortunately, libraries of longer fragments must be much larger (for ten states per residue, a seven-residue library is 100 times larger than a five-residue library). As the number of known protein native structures increases, it will be possible to construct larger libraries to better exploit this correlation between neighboring residues. Our fragment libraries, which offer a wide range of optimal fragments suited to different accuracies of fit, may prove to be useful for generating better decoy sets for ab initio protein folding and for generating accurate loop conformations in homology modeling.

Models, Molecular↗

A theoretical search for folding/unfolding nuclei in three-dimensional protein structures.

When a protein folds or unfolds, it has to pass through many half-folded microstates. Only a few of them can be seen experimentally. In a two-state transition proceeding with no accumulation of metastable intermediates [Fersht, A. R. (1995) Curr. Opin. Struct. Biol. 5, 79-84], only the semifolded microstates corresponding to the transition state can be outlined; they influence the folding/unfolding kinetics. Our aim is to calculate them, provided the three-dimensional protein structure is given. The presented approach follows from the capillarity theory of protein folding and unfolding [Wolynes, P. G. (1997) Proc. Natl. Acad. Sci. USA 94, 6170-6175]. The approach is based on a search for free-energy saddle point(s) on a network of protein unfolding pathways. Under some approximations, this search is rapidly performed by dynamic programming and, despite its relative simplicity, gives a good correlation with experiment. The computed folding nuclei look like ensembles of those compact and closely packed parts of the three-dimensional native folds that contain a small number of disordered protruding loops. Their estimated free energy is consistent with the rapid (within seconds) folding and unfolding of small proteins at the point of thermodynamic equilibrium between the native fold and the coil.

Protein Conformation↗

PROSPECT-PSPP: an automatic computational pipeline for protein structure prediction.

Knowledge of the detailed structure of a protein is crucial to our understanding of the biological functions of that protein. The gap between the number of solved protein structures and the number of protein sequences continues to widen rapidly in the post-genomics era due to long and expensive processes for solving structures experimentally. Computational prediction of structures from amino acid sequence has come to play a key role in narrowing the gap and has been successful in providing useful information for the biological research community. We have developed a prediction pipeline, PROSPECT-PSPP, an integration of multiple computational tools, for fully automated protein structure prediction. The pipeline consists of tools for (i) preprocessing of protein sequences, which includes signal peptide prediction, protein type prediction (membrane or soluble) and protein domain partition, (ii) secondary structure prediction, (iii) fold recognition and (iv) atomic structural model generation. The centerpiece of the pipeline is our threading-based program PROSPECT. The pipeline is implemented using SOAP (Simple Object Access Protocol), which makes it easier to share our tools and resources. The pipeline has an easy-to-use user interface and is implemented on a 64-node dual processor Linux cluster. It can be used for genome-scale protein structure prediction. The pipeline is accessible at http://csbl.bmb.uga.edu/protein_pipeline.

Computational Biology↗

The elastic net algorithm and protein structure prediction.

Predicting protein structures from their amino acid sequences is a problem of global optimization. Global optima (native structures) are often sought using stochastic sampling methods such as Monte Carlo or molecular dynamics, but these methods are slow. In contrast, there are fast deterministic methods that find near-optimal solutions of well-known global optimization problems such as the traveling salesman problem (TSP). But fast TSP strategies have yet to be applied to protein folding, because of fundamental differences in the two types of problems. Here, we show how protein folding can be framed in terms of the TSP, to which we apply a variation of the Durbin-Willshaw elastic net optimization strategy. We illustrate using a simple model of proteins with database-derived statistical potentials and predicted secondary structure restraints. This optimization strategy can be applied to many different models and potential functions, and can readily incorporate experimental restraint information. It is also fast; with the simple model used here, the method finds structures that are within 5-6 A all-Calpha-atom RMSD of the known native structures for 40-mers in about 8 s on a PC; 100-mers take about 20 s. The computer time tau scales as tau approximately n, where n is the number of amino acids. This method may prove to be useful for structure refinement and prediction.

Algorithms↗

The muscle-derived lens of a squid bioluminescent organ is biochemically convergent with the ocular lens. Evidence for recruitment of aldehyde dehydrogenase as a predominant structural protein.

Many of the structural proteins of ocular lenses, commonly referred to as crystallins, are identical to specific enzymes or the result of a recent gene duplication (Piatigorsky, J., and Wistow, G. (1991) Science 252, 1078-1079). One such enzyme, aldehyde dehydrogenase (ALDH), has been recruited as a lens crystallin in certain mammals (Wistow, G., and Kim, H. (1991) J. Mol. Evol. 32, 262-269) and cephalopods (Tomarev, S., Zinovieva, R., and Piatigorsky, J. (1991) J. Biol. Chem. 266, 24226-24231). We report here that a transparent tissue, derived from muscle but functioning as a lens in the light-emitting organ of a squid, Euprymna scolopes, shows striking biochemical convergence with the epidermally derived ocular lenses of some mammals and cephalopods. In the light organ lens of E. scolopes, an ALDH-like protein is the predominant molecular component. The typical muscle-specific proteins are replaced as the dominant species by a protein composed of 54-kDa subunits. This protein, which we designate as L-crystallin, constitutes approximately 70% of the total soluble protein of the light organ lens. The amino acid sequences of three peptides of L-crystallin (approximately 9% of the total protein) showed 54.5% sequence identity with human cytosolic ALDH. Using polyclonal antiserum made against L-crystallin, we found that it is present in low abundance in other tissues of the squid, including muscle and the ocular lens. This polyclonal antiserum also cross-reacted with the ALDH-like crystallins found in the ocular lenses of certain mammals and cephalopods. L-Crystallin showed no ALDH activity, which is similar to several other enzyme/crystallins, including ALDH/eta-crystallin (Wistow, G., and Kim, H. (1991) J. Mol. Evol. 32, 262-269). The characteristics of this muscle-derived lens are evidence that a common biochemical basis underlies transparency and that certain proteins may possess properties that promote their selection as lens structural proteins.

Aldehyde Dehydrogenase↗

An intelligent system for comparing protein structures.

An approach to protein structure comparison is presented which uses techniques of artificial intelligence (AI) to generate a mapping between two protein structures. The approach proceeds by first identifying the seed of a possible mapping, and then searching for ways to extend the seed by incorporating corresponding elements from the two proteins. Correspondence is judged using heuristic functions which assess the similarity of the structural environments of the elements. The search can be guided by separately encoded knowledge. A prototype has been implemented which is able to rapidly create mappings with a high degree of accuracy in test cases.

Animals↗

Characterization of the African swine fever virus structural protein p14.5: a DNA binding protein.

The gene encoding the structural protein p14.5 of African swine fever virus (ASFV) has been mapped and sequenced. This gene, designated E120R, is located in the Sa/l H/EcoRl E restriction fragment of the ASFV genome and is predicted to encode a protein of 120 amino acids with a molecular weight of 13.4 kDa. Northern-blot analysis showed that E120R is transcribed at late times during the viral replication cycle. The E120R gene product has been expressed in Escherichia coli, purified, and used as an antigen for antibody production. The antiserum anti-pE120R recognized a protein in infected cell extracts with an apparent molecular mass of 14.5 kDa, named p14.5. This antiserum also detected protein p14.5 in purified virus particles. Protein p14.5 is synthesized late in infection and is located in viral factories. Immunoprecipitation analysis and binding-assay experiments have shown that protein p14.5 interacts with a protein that could correspond to the major structural protein p72. Purified protein p14.5 interacts with DNA in a sequence-independent manner. It binds to both single-stranded and double-stranded DNA. A possible role of protein p14.5 in the encapsidation of ASFV DNA is suggested.

African Swine Fever Virus↗

Using protein structural information in evolutionary inference: transmembrane proteins.

We present a model of amino acid sequence evolution based on a hidden Markov model that extends to transmembrane proteins previous methods that incorporate protein structural information into phylogenetics. Our model aims to give a better understanding of processes of molecular evolution and to extract structural information from multiple alignments of transmembrane sequences and use such information to improve phylogenetic analyses. This should be of value in phylogenetic studies of transmembrane proteins: for example, mitochondrial proteins have acquired a special importance in phylogenetics and are mostly transmembrane proteins. The improvement in fit to example data sets of our new model relative to less complex models of amino acid sequence evolution is statistically tested. To further illustrate the potential utility of our method, phylogeny estimation is performed on primate CCR5 receptor sequences, sequences of l and m subunits of the light reaction center in purple bacteria, guinea pig sequences with respect to lagomorph and rodent sequences of calcitonin receptor and K-substance receptor, and cetacean sequences of cytochrome b.

Animals↗

Reduced representation model of protein structure prediction: statistical potential and genetic algorithms.

A reduced representation model, which has been described in previous reports, was used to predict the folded structures of proteins from their primary sequences and random starting conformations. The molecular structure of each protein has been reduced to its backbone atoms (with ideal fixed bond lengths and valence angles) and each side chain approximated by a single virtual united-atom. The coordinate variables were the backbone dihedral angles phi and psi. A statistical potential function, which included local and nonlocal interactions and was computed from known protein structures, was used in the structure minimization. A novel approach, employing the concepts of genetic algorithms, has been developed to simultaneously optimize a population of conformations. With the information of primary sequence and the radius of gyration of the crystal structure only, and starting from randomly generated initial conformations, I have been able to fold melittin, a protein of 26 residues, with high computational convergence. The computed structures have a root mean square error of 1.66 A (distance matrix error = 0.99 A) on average to the crystal structure. Similar results for avian pancreatic polypeptide inhibitor, a protein of 36 residues, are obtained. Application of the method to apamin, an 18-residue polypeptide with two disulfide bonds, shows that it folds apamin to native-like conformations with the correct disulfide bonds formed.

Algorithms↗

Highly fluctuating protein structures revealed by variable-pressure nuclear magnetic resonance.

Although our knowledge of basic folded structures of proteins has dramatically improved, the extent of our corresponding knowledge of higher-energy conformers remains extremely slim. The latter information is crucial for advancing our understanding of mechanisms of protein function, folding, and conformational diseases. Direct spectroscopic detection and analysis of structures of higher-energy conformers are limited, particularly under physiological conditions, either because their equilibrium populations are small or because they exist only transiently in the folding process. A new experimental strategy using pressure perturbation in conjunction with multidimensional NMR spectroscopy is being used to overcome this difficulty. A number of rare conformers are detected under pressure for a variety of proteins such as the Ras-binding domain of RalGDS, beta-lactoglobulin, dihydrofolate reductase, ubiquitin, apomyoglobin, p13(MTCP1), and prion, which disclose a rich world of protein structure between basically folded and globally unfolded states. Specific structures suggest that these conformers are designed for function and are closely identical to kinetic intermediates. Detailed structural determination of higher-energy conformers with variable-pressure NMR will extend our knowledge of protein structure and conformational fluctuation over most of the biologically relevant conformational space.

Animals↗

Genetic manipulation of arterivirus alternative mRNA leader-body junction sites reveals tight regulation of structural protein expression.

To express its structural proteins, the arterivirus Equine arteritis virus (EAV) produces a nested set of six subgenomic (sg) RNA species. These RNA molecules are generated by a mechanism of discontinuous transcription, during which a common leader sequence, representing the 5' end of the genomic RNA, is attached to the bodies of the sg RNAs. The connection between the leader and body parts of an mRNA is formed by a short, conserved sequence element termed the transcription-regulating sequence (TRS), which is present at the 3' end of the leader as well as upstream of each of the structural protein genes. With the exception of RNA3, only one body TRS was previously assumed to be used to join the leader and body of each EAV sg RNA. Here we show that for the synthesis of two other sg RNAs, RNA4 and RNA5, alternative leader-body junction sites that differ substantially in transcriptional activity are used. By site-directed mutagenesis of an EAV infectious cDNA clone, the alternative TRSs used to generate RNA3, -4, and -5 were inactivated, which strongly influenced the corresponding RNA levels and the production of infectious progeny virus. The relative amounts of RNA produced from alternative TRSs differed significantly and corresponded to the relative infectivities of the virus mutants. This strongly suggested that the structural proteins that are expressed from these RNAs are limiting factors during the viral life cycle and that the discontinuous step in sg RNA synthesis is crucial for the regulation of their expression. On the basis of a theoretical analysis of the predicted RNA structure of the 3' end of the EAV genome, we propose that the local secondary RNA structure of the body TRS regions is an important factor in the regulation of the discontinuous step in EAV sg mRNA synthesis.

Animals↗