PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Protein structure”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Ab-initio prediction and reliability of protein structural genomics by PROPAINOR algorithm.

We have formulated the ab-initio prediction of the 3D-structure of proteins as a probabilistic programming problem where the inter-residue 3D-distances are treated as random variables. Lower and upper bounds for these random variables and the corresponding probabilities are estimated by nonparametric statistical methods and knowledge-based heuristics. In this paper we focus on the probabilistic computation of the 3D-structure using these distance interval estimates. Validation of the predicted structures shows our method to be more accurate than other computational methods reported so far. Our method is also found to be computationally more efficient than other existing ab-initio structure prediction methods. Moreover, we provide a reliability index for the predicted structures too. Because of its computational simplicity and its applicability to any random sequence, our algorithm called PROPAINOR (PROtein structure Prediction by AI an Nonparametric Regression) has significant scope in computational protein structural genomics.

Algorithms↗

Protein structure comparison by probability-based matching of secondary structure elements.

MOTIVATION: Protein structure comparison (PSC) has been used widely in studies of structural and functional genomics. However, PSC is computationally expensive and as a result almost all of the PSC methods currently in use look only for the optimal alignment and ignore many alternative alignments that are statistically significant and that may provide insight into protein evolution or folding. RESULTS: We have developed a new PSC method with efficiency to detect potentially viable alternative alignments in all-against-all database comparisons. The efficiency of the new PSC method derives from the ability to directly home in on a limited number of viable and ranked alignment solutions based on intuitively derived SSE (secondary structure element)-matching probabilities.

Algorithms↗

Empirical relationships between protein structure and carboxyl pKa values in proteins.

Relationships between protein structure and ionization of carboxyl groups were investigated in 24 proteins of known structure and for which 115 aspartate and 97 glutamate pK(a) values are known. Mean pK(a) values for aspartates and glutamates are < or = 3.4 (+/-1.0) and 4.1 (+/-0.8), respectively. For aspartates, mean pK(a) values are 3.9 (+/-1.0) and 3.1 (+/-0.9) in acidic (pI < 5) and basic (pI > 8) proteins, respectively, while mean pK(a) values for glutamates are approximately 4.2 for acidic and basic proteins. Burial of carboxyl groups leads to dispersion in pK(a) values: pK(a) values for solvent-exposed groups show narrow distributions while values for buried groups range from < 2 to 6.7. Calculated electrostatic potentials at the carboxyl groups show modest correlations with experimental pK(a) values and these correlations are not improved by including simple surface-area-based terms to account for the effects of desolvation. Mean aspartate pK(a) values decrease with increasing numbers of hydrogen bonds but this is not observed at glutamates. Only 10 pK(a) values are > 5.5 and most are found in active sites or ligand-binding sites. These carboxyl groups are buried and usually accept no more than one hydrogen bond. Aspartates and glutamates at the N-termini of helices have mean pK(a) values of 2.8 (+/-0.5) and 3.4 (+/-0.6), respectively, about 0.6 units less than the overall mean values.

Animals↗

Sensitivity and selectivity in protein structure comparison.

Seven protein structure comparison methods and two sequence comparison programs were evaluated on their ability to detect either protein homologs or domains with the same topology (fold) as defined by the CATH structure database. The structure alignment programs Dali, Structal, Combinatorial Extension (CE), VAST, and Matras were tested along with SGM and PRIDE, which calculate a structural distance between two domains without aligning them. We also tested two sequence alignment programs, SSEARCH and PSI-BLAST. Depending upon the level of selectivity and error model, structure alignment programs can detect roughly twice as many homologous domains in CATH as sequence alignment programs. Dali finds the most homologs, 321-533 of 1120 possible true positives (28.7%-45.7%), at an error rate of 0.1 errors per query (EPQ), whereas PSI-BLAST finds 365 true positives (32.6%), regardless of the error model. At an EPQ of 1.0, Dali finds 42%-70% of possible homologs, whereas Matras finds 49%-57%; PSI-BLAST finds 36.9%. However, Dali achieves >84% coverage before the first error for half of the families tested. Dali and PSI-BLAST find 9.2% and 5.2%, respectively, of the 7056 possible topology pairs at an EPQ of 0.1 and 19.5, and 5.9% at an EPQ of 1.0. Most statistical significance estimates reported by the structural alignment programs overestimate the significance of an alignment by orders of magnitude when compared with the actual distribution of errors. These results help quantify the statistical distinction between analogous and homologous structures, and provide a benchmark for structure comparison statistics.

Computational Biology↗

Protein structure by mechanical triangulation.

Knowledge of protein structure is essential to understand protein function. High-resolution protein structure has so far been the domain of ensemble methods. Here, we develop a simple single-molecule technique to measure spatial position of selected residues within a folded and functional protein structure in solution. Construction and mechanical unfolding of cysteine-engineered polyproteins with controlled linkage topology allows measuring intramolecular distance with angstrom precision. We demonstrate the potential of this technique by determining the position of three residues in the structure of green fluorescent protein (GFP). Our results perfectly agree with the GFP crystal structure. Mechanical triangulation can find many applications where current bulk structural methods fail.

Calibration↗

Protein structure modelling from remote sequence similarity.

Many methods exist for taking a sequence that exhibits similarity to another of known structure and building a molecular model. However, when the sequence similarity is very remote and fragmentary, this 'modelling-by-homology' approach is less reliable. Current methods that tackle this problem are reviewed below, taking as an example the construction of a predicted model for the retroviral protease. This earlier work, which was only partially automatic, identified many of the outstanding difficulties that have subsequently been automated in computer programs, developed both by the author and many others. Because of the rapid proliferation of methods and their variants, an exhaustive review of the literature has not been possible and the following survey concentrates on the developments of the author and colleagues to explain the basic methods.

Algorithms↗

DMAPS: a database of multiple alignments for protein structures.

The database of multiple alignments for protein structures (DMAPS) provides instant access to pre-computed multiple structure alignments for all protein structure families in the Protein Data Bank (PDB). Protein structure families have been obtained from four distinct classification methods including SCOP, CATH, ENZYME and CE, and multiple structure alignments have been built for all families containing at least three members, using CE-MC software. Currently, multiple structure alignments are available for 3050 SCOP-, 3087 CATH-, 664 ENZYME- and 1707 CE-based families. A web-based query system has been developed to retrieve multiple alignments for these families using the PDB chain ID of any member of a family. Multiple alignments can be viewed or downloaded in six different formats, including JOY/html, TEXT, FASTA, PDB (superimposed coordinates), JOY/postscript and JOY/rtf. DMAPS is accessible online at http://bioinformatics.albany.edu/~dmaps.

Databases, Protein↗

The structure of Ski8p, a protein regulating mRNA degradation: Implications for WD protein structure.

Ski8p is a 44-kD protein that primarily functions in the regulation of exosome-mediated, 3'--> 5' degradation of damaged mRNA. It does so by forming a complex with two partner proteins, Ski2p and Ski3p, which complete a complex that is capable of recruiting and activating the exosome/Ski7p complex that functions in RNA degradation. Ski8p also functions in meiotic recombination in complex with Spo11 in yeast. It is one of the many hundreds of primarily eukaryotic proteins containing tandem copies of WD repeats (also known as WD40 or beta-transducin repeats), which are short ~40 amino acid motifs, often terminating in a Trp-Asp dipeptide. Genomic analyses have demonstrated that WD repeats are found in 1%-2% of proteins in a typical eukaryote, but are extremely rare in prokaryotes. Almost all structurally characterized WD-repeat proteins are composed of seven such repeats and fold into seven-bladed beta propellers. Ski8p was thought to contain five WD repeats on the basis of primary sequence analysis implying a five-bladed propeller. The 1.9 A crystal structure unexpectedly exhibits a seven-bladed propeller fold with seven structurally authentic WD repeats. Structure-based sequence alignments show additional sequence diversity in the two undetected repeats. This demonstrates that many WD repeats have not yet been identified in sequences and also raises the possibility that the seven-bladed propeller may be the predominant fold for this family of proteins.

Amino Acid Motifs↗

A method for simultaneous alignment of multiple protein structures.

Here, we present MultiProt, a fully automated highly efficient technique to detect multiple structural alignments of protein structures. MultiProt finds the common geometrical cores between input molecules. To date, most methods for multiple alignment start from the pairwise alignment solutions. This may lead to a small overall alignment. In contrast, our method derives multiple alignments from simultaneous superpositions of input molecules. Further, our method does not require that all input molecules participate in the alignment. Actually, it efficiently detects high scoring partial multiple alignments for all possible number of molecules in the input. To demonstrate the power of MultiProt, we provide a number of case studies. First, we demonstrate known multiple alignments of protein structures to illustrate the performance of MultiProt. Next, we present various biological applications. These include: (1) a partial alignment of hinge-bent domains; (2) identification of functional groups of G-proteins; (3) analysis of binding sites; and (4) protein-protein interface alignment. Some applications preserve the sequence order of the residues in the alignment, whereas others are order-independent. It is their residue sequence order-independence that allows application of MultiProt to derive multiple alignments of binding sites and of protein-protein interfaces, making MultiProt an extremely useful structural tool.

Algorithms↗

Identification of amino acids involved in protein structural uniqueness: implication for de novo protein design.

Structural uniqueness is characteristic of native proteins and is essential to express their biological functions. The major factors that bring about the uniqueness are specific interactions between hydrophobic residues and their unique packing in the protein core. To find the origin of the uniqueness in their amino acid sequences, we analyzed the distribution of the side chain rotational isomers (rotamers) of hydrophobic amino acids in protein tertiary structures and derived deltaS(contact), the conformational-entropy changes of side chains by residue-residue contacts in each secondary structure. The deltaS(contact) values indicate distinct tendencies of the residue pairs to restrict side chain conformation by inter-residue contacts. Of the hydrophobic residues in alpha-helices, aliphatic residues (Leu, Val, Ile) strongly restrict the side chain conformations of each other. In beta-sheets, Met is most strongly restricted by contact with Ile, whereas Leu, Val and Ile are less affected by other residues in contact than those in alpha-helices. In designed and native protein variants, deltaS(contact) was found to correlate with the folding-unfolding cooperativity. Thus, it can be used as a specificity parameter for designing artificial proteins with a unique structure.

Amino Acid Sequence↗

An integrated approach to the analysis and modeling of protein sequences and structures. I. Protein structural alignment and a quantitative measure for protein structural distance.

We have devised and implemented in PrISM (protein informatics system for modeling) a new measure of protein structural relationships, the protein structural distance (PSD). The PSD is designed to describe relationships between protein structures in quantitative rather than descriptive terms and is applicable both when two structures are very similar, and when they are very different. It is calculated with a structural alignment procedure that uses double dynamic programming to align secondary structure elements and an iterative rigid body superposition that minimizes the root-mean-square deviation of C(alpha) atoms. The alignment algorithm, as implemented on a modest workstation, is computationally efficient, allowing for large-scale structural comparisons. PSD scores for more than one and a half million pairs of proteins were calculated and compared to the discrete classification of proteins in the SCOP database. The PSD scores, which were obtained automatically, are in large part consistent with the manually derived classifications in SCOP. Discrepancies do arise, however, due, in part, to the fact that SCOP uses criteria other than structural similarity to derive classifications while the PrISM procedure is exclusively structure based. Analysis of PSD scores suggests that there is a continuous aspect of protein conformation space, even though various classification schemes are extremely useful. The use of a continuous measure for structural distance between all pairs of proteins allows us, as described in the two accompanying papers to derive sequence/structure relationships in a more quantitative way than has previously been possible. An important strength of the approach implemented in PrISM is its ability to address many different kinds of queries interactively, making its structural comparison procedure a convenient computational tool that complements structural classification databases such as SCOP and CATH.

Algorithms↗

The structure and organization of lamprin genes: multiple-copy genes with alternative splicing and convergent evolution with insect structural proteins.

Lamprin is a unique structural protein which forms the extracellular matrix of several cartilaginous structures found in the lamprey. Lamprin is noncollagenous in nature but shows sequence similarities to elastins and to insect structural proteins. Here, we characterize the structure and organization of lamprin genes, demonstrating the presence of multiple similar but not identical copies of the lamprin gene in the genome of the lamprey. In at least one species of lamprey, Lampetra richardsoni, the multiple gene copies are arranged in tandem in the genome in a head-to-tail orientation. Lamprin genes from Petromyzon marinus contain either seven or eight exons, with exon 4 being alternatively spliced in all genes, resulting in a total of six different lamprin transcripts. All exon junctions are of class 1,1. An unusual feature of the lamprin gene structure is the distribution of the 3' untranslated region sequence among multiple exons. A TATA box and cap sequence have been identified in upstream sequences in close proximity to the transcription start site, but no CAAT box could be identified. Sequence and gene structure comparisons between lamprins, elastins, and insect structural proteins suggest that the regions of sequence similarity are the result of a process of convergent evolution.

Alternative Splicing↗

Protein structure prediction in the postgenomic era.

As the number of completely sequenced genomes rapidly increases, the postgenomic problem of gene function identification becomes ever more pressing. Predicting the structures of proteins encoded by genes of interest is one possible means to glean subtle clues as to the functions of these proteins. There are limitations to this approach to gene identification and a survey of the expected reliability of different protein structure prediction techniques has been undertaken.

Animals↗

Mapping low-resolution three-dimensional protein structures using chemical cross-linking and Fourier transform ion-cyclotron resonance mass spectrometry.

Techniques in mass spectrometry (MS) combined with chemical cross-linking have proven to be efficient tools for the rapid determination of low-resolution three-dimensional (3-D) structures of proteins. The general procedure involves chemical cross-linking of a protein followed by enzymatic digestion and MS analysis of the resulting peptide mixture. These experiments are generally fast and do not require large quantities of protein. However, the large number of peptide species created from the digestion of cross-linked proteins makes it difficult to identify relevant intermolecular cross-linked peptides from MS data. We present a method for mapping low-resolution 3-D protein structures by combining chemical cross-linking with high-resolution FTICR (Fourier transform ion-cyclotron resonance) mass spectrometry using cytochrome c and hen egg lysozyme as model proteins. We applied several homo-bifunctional, amine-reactive cross-linking reagents that bridge distances from 6 to 16 A. The non-digested cross-linking reaction mixtures were monitored by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOFMS) to determine the extent of cross-linking. Enzymatically digested reaction mixtures were separated by nano-high-performance liquid chromatography (nano-HPLC) on reverse-phase columns applying water/acetonitrile gradients with flow rates of 200 nL/min. The nano-HPLC system was directly coupled to an FTICR mass spectrometer equipped with a nano-ESI (electrospray ionization) source. Cross-linking products were identified using a combination of the GPMAW software and ExPASy Proteomics tools. For correct assignment of the cross-linking products the key factor is to rely on a mass spectrometric method providing both high resolution and high mass accuracy, such as FTICRMS. By combining chemical cross-linking with FTICRMS we were able to rapidly define several intramolecular constraints for cytochrome c and lysozyme.

Animals↗

Prediction of protein structural classes by a new measure of information discrepancy.

Since it was observed that the structural class of a protein is related to its amino acid composition, various methods based on amino acid composition have been proposed to predict protein structural classes. Though those methods are effective to some degree, their predictive quality is confined because amino acid composition cannot sufficiently include the information of protein sequences. In this paper, a measure of information discrepancy is applied to the prediction of protein structural classes; different from the previous methods, this new approach is based on the comparisons of subsequence distributions; therefore, the effect of residue order on protein structure is taken into account. The predictive results of the new approach on the same data set are better than those of the previous methods. As to a data set of 1401 sequences with no more than 30% redundancy, the overall correctness rates of resubstitution test and Jackknife test are 99.4 and 75.02%, respectively, and to other data sets the similar results are also obtained. All tests demonstrate that the residue order along protein sequences plays an important role on recognition of protein structural classes, especially for alpha/beta proteins and alpha+beta proteins. In addition, the tests also show that the new method is simple and efficient.

Algorithms↗

Functional inferences from blind ab initio protein structure predictions.

Ab initio protein structure prediction methods have improved dramatically in the past several years. Because these methods require only the sequence of the protein of interest, they are potentially applicable to the open reading frames in the many organisms whose sequences have been and will be determined. Ab initio methods cannot currently produce models of high enough resolution for use in rational drug design, but there is an exciting potential for using the methods for functional annotation of protein sequences on a genomic scale. Here we illustrate how functional insights can be obtained from low-resolution predicted structures using examples from blind ab initio structure predictions from the third and fourth critical assessment of structure prediction (CASP3, CASP4) experiments.

Computational Biology↗

PALI: a database of alignments and phylogeny of homologous protein structures.

PALI is a database of structure-based sequence alignments and phylogenetic relationships derived on the basis of three-dimensional structures of homologous proteins. This database enables grouping of pairs of homologous protein structures on the basis of their sequence identity calculated from the structure-based alignment and PALI also enables association of a new sequence to a family and automatic generation of a dendrogram combining the query sequence and homologous protein structures.

Databases, Factual↗