PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

What is the probability of a chance prediction of a protein structure with an rmsd of 6 A?

BACKGROUND: The root mean square deviation (rmsd) between corresponding atoms of two protein chains is a commonly used measure of similarity between two protein structures. The smaller the rmsd is between two structures, the more similar are these two structures. In protein structure prediction, one needs the rmsd between predicted and experimental structures for which a prediction can be considered to be successful. Success is obvious only when the rmsd is as small as that for closely homologous proteins (< 3 A). To estimate the quality of the prediction in the more general case, one has to compare the native structure not only with the predicted one but also with randomly chosen protein-like folds. One can ask: how many such structures must be considered to find a structure with a given rmsd from the native structure? RESULTS: We calculated the rmsd values between native structures of 142 proteins and all compact structures obtained in the threading of these protein chains over 364 non-homologous structures. The rmsd distributions have a Gaussian form, with the average rmsd approximately proportional to the radius of gyration. CONCLUSIONS: We estimated the number of protein-like structures required to obtain a structure within an rmsd of 6 A to be 10(4)-10(5) for chains of 60-80 residues and 10(11)-10(12) structures for chains of 160-200 residues. The probability of obtaining a 6 A rmsd by chance is so remote that when such structures are obtained from a prediction algorithm, it should be considered quite successful.

Databases, Factual

Synthesis of mycolic acids of mycobacteria: an assessment of the cell-free system in light of the whole genome.

Mycolic acids are 70-90 carbon, alpha-alkyl, beta-hydroxy fatty acids constituting a major component of the cell envelope of Mycobacterium tuberculosis. The fact that the mycolic acid biosynthetic pathway is both essential in mycobacteria and the target for many first-line anti-TB drugs necessitates a detailed understanding of its biochemistry. A whole cell-free, but cell particulate- and membrane-containing enzyme preparation for mycolic acid biosynthesis was developed a few years ago and studied extensively. This system was shown to catalyze the synthesis of mature mycolic acids from [14C]acetate, but allows only minimal deposition into the cell wall proper. In the meantime the sequence of the entire genome of M. tuberculosis has been elucidated and its analysis using numerous protein sequence-based algorithms predicted cytoplasmic localization and a soluble, not a particulate, nature for the enzymes involved in the mycolic acid synthetic pathway. Accordingly, we re-assessed the 'cell-free' system for mycolic acid synthesis and concluded that it is probably due to the presence of unbroken cells, since viable cells were recovered from the cell wall preparation. The amount of whole cells depended upon the efficiency of the cell disruption method and conditions, and the amount of mycolic acid synthesized by the putative cell-free system correlated with the content of whole cells. Thus, accumulated results from the use of this 'cell-free' cell wall-based system should be re-evaluated in the light of these new data.

Cell Fractionation

Secondary-structure predictions of calcium-binding proteins.

The known tertiary structure of carp muscle parvalbumin is consistent with an "EF-hand" architecture (helix-loop-helix) for each calcium-ion binding site. Primary-sequence alignments have indicated four EF hands in rabbit skeletal muscle troponin C and in rabbit myosin alkali light chains. Five secondary-structure prediction methods, based on amino acid sequence only, have been fully computerized and used to calculate joint prediction histograms for several calcium-binding proteins. The joint histogram can suggest directly the extent and sequence of the helical- and loop-structural elements, as well as any secondary structural distortions or evolutionary developments. Since the histogram predicted well the length and sequence of secondary structural elements in carp muscle parvalbumin, it seemed reasonable to calculate the joint distribution for other proteins that might bind calcium through the EF-hand configuration. The histograms indicated the four EF-hand regions speculated fro rabbit skeletal muscle troponin C but suggested only three such hands in bovine cardiac muscle troponin C and with a distorted fourth hand. Considerable secondary structural distortion is postulated for the alkali light chains. Possible EF configurations consistent with the histogram results are speculated for Escherichia coli acyl-carrier protein and bovine prothrombin fragment 1, which have been shown to bind calcium. The secondary-structure-prediction algorithms appear to be a useful adjunct to sequence-alignment techniques, especially in cases where the primary sequence homology is weak or the evolutionary distance is large.

Amino Acid Sequence

Substrate binding and catalysis by ubiquitin C-terminal hydrolases: identification of two active site residues.

Ubiquitin C-terminal hydrolases (UCH's) are a newly-defined class of thiol proteases implicated in the proteolytic processing of polymeric ubiquitin. They are important for the generation of monomeric ubiquitin, the active component of the eukaryotic ubiquitin-dependent proteolytic system. There are at least three mammalian isozymes which are tissue specific and developmentally regulated. To study the structure and functional roles of these highly homologous enzymes, we have subcloned and overexpressed two of these isozymes, UCH-L1 and UCH-L3. Here, we report their purification, physical characteristics, and the mutagenesis of UCH-L1. Site-directed mutagenesis of UCH-L1 reveals that C90 and H161 are involved in catalytic rate enhancement. Data from circular dichroic and Raman spectroscopy, as well as secondary structure prediction algorithms, indicate that both isozymes have a significant amount of alpha-helix (> 35%), and contain no disulfide bonds. Both enzymes are reasonably stable, undergoing a reversible thermal denaturation at 52 degrees C. These transitions are characterized by thermodynamic parameters typical of single domain globular proteins. Substrate binding affinity to UCH-L3 was directly measured by equilibrium gel filtration (Kd = 0.5 microM), and the results are similar to the kinetically determined Km for ubiquitin ethyl ester (o.6 microM). The binding is primarily electrostatic in nature and indicates the existence of a specific and extensive binding site for ubiquitin on the surface of the enzyme.

Amino Acid Sequence

Solution conformation of a five-nucleotide RNA bulge loop from a group I intron.

We present the solution conformation, determined by NMR spectroscopy, of a five-nucleotide RNA bulge loop. The bulge interrupts the stem of a 25-nucleotide RNA hairpin, and its sequence and flanking sequences are those of a conserved bulge from a Group I intron. The secondary structure of the bulge loop in the hairpin context is that predicted by the secondary structure prediction algorithm of Zuker. It differs, however, from the secondary structure deduced from sequence covariation of the bulge in the context of the functionally folded Group I introns and observed in the crystal structure of an independently folding domain of the Group I intron from Tetrahymena thermophila. This difference represents an exception to the heierarchical model of RNA folding in which preformed elements of secondary structure interact to form a tertiary structure. The three-dimensional structure of the bulge loop is characterized by discontinuous base stacking. Adjacent adenines stack with each other and with the flanking double helices. However, the position of the central uracil is not well defined by NOE distance constraints and is a point of discontinuity in the base stacking.

Animals

Structural studies of detergent-solubilized and vesicle-reconstituted low-density lipoprotein (LDL) receptor.

The low-density lipoprotein (LDL) receptor plays a key role in maintaining circulating and cellular cholesterol homeostasis. The LDL receptor is a transmembrane glycoprotein whose biochemical and genetic properties have been extensively studied notably by Brown, Goldstein and colleagues [Brown, M. S., & Goldstein, J. L., (1986) Science 232, 34-47]. However, few if any structural studies of the LDL receptor have been reported, and details of its secondary and tertiary structure are lacking. In an attempt to determine the low-resolution structure of the LDL receptor, we have purified the receptor from bovine adrenal cortices using modifications of the method of Schneider et al. [Schneider, W. J., Goldstein, J. L., & Brown, M. S. (1985) Methods in Enzymol.109, 405-417]. Using circular dichroism, the secondary structure of the detergent-solubilized bovine LDL receptor at 25 degrees C was shown to be 19% alpha-helix, 42% beta-sheet, and 39% random coil. Interestingly, the detergent-solubilized receptor appeared to be quite resistant to changes in secondary structure over the temperature range 10-90 degrees C, with only minor but reversible changes being observed. In contrast, a more pronounced unfolding of the detergent-solubilized receptor was observed in the presence of guanidinium hydrochloride. Using the complete sequence of the human LDL receptor, sequence analysis by the Chou-Fasman prediction algorithm showed quite good agreement with the experimentally determined secondary structure of the bovine LDL receptor at 25 degrees C. Finally, the purified, bovine LDL receptor was reconstituted into large unilamellar vesicles of egg yolk phosphatidylcholine using a procedure exploiting preformed vesicles and detergent dialysis. We showed previously using negative stain electron microscopy that reconstituted vesicles bind LDL. Now, using cryoelectron microscopy of frozen hydrated reconstituted vesicles evidence of an extended, stick-like morphology (length approximately 120 A) for the extracellular domain of the LDL receptor has been obtained. Successful purification of the receptor, its incorporation into single bilayer vesicles, and its direct visualization by cryoelectron microscopy pave the way for more detailed structural studies of the LDL receptor and the receptor-LDL complex.

Adrenal Cortex

Identification of protein-coding regions in Arabidopsis thaliana genome based on quadratic discriminant analysis.

A new method (MZEF) for predicting internal coding exons in genomic DNA sequences has been developed. This method is based on a prediction algorithm that uses the quadratic discriminant function for multivariate statistical pattern recognition. With improved feature measures, an Arabidopsis thaliana-specific implementation of MZEF is completed and made available to the plant genome community.

Algorithms

Structural features of the human salivary mucin, MUC7.

Human salivary mucin (MUC7) is characterized by a single polypeptide chain of 357 aa. Detailed analysis of the derived MUC7 peptide sequence reveals five distinct regions or domains: (1) an N-terminal basic, histatin-like domain which has a leucine-zipper segment, (2) a moderately glycosylated domain, (3) six heavily glycosylated tandem repeats each consisting of 23 aa, (4) another heavily glycosylated MUC1- and MUC2-like domain, and (5) a C-terminal leucine-zipper segment. Chemical analysis and semi-empirical prediction algorithms for O-glycosylation suggested that 86/105 (83%) Ser/Thr residues were O-glycosylated with the majority located in the tandem repeats. The high (approximately 25%) proline content of MUC7 including 19 diproline segments suggested the presence of polyproline type structures. CD studies of natural and synthetic diproline-rich peptides and glycopeptides indicated that polyproline type structures do play a significant role in the conformational dynamics of MUC7. In addition, crystal structure analysis of a synthetic diproline segment (Boc-Ala-Pro-OBzl) revealed a polyproline type II extended structure. Collectively, the data indicate that the polyproline type II structure, dispersed throughout the tandem repeats, may impart a stiffening of the backbone and could act in consort with the glycosylated segments to keep MUC7 in a semi-rigid, rod shaped conformation resembling a 'bottle-brush' model.

Amino Acid Sequence

Beta-hairpin families in globular proteins.

Beta-hairpins, one of the simplest supersecondary structures, are widespread in globular proteins, and have often been suggested as possible sites for nucleation. Here we consider the conformation and sequences of the loop regions of beta-hairpins by analysing proteins of known structure. We find that the 'tight' beta-hairpins, classified by the length and conformations of their loop regions, form distinct families and that the loop regions of the family members have sequences which are characteristic of that family. The two-residue hairpin loops include almost entirely I' or II' beta-turns, in contrast to the general preference for type I and type II turns. These findings are being used to help define templates or consensus sequences to be incorporated into our existing supersecondary structure prediction algorithm. This information can also be used in model-building homologous proteins.

Amino Acid Sequence

Substrate recognition by proteinases.

The molecular recognition of limited proteolytic site substrates by serine proteinases has been compared and contrasted to the recognition of serine proteinase inhibitors, utilising the coordinate sets contained in the Brookhaven Protein Databank. Most families of these inhibitors are known to possess a structurally conserved recognition motif at their reactive site-binding loops. Structural comparisons with trypsin limited proteolytic sites revealed that the in situ conformation of these substrates bears little resemblance to the inhibitor-binding loops. Assuming that both inhibitors and substrates bind to the proteinase in the same manner, segmental mobility would be required to permit substrates to adopt an 'inhibitor-like' binding conformation, which is presumed to be necessary for proteolysis. Modelling experiments have been conducted to attempt to introduce such a conformation into tryptic limited proteolytic segments of the native proteins, to test the ability of the limited proteolytic sites to alter their geometry. Further to this, the conformational parameters of accessibility, protrusion, mobility and secondary structure have been analysed and incorporated into a predictive algorithm to assign likely limited proteolytic sites within native protein structures.

Binding Sites

Empirical studies of protein secondary structure by vibrational circular dichroism and related techniques. Alpha-lactalbumin and lysozyme as examples.

Vibrational circular dichroism (VCD) has been shown to be sensitive to secondary structure in proteins and peptides and has been used as the basis for quantitative secondary-structure-prediction algorithms. However, the accuracy of these algorithms is not matched by the apparent qualitative sensitivity of the VCD spectra. This report provides examples of the use of VCD to follow structural change spectrally and to clarify the qualitative nature of the structural changes underlying the spectral variation. The VCD spectra and the complementary UV electronic CD (ECD) and FTIR spectra of alpha-lactalbumin (LA) have been studied as a function of pH, denaturation, Ca2+ ion and solvent conditions for several species. Spectral data for lysozyme were compared with those of LA because of their very similar crystal structures. In fact, these proteins in D2O-based pH 7 solution have quite different spectra using these optical techniques. Even for the LA proteins, the human differs from the bovine and goat species. Furthermore, under low pH conditions, where the LAs are in a reversibly denatured, molten globule form, the spectra are more similar, species variation is minimal and the spectral differences from lysozyme are in fact smaller. Our data are consistent with native, pH 7, alpha-lactalbumin having a less well organized structure than lysozyme, possibly in a dynamic sense. Conversely, in the low-pH, molten globule form of LA, tertiary structure is lost which could relax constraints that might distort the helical segments in the native form. The differences between the interpretation of our results and those from X-ray and NMR data may be due to motional sampling of various geometries in LA which all contribute to the spectral signatures seen in optical spectra but whose contributions are washed out in NMR or frozen out in the crystal structure. Part of this flexibility may relate to the rather large 3(10)-helical content in the LA protein structure. Fluctionality may have specific functional effects, perhaps allowing LA to bind better to beta-galactosyl transferase and form the biologically active lactose synthetase complex.

Amino Acid Sequence

Identification of heparin-binding domains in the amyloid precursor protein of Alzheimer's disease by deletion mutagenesis and peptide mapping.

Recent studies have shown that the binding of the amyloid protein precursor (APP) of Alzheimer's disease to heparan sulfate proteoglycans (HSPGs) can modulate a neurite outgrowth-promoting function associated with APP. We used three different approaches to identify heparin-binding domains in APP. First, as heparin-binding domains are likely to be within highly folded regions of proteins, we analyzed the secondary structure of APP using several predictive algorithms. This analysis showed that two regions of APP695 contain a high degree of secondary structure, and clusters of basic residues, considered mandatory for heparin binding, were found, principally within these regions. To determine which domains of APP bind heparin, deletion mutants of APP695 were prepared and analyzed for binding to a heparin affinity column. The results suggested that there must be at least two distinct heparin-binding regions in APP. To identify novel heparin-binding regions, peptides homologous to candidate heparin-binding domains were analyzed for their ability to bind heparin. These experiments suggested that APP contains at least four heparin-binding domains. The presence of more than one heparin-binding domain on APP suggests the possibility that APP may interact with more than one type of glycosaminoglycan.

Alzheimer Disease

The leptin haemopoietic cytokine fold is stabilized by an intrachain disulfide bond.

Structure prediction algorithms have tagged leptin as the newest member of the haemopoietic cytokine family, a diverse class of secreted hormone-like factors with pleiotropic effects in immunity and haemopoietic development. While haemopoietic cytokines typically lack sequence similarity, they conserve a distinctive three-dimensional fold, a four-alpha-helix bundle structure that is recognized by the cognate family of haemopoietic cellular receptors. We have constructed a detailed molecular model of the human leptin helical fold that places the two cysteine residues of the leptin chain, Cys96 and Cys146, in close spatial proximity to each other. In this report, we present evidence that these cysteines are involved in an intrachain disulfide bridge that is critical for the structural integrity and stability of leptin. A leptin variant that is unable to form the disulfide link shows a reduced biological response when administered to leptin-deficient, ob/ob mice.

Algorithms

Insertion mutagenesis as a tool to predict the secondary structure of a muscarinic receptor domain determining specificity of G-protein coupling.

The N-terminal segment of the third intracellular loop (i3) of muscarinic acetylcholine receptors and other G protein-coupled receptors has been shown to largely determine the G-protein coupling selectivity displayed by a given receptor subtype. Based on secondary-structure prediction algorithms, we have tested the hypothesis that this region adopts an alpha-helical secondary structure. Using the rat m3 muscarinic receptor as a model system, a series of five mutant receptors, m3(+1A) to m3(+5A) were created in which one to five additional alanine residues were inserted between the end of the fifth transmembrane domain and the beginning of i3. We speculated that this manipulation should lead to a rotation of the N-terminal segment of the i3 domain (if it is in fact alpha-helically arranged), thus producing pronounced effects on receptor/G protein coupling. Pharmacological analysis of the various mutant receptors expressed in COS-7 cells showed that m3(+1A), m3(+3A), and m3(+4A) retained strong functional activity, whereas m3(+2A) and m3(+5A) proved to be virtually inactive. Helical wheel models show that this pattern is fully consistent with the notion that the N-terminal portion of i3 forms an amphiphilic alpha-helix and that the hydrophobic side of this helix represents the G-protein recognition surface.

Acetylcholine

Rat skeletal muscle selenoprotein W: cDNA clone and mRNA modulation by dietary selenium.

Rat skeletal muscle selenoprotein W cDNA was isolated and sequenced. The isolation strategy involved design of degenerate PCR primers from reverse translation of a partial peptide sequence. A reverse transcription-coupled PCR product from rat muscle mRNA was used to screen a muscle cDNA library prepared from selenium-supplemented rats. The cDNA sequence confirmed the known protein primary sequence, including a selenocysteine residue encoded by TGA, and identified residues needed to complete the protein sequence. RNA folding algorithms predict a stem-loop structure in the 3' untranslated region of the selenoprotein W mRNA that resembles selenocysteine insertion sequence (SE-CIS) elements identified in other selenocysteine coding cDNAs. Dietary regulation of selenoprotein W mRNA was examined in rat muscle. Dietary selenium at 0.1 ppm as selenite increased muscle mRNA 4-fold relative to a selenium-deficient diet. Higher dietary selenium produced no further increase in mRNA levels.

Amino Acid Sequence

Smoothness within ruggedness: the role of neutrality in adaptation.

RNA secondary structure folding algorithms predict the existence of connected networks of RNA sequences with identical structure. On such networks, evolving populations split into subpopulations, which diffuse independently in sequence space. This demands a distinction between two mutation thresholds: one at which genotypic information is lost and one at which phenotypic information is lost. In between, diffusion enables the search of vast areas in genotype space while still preserving the dominant phenotype. By this dynamic the success of phenotypic adaptation becomes much less sensitive to the initial conditions in genotype space.

Adaptation, Biological

Identification of protein coding regions in the human genome by quadratic discriminant analysis.

A new method for predicting internal coding exons in genomic DNA sequences has been developed. This method is based on a prediction algorithm that uses the quadratic discriminant function for multivariate statistical pattern recognition. Substantial improvements have been made (with only 9 discriminant variables) when compared with existing methods: HEXON [Solovyev, V. V., Salamov, A. A. & Lawrence, C. B. (1994) Nucleic Acids Res. 22, 5156-5163] (based on linear discriminant analysis) and GRAIL2 [Uberbacher, E. C. & Mural, R. J. (1991) Proc. Natl. Acad. Sci. USA 88, 11261-11265] (based on neural networks). A computer program called MZEF is freely available to the genome community and allows users to adjust prior probability and to output alternative overlapping exons.

Base Sequence

The A-kinase anchoring domain of type IIalpha cAMP-dependent protein kinase is highly helical.

Subcellular localization of the type II cAMP-dependent protein kinase is controlled by interaction of the regulatory subunit with A-Kinase Anchoring Proteins (AKAPs). This contribution examines the solution structure of a 44-residue region that is sufficient for high affinity binding to AKAPs. The N-terminal dimerization domain of the type IIalpha regulatory subunit of cAMP-dependent protein kinase was expressed to high levels on minimal media and uniformly isotopically enriched with 15N and 13C nuclei. Sequence-specific backbone and side chain resonance assignments have been made for greater than 95% of the amino acids in the free dimerization domain using high resolution multidimensional heteronuclear NMR techniques. Contrary to the results from secondary structure prediction algorithms, our analysis indicates that the domain is highly helical with a single 3-5-residue sequence involved in a beta-strand. The assignments and secondary structure analysis provide the basis for analyzing the structure and dynamics of the dimerization domain both free and complexed with specific anchoring proteins.

Amino Acid Sequence