PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Secondary-structure predictions of calcium-binding proteins.

The known tertiary structure of carp muscle parvalbumin is consistent with an "EF-hand" architecture (helix-loop-helix) for each calcium-ion binding site. Primary-sequence alignments have indicated four EF hands in rabbit skeletal muscle troponin C and in rabbit myosin alkali light chains. Five secondary-structure prediction methods, based on amino acid sequence only, have been fully computerized and used to calculate joint prediction histograms for several calcium-binding proteins. The joint histogram can suggest directly the extent and sequence of the helical- and loop-structural elements, as well as any secondary structural distortions or evolutionary developments. Since the histogram predicted well the length and sequence of secondary structural elements in carp muscle parvalbumin, it seemed reasonable to calculate the joint distribution for other proteins that might bind calcium through the EF-hand configuration. The histograms indicated the four EF-hand regions speculated fro rabbit skeletal muscle troponin C but suggested only three such hands in bovine cardiac muscle troponin C and with a distorted fourth hand. Considerable secondary structural distortion is postulated for the alkali light chains. Possible EF configurations consistent with the histogram results are speculated for Escherichia coli acyl-carrier protein and bovine prothrombin fragment 1, which have been shown to bind calcium. The secondary-structure-prediction algorithms appear to be a useful adjunct to sequence-alignment techniques, especially in cases where the primary sequence homology is weak or the evolutionary distance is large.

Amino Acid Sequence

Beta-hairpin families in globular proteins.

Beta-hairpins, one of the simplest supersecondary structures, are widespread in globular proteins, and have often been suggested as possible sites for nucleation. Here we consider the conformation and sequences of the loop regions of beta-hairpins by analysing proteins of known structure. We find that the 'tight' beta-hairpins, classified by the length and conformations of their loop regions, form distinct families and that the loop regions of the family members have sequences which are characteristic of that family. The two-residue hairpin loops include almost entirely I' or II' beta-turns, in contrast to the general preference for type I and type II turns. These findings are being used to help define templates or consensus sequences to be incorporated into our existing supersecondary structure prediction algorithm. This information can also be used in model-building homologous proteins.

Amino Acid Sequence

Substrate recognition by proteinases.

The molecular recognition of limited proteolytic site substrates by serine proteinases has been compared and contrasted to the recognition of serine proteinase inhibitors, utilising the coordinate sets contained in the Brookhaven Protein Databank. Most families of these inhibitors are known to possess a structurally conserved recognition motif at their reactive site-binding loops. Structural comparisons with trypsin limited proteolytic sites revealed that the in situ conformation of these substrates bears little resemblance to the inhibitor-binding loops. Assuming that both inhibitors and substrates bind to the proteinase in the same manner, segmental mobility would be required to permit substrates to adopt an 'inhibitor-like' binding conformation, which is presumed to be necessary for proteolysis. Modelling experiments have been conducted to attempt to introduce such a conformation into tryptic limited proteolytic segments of the native proteins, to test the ability of the limited proteolytic sites to alter their geometry. Further to this, the conformational parameters of accessibility, protrusion, mobility and secondary structure have been analysed and incorporated into a predictive algorithm to assign likely limited proteolytic sites within native protein structures.

Binding Sites

The insulin receptor juxtamembrane region contains two independent tyrosine/beta-turn internalization signals.

We have investigated the role of tyrosine residues in the insulin receptor cytoplasmic juxtamembrane region (Tyr953 and Tyr960) during endocytosis. Analysis of the secondary structure of the juxtamembrane region by the Chou-Fasman algorithms predicts that both the sequences GPLY953 and NPEY960 form tyrosine-containing beta-turns. Similarly, analysis of model peptides by 1-D and 2-D NMR show that these sequences form beta-turns in solution, whereas replacement of the tyrosine residues with alanine destabilizes the beta-turn. CHO cell lines were prepared expressing mutant receptors in which each tyrosine was mutated to phenylalanine or alanine, and an additional mutant contained alanine at both positions. These mutations had no effect on insulin binding or receptor autophosphorylation. Replacements with phenylalanine had no effect on the rate of [125I]insulin endocytosis, whereas single substitutions with alanine reduced [125I]insulin endocytosis by 40-50%. Replacement of both tyrosines with alanine reduced internalization by 70%. These data suggest that the insulin receptor contains two tyrosine/beta-turns which contribute independently and additively to insulin-stimulated endocytosis.

Amino Acid Sequence

Immunoprotective Leishmania major synthetic T cell epitopes.

Using the predictive algorithm of Rothbard and Taylor (1988. EMBO J. 7:93) and the primary structure of gp63 (Button, L., and M.R. McMaster. 1988. J. Exp. Med. 167:724; Miller, R.A., S.G. Reed, and M. Parsons. 1990. Mol. Biochem. Parasitol. 39:267) we have been able to delineate the structures of a number of gp63 T-cell epitopes which stimulate the proliferation of CD4+ cells. One of these synthetic antigens, inoculated subcutaneously with adjuvant, was shown to specifically induce proliferation of the Th1 subset and provided immunoprotection against two species of Leishmania parasites.

Algorithms

T cell determinant structure: cores and determinant envelopes in three mouse major histocompatibility complex haplotypes.

T lymphocytes recognize discrete regions on an antigen. The specificity of the T cell responses in three mouse strains of differing major histocompatibility complex (MHC) haplotype to a protein antigen, lysozyme, was analyzed using a series of peptides that walk the antigen in single amino acid steps. These peptide series were synthesized using the pin synthesis system, which was modified to allow the peptides to be cleaved from the pins into a physiological buffer free of toxic compounds. This methodology overcomes many of the problems associated with the production of peptides for screening proteins for antigenic determinants. The T cell determinants for the three strains were markedly different. This result points out the limitations of algorithms predicting determinants without reference to the MHC, and the importance of the empirical methodology. This analysis of the T cell response to lysozyme constitutes the most complete study of reactivity to a foreign protein to date and illustrates many important features of antigen recognition by T cells, e.g., presence of major and minor determinant regions. The outer boundaries of each immunogenic region, the determinant envelope, are difficult to define from recently immunized lymph nodes because of the heterogeneity in T cell recognition. However, core sequences common to all the immunogenic peptides in a continuous sequence can be easily defined.

Amino Acid Sequence

Identification of T cell epitopes occurring in a meningococcal class 1 outer membrane protein using overlapping peptides assembled with simultaneous multiple peptide synthesis.

The meningococcal class 1 outer membrane protein (OMP) plays an important role in the development of protective immunity against meningococcal infection, and is therefore considered to be a promising candidate antigen (Ag) for a meningococcal vaccine. The induction of an effective antibody response entirely depends upon T helper cells. To identify T cell epitopes of the OMP, we prepared 45 overlapping synthetic peptides representing the entire sequence of the class 1 protein of reference strain H44/76. Fully automated simultaneous multiple peptide synthesis (SMPS) was used to assemble the 45 twenty mer which overlapped by 12 amino acid residues on a 12 mumol scale. The peptides were tested for recognition by peripheral blood mononuclear cells (PBMC) obtained from 34 volunteers. Surprisingly, all synthetic peptides induced proliferative responses of PBMC isolated from one or more human histocompatibility leukocyte antigen (HLA)-typed immune adults. With PBMC from seven nonimmune donors, no proliferative response was observed. Immunodominant regions were found, recognized by PBMC from many volunteers, irrespective of their HLA type. Most of the immunodominant T cell epitopes are located outside the variable regions and, thus, will be conserved among different meningococcal (and gonococcal) strains. Furthermore, the overlapping peptides could be used to identify the epitopes recognized by OMP-specific T cell clones with known HLA restriction. It is interesting that the epitopes defined with the clones occur in highly conserved areas, shared by all neisserial porin proteins. In summary, this analysis of the T cell response to the meningococcal class 1 OMP constitutes a complete study of reactivity to a foreign protein, and illustrates some important features of Ag recognition by T cells. Our data demonstrate unexpected diversity in the T cell recognition of the OMP, and imply that the T cell repertoire against foreign Ag may be greater than previously assumed. This observation is supported by recent data on the interaction of peptide and major histocompatibility complex (MHC) class II, the latter being much less selective than MHC class I. Finally, a comparative analysis pointed out the limitations of algorithms predicting T cell determinants, and the importance of the empirical methodology provided by SMPS.

Adult

Definition of receptor binding domains in interferon-alpha.

Earlier studies from this laboratory had identified three regions in interferon-alpha (IFN-alpha) that influence the active conformation of the molecule. These domains are associated with the amino acid residues 10-35, 78-107, and 123-166. In this report, we define these domains more accurately by identifying their critical clusters of amino acids. Using a panel of IFN-alpha 2a variants in antiviral, growth inhibitory, and receptor binding studies, we are able to show that these three domains, defined by residues 29-35, 78-95, and 123-140, are likely located on the surface of the molecule, with domains 29-35 and 123-140 in close spatial proximity. We conclude that the 29-35 and 123-140 domains are responsible for IFN-alpha receptor binding interactions and constitute receptor recognition sites in IFN-alpha. Extrapolating from our biological activity data, in the context of a number of predictive algorithms that provide insights into the hydrophobicity/hydrophilicity, surface probability, and flexibility of amino acid clusters, we infer that the residues 29-35 influence the active configuration of IFN-alpha most significantly. This region likely represents a loop structure that is relatively rigid in configuration. The carboxy-terminally located strategic domain, 123-140, is comprised of two clusters of amino acid residues, one that forms part of a rigid alpha-helix, the other a more flexible loop structure. Similarly, the 78-95 domain comprises a portion of an alpha-helical structure that is followed by a loop structure. Close examination of the amino acid sequences in all three regions among the different species of IFN-alpha s and human IFN-beta indicate that the 29-35 and 123-140 domains are most highly conserved, yet some variance is apparent in the 78-95 domain. We propose that the 78-95 region influences species specificity among the murine and human IFN-alpha s and determines the differential specificity of action between human IFN-alpha and human IFN-beta.

Amino Acid Sequence

A simple method for predicting the secondary structure of globular proteins: implications and accuracy.

A method is presented for predicting the secondary structure of globular proteins from their amino acid sequence. It is based on a rigorous statistical exploitation of the well-known biological fact that the amino acid compositions of each secondary structure are different. We also propose an evaluation process that allows us to estimate the capacity of a method to predict the secondary structure of a new protein which does not have any homologous proteins whose structure is already known. This evaluation process shows that our method has a prediction accuracy of 58.7% over three states for the 62 proteins of the Kabsch and Sander (1983a) data bank. This result is better than that obtained by the most widely used methods--Lim (1974), Chou and Fasman (1978) and Garnier et al. (1978)--and also than that obtained by a recent method based on local homologies (Levin et al., 1986). Our prediction method is very simple and may be implemented on any microcomputer and even on programmable pocket calculators. A simple Pascal implementation of the method prediction algorithm is given. The interpretation of our results in terms of protein folding and directions for further work are discussed.

Algorithms

Statistical evaluation and biological interpretation of non-random abundance in the E. coli K-12 genome of tetra- and pentanucleotide sequences related to VSP DNA mismatch repair.

The abundance of all tetra- and pentanucleotide sequences is calculated for a set of DNA sequence data comprising 767,393 nucleotides of the E. coli K-12 genome. Observed frequencies are compared to those expected from a Markov chain prediction algorithm. Systematic and extreme non-random representations are found for special sets of sequences. These are interpreted as arising from incorporation of a 2'-deoxyguanosine residue opposite thymidine during replication which, in special sequence contexts, leads to a T/G mismatch that is simultaneously substrate for two competing DNA mismatch repair systems: the mutHLS and the VSP pathway. Processing by the former leads to error correction, by the latter to mutation fixation. The significance of the latter process, as demonstrated here, makes it unlikely that VSP repair has evolved mainly as a mutation avoidance mechanism. It is proposed that in E. coli K-12, VSP repair, together with DNA cytosine methylation, constitutes a mutagenesis/recombination system capable of promoting gene-conversion-like unidirectional transfer of short stretches of DNA sequence.

Algorithms

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans

Primary and predicted secondary structures of the caseins in relation to their biological functions.

In free solution, the caseins behave as non-compact and largely flexible molecules with a high proportion of residues accessible to solvent. Historically, they have been described as random coil-type proteins with only a nutritional function. Nevertheless, secondary structure prediction algorithms indicate that many parts of the (unphosphorylated, unglycosylated) polypeptide chains can form regular structures. In particular, a recurrent motif of the Ca2+-sensitive caseins in man, rat, mouse, guinea pig and ruminant species is an alpha-helix--loop--alpha-helix conformation in which the loop region typically contains a cluster of sites of phosphorylation. The biological function of the caseins is considered and it is suggested that the potential or actual conformations of the group of Ca2+-sensitive caseins are suited to the function of modulating the precipitation of calcium phosphate from solution. Either they can act as sites for nucleation or they can bind rapidly to calcium phosphate nuclei as they form spontaneously from supersaturated solution.

Amino Acid Sequence

Nuclear magnetic resonance studies and molecular dynamics simulations of the solution conformation of a 'designed', alpha-helical peptide.

The complete three-dimensional structure in methanol of an amphipathic alpha-helical peptide, that has been designed by taking into account the three-dimensional structures of small haemolytic peptides, secondary structure prediction algorithms and the well documented literature on alpha-helix stabilizing factors, has been elucidated by two-dimensional NMR spectroscopy. Initially various two-dimensional spectra (COSY, TOCSY, and NOESY) allowed the complete sequence specific assignment of all signals in the 1H spectrum. Consequently trial structures were generated which were then subjected to molecular dynamics simulations using 121 NOE-derived distances and 25 vicinal coupling constant values as structural restraints to give a final set of calculated structures. These structures are in complete agreement with the results of a circular dichroism study and reveal that the peptide adopted a highly ordered alpha-helical conformation. Details of the structure which throw light on future peptide/protein design are discussed.

Amino Acid Sequence

Automatic sleep/wake identification from wrist activity.

The purpose of this study was to develop and validate automatic scoring methods to distinguish sleep from wakefulness based on wrist activity. Forty-one subjects (18 normals and 23 with sleep or psychiatric disorders) wore a wrist actigraph during overnight polysomnography. In a randomly selected subsample of 20 subjects, candidate sleep/wake prediction algorithms were iteratively optimized against standard sleep/wake scores. The optimal algorithms obtained for various data collection epoch lengths were then prospectively tested on the remaining 21 subjects. The final algorithms correctly distinguished sleep from wakefulness approximately 88% of the time. Actigraphic sleep percentage and sleep latency estimates correlated 0.82 and 0.90, respectively, with corresponding parameters scored from the polysomnogram (p < 0.0001). Automatic scoring of wrist activity provides valuable information about sleep and wakefulness that could be useful in both clinical and research applications.

Adult

Aminoglycoside dosing in pediatric patients.

We assessed the performance of a predictive algorithm for dosing aminoglycoside antibiotics in 75 pediatric patients and Bayesian feedback in 36. The absolute errors for peak and trough concentrations were 1.83 and 0.80 micrograms/ml, respectively, which seem clinically acceptable for most patients. However, the algorithm had significant negative bias for both peaks and troughs. Implementation of Bayesian feedback eliminated bias in a second set of concentrations and significantly decreased its magnitude for both peaks (p = 0.028) and troughs (p = 0.005). This method may allow more accurate dosing of aminoglycoside antibiotics in pediatric patients, though it would most likely be improved by better definition of population parameters and their variability.

Algorithms

Definition of murine T helper cell determinants in the major capsid protein of human papillomavirus type 16.

Three murine major histocompatibility complex (MHC) class II-restricted T cell determinants were identified in the major capsid protein L1 of human papillomavirus (HPV) type 16. Peptides derived from HPV-16 L1, which contain putative T cell epitopes located by a predictive algorithm, were synthesized and tested for lymphoproliferative activity by direct immunization, followed by in vitro assay of responses to peptides or recombinant HPV-16 L1. The MHC restriction of the stimulatory peptides was determined using blocking monoclonal antibodies against class II molecules. The responses, which were specific for the priming peptides alone, cross-reacted with recombinant L1 but not with analogous peptides derived from other HPV types.

Amino Acid Sequence

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Journal Article

Complete sequence and model for the A2 subunit of the carotenoid pigment complex, crustacyanin.

The complete sequence has been determined for the A2 subunit of crustacyanin, an astaxanthin-binding protein from the carapace of the lobster Homarus gammarus. The polypeptide chain is 174 residues long and is similar to proteins of the retinol-binding protein superfamily. Some regions of the sequence are most similar to the retinol-binding protein, beta-lactoglobulin subgroup, while the disulphide bonding pattern is more akin to that seen in the porphyrin binding proteins insecticyanin and bilin-binding protein. It is beginning to appear as though this superfamily of proteins, characterized by a similar gross structural framework, may be further subdivided into interrelated subclasses. Model building based on the coordinates of the known structure of human plasma retinol-binding protein and on empirical prediction algorithms has allowed the putative identification of side-chains which line the binding cavity. This pocket is larger than in retinol binding protein and beta-lactoglobulin but does not allow the carotenoid to adopt a folded conformation. The amino acid composition of the pocket does not support a 'charge-shift'-type hypothesis to support the bathochromic shift phenomenon which takes place on interaction of the chromophore with the protein. Instead aromatic side-chains may play a prominent role.

Amino Acid Sequence