PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Computer-aided search for effective antisense RNA target sequences of the human immunodeficiency virus type 1.

For the biological and therapeutic application of antisense nucleic acids, there is a need to identify effective local target regions of given cellular target mRNAs or viral single-stranded nucleic acids. One critical parameter for the effectiveness of antisense nucleic acids could be the potential of intramolecular folding of a given sequence element of the target strand and the antisense strand, respectively. The folding potential of such subsequences was calculated by using an established secondary structure prediction algorithm. For the genomic RNA and the complementary RNA strand of the human immunodeficiency virus type 1 (HIV-1), an energy profile was calculated that monitors the local folding potential of each sequence position surrounded by a window of given length ranging from 50 to 400 nucleotides. The resulting energy profile was compared to the effectiveness of HIV-1-directed antisense RNAs. It was found that significant minima of the local folding potential (high delta G values) correlated with antisense RNA target regions involved in strong inhibition of HIV-1 replication that had been measured independently in two earlier studies by using different experimental approaches. Conversely, antisense RNAs directed against local subregions with a high folding potential (low delta G values) showed weak or no antiviral effect in human cells. The results indicate that analyses of the local folding potential of a given target RNA can support the selection of effective target sequences for antisense RNA.

Antiviral Agents

Definition of receptor binding domains in interferon-alpha.

Earlier studies from this laboratory had identified three regions in interferon-alpha (IFN-alpha) that influence the active conformation of the molecule. These domains are associated with the amino acid residues 10-35, 78-107, and 123-166. In this report, we define these domains more accurately by identifying their critical clusters of amino acids. Using a panel of IFN-alpha 2a variants in antiviral, growth inhibitory, and receptor binding studies, we are able to show that these three domains, defined by residues 29-35, 78-95, and 123-140, are likely located on the surface of the molecule, with domains 29-35 and 123-140 in close spatial proximity. We conclude that the 29-35 and 123-140 domains are responsible for IFN-alpha receptor binding interactions and constitute receptor recognition sites in IFN-alpha. Extrapolating from our biological activity data, in the context of a number of predictive algorithms that provide insights into the hydrophobicity/hydrophilicity, surface probability, and flexibility of amino acid clusters, we infer that the residues 29-35 influence the active configuration of IFN-alpha most significantly. This region likely represents a loop structure that is relatively rigid in configuration. The carboxy-terminally located strategic domain, 123-140, is comprised of two clusters of amino acid residues, one that forms part of a rigid alpha-helix, the other a more flexible loop structure. Similarly, the 78-95 domain comprises a portion of an alpha-helical structure that is followed by a loop structure. Close examination of the amino acid sequences in all three regions among the different species of IFN-alpha s and human IFN-beta indicate that the 29-35 and 123-140 domains are most highly conserved, yet some variance is apparent in the 78-95 domain. We propose that the 78-95 region influences species specificity among the murine and human IFN-alpha s and determines the differential specificity of action between human IFN-alpha and human IFN-beta.

Amino Acid Sequence

A simple method for predicting the secondary structure of globular proteins: implications and accuracy.

A method is presented for predicting the secondary structure of globular proteins from their amino acid sequence. It is based on a rigorous statistical exploitation of the well-known biological fact that the amino acid compositions of each secondary structure are different. We also propose an evaluation process that allows us to estimate the capacity of a method to predict the secondary structure of a new protein which does not have any homologous proteins whose structure is already known. This evaluation process shows that our method has a prediction accuracy of 58.7% over three states for the 62 proteins of the Kabsch and Sander (1983a) data bank. This result is better than that obtained by the most widely used methods--Lim (1974), Chou and Fasman (1978) and Garnier et al. (1978)--and also than that obtained by a recent method based on local homologies (Levin et al., 1986). Our prediction method is very simple and may be implemented on any microcomputer and even on programmable pocket calculators. A simple Pascal implementation of the method prediction algorithm is given. The interpretation of our results in terms of protein folding and directions for further work are discussed.

Algorithms

Statistical evaluation and biological interpretation of non-random abundance in the E. coli K-12 genome of tetra- and pentanucleotide sequences related to VSP DNA mismatch repair.

The abundance of all tetra- and pentanucleotide sequences is calculated for a set of DNA sequence data comprising 767,393 nucleotides of the E. coli K-12 genome. Observed frequencies are compared to those expected from a Markov chain prediction algorithm. Systematic and extreme non-random representations are found for special sets of sequences. These are interpreted as arising from incorporation of a 2'-deoxyguanosine residue opposite thymidine during replication which, in special sequence contexts, leads to a T/G mismatch that is simultaneously substrate for two competing DNA mismatch repair systems: the mutHLS and the VSP pathway. Processing by the former leads to error correction, by the latter to mutation fixation. The significance of the latter process, as demonstrated here, makes it unlikely that VSP repair has evolved mainly as a mutation avoidance mechanism. It is proposed that in E. coli K-12, VSP repair, together with DNA cytosine methylation, constitutes a mutagenesis/recombination system capable of promoting gene-conversion-like unidirectional transfer of short stretches of DNA sequence.

Algorithms

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans

Primary and predicted secondary structures of the caseins in relation to their biological functions.

In free solution, the caseins behave as non-compact and largely flexible molecules with a high proportion of residues accessible to solvent. Historically, they have been described as random coil-type proteins with only a nutritional function. Nevertheless, secondary structure prediction algorithms indicate that many parts of the (unphosphorylated, unglycosylated) polypeptide chains can form regular structures. In particular, a recurrent motif of the Ca2+-sensitive caseins in man, rat, mouse, guinea pig and ruminant species is an alpha-helix--loop--alpha-helix conformation in which the loop region typically contains a cluster of sites of phosphorylation. The biological function of the caseins is considered and it is suggested that the potential or actual conformations of the group of Ca2+-sensitive caseins are suited to the function of modulating the precipitation of calcium phosphate from solution. Either they can act as sites for nucleation or they can bind rapidly to calcium phosphate nuclei as they form spontaneously from supersaturated solution.

Amino Acid Sequence

Nuclear magnetic resonance studies and molecular dynamics simulations of the solution conformation of a 'designed', alpha-helical peptide.

The complete three-dimensional structure in methanol of an amphipathic alpha-helical peptide, that has been designed by taking into account the three-dimensional structures of small haemolytic peptides, secondary structure prediction algorithms and the well documented literature on alpha-helix stabilizing factors, has been elucidated by two-dimensional NMR spectroscopy. Initially various two-dimensional spectra (COSY, TOCSY, and NOESY) allowed the complete sequence specific assignment of all signals in the 1H spectrum. Consequently trial structures were generated which were then subjected to molecular dynamics simulations using 121 NOE-derived distances and 25 vicinal coupling constant values as structural restraints to give a final set of calculated structures. These structures are in complete agreement with the results of a circular dichroism study and reveal that the peptide adopted a highly ordered alpha-helical conformation. Details of the structure which throw light on future peptide/protein design are discussed.

Amino Acid Sequence

Global folding of proteins using a limited number of distance constraints.

A Monte Carlo method is presented which can obtain the correct tertiary fold of a protein given the secondary structure and as few as three interactions between each secondary structure unit. This method was used to fold hemerythrin, flavodoxin, bovine pancreatic trypsin inhibitor and a variable light domain from an immunoglobulin using the known secondary structures of these proteins. Each of the proteins was successfully folded to obtain a structure resembling the initial X-ray structure. Reasonable success was also achieved when using a secondary structure prediction algorithm to assign secondary structure. The r.m.s. deviations between the folded proteins and the crystal structures are in the order of 3-5 A for the backbone coordinates. Evaluation of the r.m.s. deviations between members of the globin family indicates that two equivalent overall folds may have r.m.s. deviations of this or even larger magnitude. The limiting number of constraints necessary to achieve the correct fold is discussed.

Algorithms

Secondary structure formation in model polypeptide chains.

Model polypeptide chains were folded into 3-D compact conformations using distance geometry techniques. Interresidue distances were predicted from the hydrophobicity of the monomers and were refined by repeated projections into lower-dimensional spaces. Main-chain hydrogen bond networks were constructed and propagated through the structure by adjusting local conformations to comply with ideal distance constraints around hydrogen bonds. The resulting folds were compact globules with distinct hydrophobic cores and contained secondary structure elements like real protein molecules. Apart from similarity in appearance, several properties of the model chains were also very close to those of native folded polypeptides. The method in its present form can serve as a starting point for the development of a novel structure prediction algorithm.

Chemical Phenomena

Automatic sleep/wake identification from wrist activity.

The purpose of this study was to develop and validate automatic scoring methods to distinguish sleep from wakefulness based on wrist activity. Forty-one subjects (18 normals and 23 with sleep or psychiatric disorders) wore a wrist actigraph during overnight polysomnography. In a randomly selected subsample of 20 subjects, candidate sleep/wake prediction algorithms were iteratively optimized against standard sleep/wake scores. The optimal algorithms obtained for various data collection epoch lengths were then prospectively tested on the remaining 21 subjects. The final algorithms correctly distinguished sleep from wakefulness approximately 88% of the time. Actigraphic sleep percentage and sleep latency estimates correlated 0.82 and 0.90, respectively, with corresponding parameters scored from the polysomnogram (p < 0.0001). Automatic scoring of wrist activity provides valuable information about sleep and wakefulness that could be useful in both clinical and research applications.

Adult

Aminoglycoside dosing in pediatric patients.

We assessed the performance of a predictive algorithm for dosing aminoglycoside antibiotics in 75 pediatric patients and Bayesian feedback in 36. The absolute errors for peak and trough concentrations were 1.83 and 0.80 micrograms/ml, respectively, which seem clinically acceptable for most patients. However, the algorithm had significant negative bias for both peaks and troughs. Implementation of Bayesian feedback eliminated bias in a second set of concentrations and significantly decreased its magnitude for both peaks (p = 0.028) and troughs (p = 0.005). This method may allow more accurate dosing of aminoglycoside antibiotics in pediatric patients, though it would most likely be improved by better definition of population parameters and their variability.

Algorithms

Definition of murine T helper cell determinants in the major capsid protein of human papillomavirus type 16.

Three murine major histocompatibility complex (MHC) class II-restricted T cell determinants were identified in the major capsid protein L1 of human papillomavirus (HPV) type 16. Peptides derived from HPV-16 L1, which contain putative T cell epitopes located by a predictive algorithm, were synthesized and tested for lymphoproliferative activity by direct immunization, followed by in vitro assay of responses to peptides or recombinant HPV-16 L1. The MHC restriction of the stimulatory peptides was determined using blocking monoclonal antibodies against class II molecules. The responses, which were specific for the priming peptides alone, cross-reacted with recombinant L1 but not with analogous peptides derived from other HPV types.

Amino Acid Sequence

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Journal Article

Complete sequence and model for the A2 subunit of the carotenoid pigment complex, crustacyanin.

The complete sequence has been determined for the A2 subunit of crustacyanin, an astaxanthin-binding protein from the carapace of the lobster Homarus gammarus. The polypeptide chain is 174 residues long and is similar to proteins of the retinol-binding protein superfamily. Some regions of the sequence are most similar to the retinol-binding protein, beta-lactoglobulin subgroup, while the disulphide bonding pattern is more akin to that seen in the porphyrin binding proteins insecticyanin and bilin-binding protein. It is beginning to appear as though this superfamily of proteins, characterized by a similar gross structural framework, may be further subdivided into interrelated subclasses. Model building based on the coordinates of the known structure of human plasma retinol-binding protein and on empirical prediction algorithms has allowed the putative identification of side-chains which line the binding cavity. This pocket is larger than in retinol binding protein and beta-lactoglobulin but does not allow the carotenoid to adopt a folded conformation. The amino acid composition of the pocket does not support a 'charge-shift'-type hypothesis to support the bathochromic shift phenomenon which takes place on interaction of the chromophore with the protein. Instead aromatic side-chains may play a prominent role.

Amino Acid Sequence

Effect of indoor lighting on normal skin.

A small but measurable component of some indoor lighting is ultraviolet radiation (UVR); whether it is sufficient to modify the indoor worker's risk for chronic skin changes is not directly answerable with available technology. A first approach to this question involves a) estimating a range of annual background solar exposure for indoor workers currently at risk; b) determining whether, and at what levels, UVR exposure is a part of specified indoor lighting; and c) calculating the increment in risk implied by a and b. This algorithm predicts that some lighting conditions that meet NIOSH recommended standards would still result in significant increases in the risk of cumulative UVR damage, including skin cancer. More information concerning actual exposure conditions, the relation of spectral effectiveness for luminosity and UVR production, and dose-time reciprocity are required to improve our predictions of long-term cutaneous effects of indoor lighting.

Erythema

Acoustic invariance in speech production: evidence from measurements of the spectral characteristics of stop consonants.

On the basis of theoretical considerations and the results of experiments with synthetic consonant-vowel syllables, it has been hypothesized that the short-time spectrum sampled at the onset of a stop consonant should exhibit gross properties that uniquely specify the consonantal place of articulation independent of the following vowel. The aim of this paper is to test this hypothesis by measuring the spectrum sampled at the onsets and offsets of a large number of consonant-vowel (CV) and vowel-consonant (VC) syllables containing both voiced and voiceless stops produced by several speakers. Templates were devised in an attempt to capture three classes of spectral shapes: diffuse-rising, diffuse-falling, and compact, corresponding to alveolar, labial, and velar consonants, respectively. Spectra were derived from the utterances by sampling at the consonantal release of CV syllables and at the implosion and burst release of VC syllables, and these spectra (smoothed by a linear prediction algorithm) were matched against the templates. It was found that about 85% of the spectra at initial consonant release and at final burst release were correctly classified by the templates, although there was some variability across vowel contexts. The spectra sampled at the implosion were not consistently classified. A preliminary examination of spectra sampled at the release of nasal consonants in CV syllables showed a somewhat lower accuracy of classification by the same templates. Overall, the results support an hypothesis that, in natural speech, the acoustic characteristics of stop consonants, specified in terms of the gross spectral shape sampled at the discontinuity in the acoustic signal, show invariant properties independent of the adjacent vowel or of the voicing characteristics of the consonant. The implication is that the auditory system is endowed with detectors that are sensitive to these kinds of gross spectral shapes, and that the existence of these detectors helps the infant to organize the sounds of speech into their natural classes.

Humans

Stoichiometric flux balance models quantitatively predict growth and metabolic by-product secretion in wild-type Escherichia coli W3110.

Flux balance models of metabolism use stoichiometry of metabolic pathways, metabolic demands of growth, and optimality principles to predict metabolic flux distribution and cellular growth under specified environmental conditions. These models have provided a mechanistic interpretation of systemic metabolic physiology, and they are also useful as a quantitative tool for metabolic pathway design. Quantitative predictions of cell growth and metabolic by-product secretion that are experimentally testable can be obtained from these models. In the present report, we used independent measurements to determine the model parameters for the wild-type Escherichia coli strain W3110. We experimentally determined the maximum oxygen utilization rate (15 mmol of O2 per g [dry weight] per h), the maximum aerobic glucose utilization rate (10.5 mmol of Glc per g [dry weight] per h), the maximum anaerobic glucose utilization rate (18.5 mmol of Glc per g [dry weight] per h), the non-growth-associated maintenance requirements (7.6 mmol of ATP per g [dry weight] per h), and the growth-associated maintenance requirements (13 mmol of ATP per g of biomass). The flux balance model specified by these parameters was found to quantitatively predict glucose and oxygen uptake rates as well as acetate secretion rates observed in chemostat experiments. We have formulated a predictive algorithm in order to apply the flux balance model to describe unsteady-state growth and by-product secretion in aerobic batch, fed-batch, and anaerobic batch cultures. In aerobic experiments we observed acetate secretion, accumulation in the culture medium, and reutilization from the culture medium. In fed-batch cultures acetate is cometabolized with glucose during the later part of the culture period.(ABSTRACT TRUNCATED AT 250 WORDS)

Acetates

Selection criteria for upper gastrointestinal examinations: attempts at improvement.

An attempt was made to improve upon selection criteria for the performance of upper gastrointestinal (UGI) series in three settings: a teaching hospital, a community hospital, and a health maintenance organization. Two statistical techniques, the polychotomous logistic model (to develop predictive algorithms for the identification of specific diseases) and the maximum attainable discrimination technique, were used to show the relationship between the percentage of patients with any disease detected and the percentage of UGI examinations performed. Results showed that neither technique improved significantly upon selection criteria for identifying patients with abnormal UGI series.

Decision Theory