PubMed HealthSearch

SEARCH · PubMed Health

Results for “Prediction Algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Structure from function: screening structural models with functional data.

Structural constraints derived from different antibody epitopes on human growth hormone (hGH) were used to screen three-dimensional models of hGH that were generated by computer algorithms. Previously, alanine-scanning mutagenesis defined the residues that modulate binding to 21 different monoclonal antibodies to hGH. These functional epitopes were composed of 4-14 side chains whose alpha-carbons clustered within 4-23 A. Distance and topographic constraints for these functional epitopes were virtually the same as constraints derived from known x-ray structures of protein-antigen complexes. The constraints were used to evaluate about 1400 models of hGH that were computer-generated by a secondary-structure prediction and packing algorithm. On average each functional epitope reduced the number of models in the pool by a factor of 2, so that 8 monoclonal antibodies could reduce the number of possible models to < 10. The average root-mean-square deviation of alpha-carbon coordinates between the x-ray structure and either the pool of starting models or final models ranged from 13 to 16 A or 4 to 7 A, respectively, depending on the pool of starting models and the level of constraints imposed. All of the final models had the correct folding topography, and the best model was within 3.8 A root-mean-square deviation of the x-ray coordinates. This model was as close as it could have been because the models were built by using ideal helices and those in the x-ray structure are not. Our studies suggest that epitope mapping data can effectively screen structural models and, when coupled to predictive algorithms, can help to generate low-resolution models of a protein.

Amino Acid Sequence

Pioneer in Molecular Biology: Conformational Ensembles in Molecular Recognition, Allostery, and Cell Function.

In 1978, for my PhD, I developed the efficient O(n3) dynamic programming algorithm for the-then open problem of RNA secondary structure prediction. This algorithm, now dubbed the "Nussinov algorithm", "Nussinov plots", and "Nussinov diagrams", is still taught across Europe and the U.S. As sequences started coming out in the 1980s, I started seeking genome-encoded functional signals, later becoming a bioinformatics trend. In the early 1990s I transited to proteins, co-developing a powerful computer vision-based docking algorithm. In the late 1990s, I proposed the foundational role of conformational ensembles in molecular recognition and allostery. At the time, conformational ensembles and free energy landscapes were viewed as physical properties of proteins but were not associated with function. The classical view of molecular recognition and binding was based on only two conformations captured by crystallography: open and closed. I proposed that all conformational states preexist. Proteins always have not one folded form-nor two-but many folded forms. Thus, rather than inducing fit, binding can work by shifting the ensembles between states, and this shifting, or redistributing the ensembles to maintain equilibrium, is the origin of the allosteric effect and protein, thus cell, function. This transformative paradigm impacted community views in allosteric drug design, catalysis, and regulation. Dynamic conformational ensemble shifts are now acknowledged as the origin of recognition, allostery, and signaling, underscoring that conformational ensembles-not proteins-are the workhorses of the cell, pioneering the fundamental idea that dynamic ensembles are the driving force behind cellular processes. Nussinov was recognized as pioneer in molecular biology by JMB.

Molecular Biology

A pre-search estimation algorithm for MEDLINE strategies with qualifiers.

Inexperienced users of online medical databases often have difficulty formulating their queries. Systems designed to assist them usually do not estimate how effective the initial search strategy will be before performing an actual search. Consequently, the search may find an overwhelming number of citations, or retrieve nothing at all. We have developed an estimation algorithm to predict the outcome of a MEDLINE search. The portion of the algorithm described here estimates retrieval for strategies containing qualifiers. In test searches, the estimate reduced the trial-and-error of strategy formulation. However, the accuracy of the estimate fell short of expectations. Our results show that pre-search estimation for strategies with qualifiers cannot be performed effectively with only the occurrence data that is presently available. They further imply that automated search intermediaries can benefit from medical knowledge which expresses the relationships that exist between terms.

Algorithms

Prediction of protein structure from amino acid sequence.

Methods of predicting protein conformation from amino acid sequence are reviewed. Several widely used algorithms to predict local secondary structure are first discussed. Four general approaches to predict the tertiary structure are then described: (1) energy calculations from an open chain; (2) recognition of sequence homology with a known structure; (3) uses of a sequence template that dictates a particular fold; (4) docking alpha-helices and beta-strands. Throughout this review, the likely success of these methods is considered.

Amino Acid Sequence

A comparison of optimal and suboptimal RNA secondary structures predicted by free energy minimization with structures determined by phylogenetic comparison.

This article describes the latest version of an RNA folding algorithm that predicts both optimal and suboptimal solutions based on free energy minimization. A number of RNA's with known structures deduced from comparative sequence analysis are folded to test program performance. The group of solutions obtained for each molecule is analysed to determine how many of the known helixes occur in the optimal solution and in the best suboptimal solution. In most cases, a structure about 80% correct is found with a free energy within 2% of the predicted lowest free energy structure.

Algorithms

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense

A dynamic programming algorithm for finding alternative RNA secondary structures.

Dynamic programming algorithms that predict RNA secondary structure by minimizing the free energy have had one important limitation. They were able to predict only one optimal structure. Given the uncertainties of the thermodynamic data and the effects of proteins and other environmental factors on structure, the optimal structure predicted by these methods may not have biological significance. We present a dynamic programming algorithm that can determine optimal and suboptimal secondary structures for an RNA. The power and utility of the method is demonstrated in the folding of the intervening sequence of the rRNA of Tetrahymena. By first identifying the major secondary structures corresponding to the lowest free energy minima, a secondary structure of possible biological significance is derived.

Animals

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

[Clinico-mathematical methods of diagnosis, prognosis of outcome and selection of optimal variant of treatment of myocardial infarction during hospital stay].

The authors have developed methods for diagnosing and forecasting myocardial infarction outcomes with the use of current mathematical approaches. Studied the problems of simulating appropriate conditions and processes. Developed some effective algorithms of predicting myocardial infarction outcomes on the basis of the models obtained. Using the approaches (algorithms) suggested the authors have reviewed the methods for finding optimal correction of myocardial infarction treatment, based on the idea of maximization of the survival probability as one of the outcomes of the disease in question. The computer-aided methods for diagnosing, forecasting outcomes and choice of optimal treatment tactics for myocardial infarction during stay at a hospital can be very instrumental in raising the efficacy of the treatment and diagnosis of the disease under consideration.

Adult

A simple technic for predicting daily maintenance dose of warfarin.

Warfarin anticoagulation data from forty-seven patients studied retrospectively was used to develop an algorithm for predicting daily maintenance dose of warfarin. A plot of prothrombin time versus cumulative warfarin loading dose was made for each of the forty-seven patients and the aera under the curve (AREA) measured from the baseline prothrombin time to a prothrombin time of 20 seconds. A linear correlation between the established daily maintenance dose of warfarin and AREA was derived regression analysis: Daily maintenance dose (mg/day) = 0.0465 x (AREA) + 1.5. The correlation was then used to predict daily maintenance dose in twenty-four patients studied prospectively. The mean prothrombin time for a seven day stabilization period after loading for all patients in the prospective study was 21.5 +/- 2.2 seconds and the seven day mean prothrombin time for each patient fell between 19 and 24 seconds. The results of the prospective study indicate that this technic is useful in the early, accurate prediction of a daily maintenance dose of warfarin.

Adult

Protein structure prediction: selecting salient features from large candidate pools.

We introduce a parallel approach, "DT-SELECT," for selecting features used by inductive learning algorithms to predict protein secondary structure. DT-SELECT is able to rapidly choose small, nonredundant feature sets from pools containing hundreds of thousands of potentially useful features. It does this by building a decision tree, using features from the pool, that classifies a set of training examples. The features included in the tree provide a compact description of the training data and are thus suitable for use as inputs to other inductive learning algorithms. Empirical experiments in the protein secondary-structure task, in which sets of complex features chosen by DT-SELECT are used to augment a standard artificial neural network representation, yield surprisingly little performance gain, even though features are selected from very large feature pools. We discuss some possible reasons for this result.

Algorithms

Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders.

Advancements in whole genome sequencing have increased the number of variants of uncertain significance (VUS) identified in patient genomes. This has created a diagnostic bottleneck for genetic counselors tasked with sifting through these variants and determining those most likely to be causative for a patient's clinical presentation. Machine learning (ML) tools can aid in identifying pathogenic variants from VUS, but there is a need for gene-specific algorithms that predict pathogenic variants with high accuracy. To address this need, we present a workflow for developing gene-specific, ensemble-learning ML tools, that leverage outputs from other algorithms, locations of variants within the gene, and evolutionary conservation data to make a prediction of pathogenicity. Variants in SMARCA2 and SMARCA4 that are associated with rare neurodevelopmental diseases were used to screen 15 ML algorithms. A random forest learner was tuned to yield a final accuracy of 0.93 on holdout data. Generalizing this predictor to other BAF complex proteins resulted in a sharp decline in performance. We trained a final predictor for all genes in the study to create a predictor that identifies pathogenic variants in these BAF subunits with an accuracy of 0.91 on holdout data. This predictor specific to BAF complex proteins performs with higher accuracy and AUROC than any other predictor. The decline in performance when generalized to other proteins emphasizes the need for the gene-specific calibration of predictors. Our workflow for the development of such models provides a quick, computationally inexpensive route for improving the ML tools available to genetic counselors.

Journal Article

Predictive Models for Hypoglycemia Risk in Haemodialysis Patients With Diabetic Kidney Disease: Systematic Review and Meta-Analysis.

AIM: To provide evidence for selecting and developing reliable clinical assessment tools for hypoglycemia in diabetic kidney disease patients during haemodialysis. DESIGN: Review. METHODS: Systematic searches were performed in 9 Chinese and English databases to collect literature regarding the development of hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease. Two reviewers independently performed literature screening, data extraction, risk-of-bias assessment, and applicability evaluation. The Prediction Model Risk of Bias Assessment Tool was used to assess the risk of bias and applicability of the included studies. Meta-analysis was conducted using R software. DATA SOURCES: CNKI, Wanfang, VIP, CBM, PubMed, Cochrane Library, EMbase, Web of Science, and CINAHL. The search period covered from the establishment date of each database to December 2025. RESULTS: Six studies, comprising six prediction models, were included. Two studies performed internal validation, and three conducted external validation. All models reported the area under the curve, ranging from 0.813 to 0.866, and calibration measures. Four studies were rated as having a high risk of bias, while all six demonstrated good overall applicability. The meta-analysis showed that the pooled AUC value of the six studies was 0.846 (95% CI: 0.823-0.867). CONCLUSION: Research on hypoglycemia risk prediction models in haemodialysis patients with diabetic kidney disease remains in the developmental stage. Although the included prediction models exhibited satisfactory apparent discriminatory ability and clinical applicability, most of the original studies suffered from a high risk of bias and lacked adequate validation. The true predictive performance and clinical application value of these models remain to be further verified. Accordingly, routine and unconditional clinical application is not recommended at this stage. Future studies should include more high-quality, multicenter external validation and develop models with high generalizability, favourable clinical applicability, and robust predictive performance to facilitate early identification of hypoglycemia risk in this population. IMPACT: This study systematically evaluated the hypoglycemia risk prediction models for diabetic kidney disease patients during haemodialysis, and the research on hypoglycemia risk prediction models for maintenance haemodialysis patients during dialysis is still in the development stage. This study provides a reference for clinical medical staff to select or develop hypoglycemia risk prediction and assessment tools for diabetic kidney disease patients during haemodialysis. REPORTING METHOD: This study was conducted in accordance with the relevant guidelines of the EQUATOR Network and followed the TRIPOD-SRMA Checklist. PATIENT OR PUBLIC CONTRIBUTION: No patient or public contribution. TRIAL REGISTRATION: PROSPERO: CRD420251243352.

Humans

An algorithm for the bonding-probability map of nucleic acid secondary structure.

We report a more efficient and well-defined algorithm for predicting a secondary structure of single-stranded nucleic acid from a primary nucleotide sequence. Using this algorithm, one- and two-dimensional bonding-probability maps of 5S rRNA of thermus thermophilus HB8 were calculated. These maps well express the stability of the secondary structure.

Chemical Phenomena

Prediction of the three-dimensional structure of human growth hormone.

In recent years, the protein-folding problem has attracted the attention of molecular biologists. Efforts have focused on developing heuristic and energy-based algorithms to predict the three-dimensional structure of a protein from its amino acid sequence. We have applied a series of heuristic algorithms to the sequence of human growth hormone. A family of five structures which are generically right-handed fourfold alpha-helical bundles are found from an investigation of approximately 10(8) structures. A plausible receptor binding site is suggested. Independent crystallographic analysis confirms some aspects of these predictions. These methods only deal with the "core" structure, and conformations of many residues are not defined. Further work is required to identify a unique set of coordinates and to clarify the topological alternative available to alpha-helical proteins.

Growth Hormone

The rational design of highly stable, amphiphilic helical peptides.

A computer algorithm was devised for the evaluation of helical stability of potentially amphiphilic peptide sequences of specified length containing a set number of leucines in the hydrophobic region. All possible combinations of Glu, Lys and Gln in the hydrophilic region are rated using a set of empirical rules for salt bridge formation in alpha-helices, and the sequences which rate the highest are displayed. The rules for salt bridge formation were largely derived from published studies on the effects of salt bridges on helical stability. The algorithm was tested by redesigning a known amphiphilic alpha-helical peptide, alpha 1B or 1, which has been shown to aggregate into four-helix bundles. Comparison of the circular dichroism spectra of two peptides, 2 and 3, to 1 demonstrated that the redesigned peptide with the highest priority score from the algorithm, 2, was more helical when aggregated and slightly more helical as a monomer, whereas the peptide with the low priority score, 3, was somewhat less helical when aggregated and much less helical when monomeric. These results support the design of the algorithm, although conclusions based on aggregation data are complicated by the importance of interhelix contacts in the bundle. Further studies are underway to examine the reliability of the algorithm's predictions regarding the design of other helical peptides.

Algorithms

Seeing "ghost" planes in stereo vision.

I have studied particular ambiguous random dot stereograms where multiple matches (that are equally possible) are available at each point. The human visual system resolves these ambiguities in two qualitatively different ways. In some cases a few transparent surfaces are perceived corresponding to all the ambiguous matches. In other cases a single dominant opaque surface is perceived. The conditions under which each behavior occurs are described. Additional experiments, designed to explore whether a number of modified stereo matching algorithms can predict human perception, are described, and their theoretical implications are discussed.

Algorithms

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue C&#x3b1;-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins