PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Predictive Learning Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding↗

Patient-specific models for predicting the outcomes of patients with community acquired pneumonia.

We investigated two patient-specific and four population-wide machine learning methods for predicting dire outcomes in community acquired pneumonia (CAP) patients. Predicting dire outcomes in CAP patients can significantly influence the decision about whether to admit the patient to the hospital or to treat the patient at home. Population-wide methods induce models that are trained to perform well on average on all future cases. In contrast, patient-specific methods specifically induce a model for a particular patient case. We trained the models on a set of 1601 patient cases and evaluated them on a separate set of 686 cases. One patient-specific method performed better than the population-wide methods when evaluated within a clinically relevant range of the ROC curve. Our study provides support for patient-specific methods being a promising approach for making clinical predictions.

Algorithms↗

Prediction of alpha-turns in proteins using PSI-BLAST profiles and secondary structure information.

In this paper a systematic attempt has been made to develop a better method for predicting alpha-turns in proteins. Most of the commonly used approaches in the field of protein structure prediction have been tried in this study, which includes statistical approach "Sequence Coupled Model" and machine learning approaches; i) artificial neural network (ANN); ii) Weka (Waikato Environment for Knowledge Analysis) Classifiers and iii) Parallel Exemplar Based Learning (PEBLS). We have also used multiple sequence alignment obtained from PSIBLAST and secondary structure information predicted by PSIPRED. The training and testing of all methods has been performed on a data set of 193 non-homologous protein X-ray structures using five-fold cross-validation. It has been observed that ANN with multiple sequence alignment and predicted secondary structure information outperforms other methods. Based on our observations we have developed an ANN-based method for predicting alpha-turns in proteins. The main components of the method are two feed-forward back-propagation networks with a single hidden layer. The first sequence-structure network is trained with the multiple sequence alignment in the form of PSI-BLAST-generated position specific scoring matrices. The initial predictions obtained from the first network and PSIPRED predicted secondary structure are used as input to the second structure-structure network to refine the predictions obtained from the first net. The final network yields an overall prediction accuracy of 78.0% and MCC of 0.16. A web server AlphaPred (http://www.imtech.res.in/raghava/alphapred/) has been developed based on this approach.

Amino Acid Sequence↗

A computational model of four regions of the cerebellum based on feedback-error learning.

We propose a computationally coherent model of cerebellar motor learning based on the feedback-error-learning scheme. We assume that climbing fiber responses represent motor-command errors generated by some of the premotor networks such as the feedback controllers at the spinal-, brain stem- and cerebral levels. Thus, in our model, climbing fiber responses are considered to convey motor errors in the motor-command coordinates rather than in the sensory coordinates. Based on the long-term depression in Purkinje cells each corticonuclear microcomplex in different regions of the cerebellum learns to execute predictive and coordinative control of different types of movements. Ultimately, it acquires an inverse model of a specific controlled object and complements crude control by the premotor networks. This general model is developed in detail as a specific neural circuit model for the lateral hemisphere. A new experiment is suggested to elucidate the coordinate frame in which climbing fiber responses are represented.

Animals↗

Child maltreatment, parent alcohol and drug-related problems, polydrug problems, and parenting practices: a test of gender differences and four theoretical perspectives.

The authors tested how adverse childhood experiences (child maltreatment and parent alcohol- and drug-related problems) and adult polydrug use (as a mediator) predict poor parenting in a community sample (237 mothers and 81 fathers). These relationships were framed within several theoretical perspectives, including observational learning, impaired functioning, self-medication, and parentification-pseudomaturity. Structural models revealed that child maltreatment predicted poor parenting practices among mothers. Parent alcohol- and drug-related problems had an indirect detrimental influence on mothers' parenting and practices through self-drug problems. Among fathers, emotional neglect experienced as a child predicted lack of parental warmth more parental neglect, and sexual abuse experienced as a child predicted a rejecting style of parenting.

Adult↗

Biomechanics of atherosclerotic plaque.

Atherosclerosis is the leading cause of death in the U.S. In balloon angioplasty, pressure is applied directly to atherosclerotic plaque to reopen the occluded blood vessel. The mechanical behavior of the plaque often determines the outcome of the angioplasty. Little information on the material properties of atherosclerotic plaque is available, yet the properties govern the plaque's behavior. Our discussion of the experimental testing and numerical analysis of plaque is directed toward summarizing the current knowledge of plaque material properties. Atherosclerotic plaque exhibits a wide range of behaviors consistent with the variability in the underlying composition. Overall, plaques exhibit nonlinear and inelastic mechanical behavior, although geometry and material properties are not well known. The histomorphological composition is critical in determining the plaque's mechanical response. Finite element approximations have been used to study the stresses developed in the diseased vessel; however, material properties are a critical component of a finite element analysis: the predictive capabilities depend on how accurately the material is modeled. When more information on plaque behavior is generated through careful and extensive experimental investigations, better models will be constructed to more accurately predict plaque responses. As the biomechanics community learns about plaque mechanics, we can use the knowledge to enhance the reliability of interventional procedures.

Arteriosclerosis↗

Bio-basis function neural network for prediction of protease cleavage sites in proteins.

The prediction of protease cleavage sites in proteins is critical to effective drug design. One of the important issues in constructing an accurate and efficient predictor is how to present nonnumerical amino acids to a model effectively. As this issue has not yet been paid full attention and is closely related to model efficiency and accuracy, we present a novel neural learning algorithm aimed at improving the prediction accuracy and reducing the time involved in training. The algorithm is developed based on the conventional radial basis function neural networks (RBFNNs) and is referred to as a bio-basis function neural network (BBFNN). The basic principle is to replace the radial basis function used in RBFNNs by a novel bio-basis function. Each bio-basis is a feature dimension in a numerical feature space, to which a nonnumerical sequence space is mapped for analysis. The bio-basis function is designed using an amino acid mutation matrix verified in biology. Thus, the biological content in protein sequences can be maximally utilized for accurate modeling. Mutual information (MI) is used to select the most informative bio-bases and an ensemble method is used to enhance a decision-making process, hence, improving the prediction accuracy further. The algorithm has been successfully verified in two case studies, namely the prediction of Human Immunodeficiency Virus (HIV) protease cleavage sites and trypsin cleavage sites in proteins.

Artificial Intelligence↗

Colorectal Liver Metastasis Pathomics Model: Integrating Single-Cell and Spatial Transcriptome Analysis With Pathomics for Predicting Liver Metastasis in Colorectal Cancer.

The liver is the primary target organ for hematologic metastasis of colorectal cancer (CRC), and CRC liver metastasis (CRLM) often precludes radical resection, making it the leading cause of death in patients with CRC. To improve the identification and prediction of liver metastasis risk, we identified a cell type of liver metastasis--triggering malignant cells (LMTMCs) through integrating single-cell RNA sequencing and spatial transcriptome analysis. Multiomics cell communication analysis indicated that the interaction between fibroblasts and LMTMCs through the COL1A1-CD44/SDC4 and LAMA4-CD44 signaling axes could promote CRLM. By applying the one-class logistic regression algorithm, we developed a CRLM scoring system in the bulk RNA-sequencing data according to the abundance of LMTMCs in each individual. Using the grouping labels derived from the CRLM scoring system in the bulk data and the corresponding whole-slide images without any manual annotations at the region or pixel level, processed via slide-level weakly supervised learning, a deep-learning model based on the ResNet18 architecture, called Colorectal Liver Metastasis Pathomics Model, was developed to predict the risk of liver metastasis in patients with CRC. The Colorectal Liver Metastasis Pathomics Model achieved an area under the curve of 0.84 at the internal test set of The Cancer Genome Atlas-CRC histology images. In the external independent validation sets, namely the Affiliated Hospital of Southwest Medical University and the Affiliated Traditional Chinese Medicine Hospital of Southwest Medical University cohorts, the areas under the curve were 0.89 and 0.72, respectively, indicating effective classification performances. This study provided new insights and tools for the early identification of CRLM and demonstrated the potential of combining multiomics with deep learning-based pathomics in cancer research.

Humans↗

AlphaGenome Enhances Personal Gene Expression Prediction but Retains Key Limitations.

In recent years, numerous genome AI models have been developed to elucidate the relationship between DNA sequence and gene expression. However, these models have faced criticism for their limited accuracy in predicting individual-specific gene expression. AlphaGenome, the current state-of-the-art in genome AI, achieves exceptional performance across a range of sequence-based predictive tasks, but its utility for personal expression prediction has not yet been assessed. In this study, we evaluate AlphaGenome's ability to predict personal gene expression and find that it significantly outperforms its predecessor. Using GTEx data, AlphaGenome improves the prediction of expression direction over Enformer, achieving an odds ratio of 3.0. In some cases, it even reverses previously observed negative correlations into positive ones. Moreover, AlphaGenome demonstrates improved performance for genes with known nonlinear sequence-expression relationships, though it uncovers mechanisms distinct from those identified by tree-based models.

deep learning↗

Improving RNA Secondary Structure Prediction Through Expanded Training Data.

In recent years, deep learning has revolutionized protein structure prediction, achieving remarkable speed and accuracy. RNA structure prediction, however, has lagged behind. Although several methods have shown some success in predicting RNA secondary and tertiary structures, none have reached the accuracy observed with contemporary protein models. The lack of success of these RNA structure prediction models has been proposed to be due to limited high-quality structural information that can be used as training data. To probe this proposed limitation, we developed a large and diverse dataset comprising paired RNA sequences and their corresponding secondary structures. We assess the utility of this enhanced dataset by retraining on a deep learning model, SincFold. We find that SincFold exhibited improved generalization to some previously unseen RNA families, enhancing its capability to predict accurate de novo RNA secondary structures. The RNASSTR dataset provides a substantial advance for RNA structure modeling, laying a strong foundation for the development of future RNA secondary structure prediction algorithms.

Journal Article↗

An associational model of birdsong sensorimotor learning I. Efference copy and the learning of song syllables.

Birdsong learning provides an ideal model system for studying temporally complex motor behavior. Guided by the well-characterized functional anatomy of the song system, we have constructed a computational model of the sensorimotor phase of song learning. Our model uses simple Hebbian and reinforcement learning rules and demonstrates the plausibility of a detailed set of hypotheses concerning sensory-motor interactions during song learning. The model focuses on the motor nuclei HVc and robust nucleus of the archistriatum (RA) of zebra finches and incorporates the long-standing hypothesis that a series of song nuclei, the Anterior Forebrain Pathway (AFP), plays an important role in comparing the bird's own vocalizations with a previously memorized song, or "template." This "AFP comparison hypothesis" is challenged by the significant delay that would be experienced by presumptive auditory feedback signals processed in the AFP. We propose that the AFP does not directly evaluate auditory feedback, but instead, receives an internally generated prediction of the feedback signal corresponding to each vocal gesture, or song "syllable." This prediction, or "efference copy," is learned in HVc by associating premotor activity in RA-projecting HVc neurons with the resulting auditory feedback registered within AFP-projecting HVc neurons. We also demonstrate how negative feedback "adaptation" can be used to separate sensory and motor signals within HVc. The model predicts that motor signals recorded in the AFP during singing carry sensory information and that the primary role for auditory feedback during song learning is to maintain an accurate efference copy. The simplicity of the model suggests that associational efference copy learning may be a common strategy for overcoming feedback delay during sensorimotor learning.

Algorithms↗

Base-rate effects in category learning: a comparison of parallel network and memory storage-retrieval models.

Exemplar-memory and adaptive network models were compared in application to category learning data, with special attention to base rate effects on learning and transfer performance. Subjects classified symptom charts of hypothetical patients into disease categories, with informative feedback on learning trials and with the feedback either given or withheld on test trials that followed each fourth of the learning series. The network model proved notably accurate and uniformly superior to the exemplar model in accounting for the detailed course of learning; both the parallel, interactive aspect of the network model and its particular learning algorithm contribute to this superiority. During learning, subjects' performance reflected both category base rates and feature (symptom) probabilities in a nearly optimal manner, a result predicted by both models, though more accurately by the network model. However, under some test conditions, the data showed substantial base-rate neglect, in agreement with Gluck and Bower (1988b).

Adult↗

Metalearning and neuromodulation.

This paper presents a computational theory on the roles of the ascending neuromodulatory systems from the viewpoint that they mediate the global signals that regulate the distributed learning mechanisms in the brain. Based on the review of experimental data and theoretical models, it is proposed that dopamine signals the error in reward prediction, serotonin controls the time scale of reward prediction, noradrenaline controls the randomness in action selection, and acetylcholine controls the speed of memory update. The possible interactions between those neuromodulators and the environment are predicted on the basis of computational theory of metalearning.

Algorithms↗

Models to predict cardiovascular risk: comparison of CART, multilayer perceptron and logistic regression.

The estimate of a multivariate risk is now required in guidelines for cardiovascular prevention. Limitations of existing statistical risk models lead to explore machine-learning methods. This study evaluates the implementation and performance of a decision tree (CART) and a multilayer perceptron (MLP) to predict cardiovascular risk from real data. The study population was randomly splitted in a learning set (n = 10,296) and a test set (n = 5,148). CART and the MLP were implemented at their best performance on the learning set and applied on the test set and compared to a logistic model. Implementation, explicative and discriminative performance criteria are considered, based on ROC analysis. Areas under ROC curves and their 95% confidence interval are 0.78 (0.75-0.81), 0.78 (0.75-0.80) and 0.76 (0.73-0.79) respectively for logistic regression, MLP and CART. Given their implementation and explicative characteristics, these methods can complement existing statistical models and contribute to the interpretation of risk.

Artificial Intelligence↗

A neural network model with dopamine-like reinforcement signal that learns a spatial delayed response task.

This study investigated how the simulated response of dopamine neurons to reward-related stimuli could be used as reinforcement signal for learning a spatial delayed response task. Spatial delayed response tasks assess the functions of frontal cortex and basal ganglia in short-term memory, movement preparation and expectation of environmental events. In these tasks, a stimulus appears for a short period at a particular location, and after a delay the subject moves to the location indicated. Dopamine neurons are activated by unpredicted rewards and reward-predicting stimuli, are not influenced by fully predicted rewards, and are depressed by omitted rewards. Thus, they appear to report an error in the prediction of reward, which is the crucial reinforcement term in formal learning theories. Theoretical studies on reinforcement learning have shown that signals similar to dopamine responses can be used as effective teaching signals for learning. A neural network model implementing the temporal difference algorithm was trained to perform a simulated spatial delayed response task. The reinforcement signal was modeled according to the basic characteristics of dopamine responses to novel stimuli, primary rewards and reward-predicting stimuli. A Critic component analogous to dopamine neurons computed a temporal error in the prediction of reinforcement and emitted this signal to an Actor component which mediated the behavioral output. The spatial delayed response task was learned via two subtasks introducing spatial choices and temporal delays, in the same manner as monkeys in the laboratory. In all three tasks, the reinforcement signal of the Critic developed in a similar manner to the responses of natural dopamine neurons in comparable learning situations, and the learning curves of the Actor replicated the progress of learning observed in the animals. Several manipulations demonstrated further the efficacy of the particular characteristics of the dopamine-like reinforcement signal. Omission of reward induced a phasic reduction of the reinforcement signal at the time of the reward and led to extinction of learned actions. A reinforcement signal without prediction error resulted in impaired learning because of perseverative errors. Loss of learned behavior was seen with sustained reductions of the reinforcement signal, a situation in general comparable to the loss of dopamine innervation in Parkinsonian patients and experimentally lesioned animals. The striking similarities in teaching signals and learning behavior between the computational and biological results suggest that dopamine-like reward responses may serve as effective teaching signals for learning behavioral tasks that are typical for primate cognitive behavior, such as spatial delayed responding.

Animals↗

Ockham's razor modeling of the matrisome channels of the basal ganglia thalamocortical loops.

A functional model of the basal ganglia-thalamocortical (BTC) loops is described. In our modeling effort, we try to minimize the complexity of our starting hypotheses. For that reason, we call this type of modeling Ockham's razor modeling. We have the additional constraint that the starting assumptions should not contradict experimental findings about the brain. First assumption: The brain lacks direct representation of paths but represents directions (called speed fields in control theory). Then control should be concerned with speed-field tracking (SFT). Second assumption: Control signals are delivered upon differencing in competing parallel channels of the BTC loops. This is modeled by extending SFT with differencing that gives rise to the robust Static and Dynamic State (SDS) feedback-controlling scheme. Third assumption: Control signals are expressed in terms of a gelatinous medium surrounding the limbs. This is modeled by expressing parameters of motion in parameters of the external space. We show that corollaries of the model fit properties of the BTC loops. The SDS provides proper identification of motion related neuronal groups of the putamen. Local minima arise during the controlling process that works in external space. The model explains the presence of parallel channels as the means to avoiding such local minima. Stability conditions of the SDS predict that the initial phase of learning is mostly concerned with selection of sign for the inverse dynamics. The model provides a scalable controller. State description in external space instead of configurational space reduces the dimensionality problem. Falsifying experiment is suggested. Computer experiments demonstrate the feasibility of the approach. We argue that the resulting scheme has a straightforward connectionist representation exhibiting population coding and Hebbian learning properties.

Basal Ganglia↗

PredIL13: Stacking a variety of machine and deep learning methods with ESM-2 language model for identifying IL13-inducing peptides.

Interleukin (IL)-13 has emerged as one of the recently identified cytokine. Since IL-13 causes the severity of COVID-19 and alters crucial biological processes, it is urgent to explore novel molecules or peptides capable of including IL-13. Computational prediction has received attention as a complementary method to in-vivo and in-vitro experimental identification of IL-13 inducing peptides, because experimental identification is time-consuming, laborious, and expensive. A few computational tools have been presented, including the IL13Pred and iIL13Pred. To increase prediction capability, we have developed PredIL13, a cutting-edge ensemble learning method with the latest ESM-2 protein language model. This method stacked the probability scores outputted by 168 single-feature machine/deep learning models, and then trained a logistic regression-based meta-classifier with the stacked probability score vectors. The key technology was to implement ESM-2 and to select the optimal single-feature models according to their absolute weight coefficient for logistic regression (AWCLR), an indicator of the importance of each single-feature model. Especially, the sequential deletion of single-feature models based on the iterative AWCLR ranking (SDIWC) method constructed the meta-classifier consisting of the top 16 single-feature models, named PredIL13, while considering the model's accuracy. The PredIL13 greatly outperformed the-state-of-the-art predictors, thus is an invaluable tool for accelerating the detection of IL13-inducing peptide within the human genome.

Humans↗

Integrating histology and spatial transcriptomics via multimodal transformers and contrastive representation learning for accurate gene expression prediction.

Predicting spatial gene expression from Histological images is a fundamental task in understanding tissue organization and molecular phenotypes. However, existing methods often rely on single-model representations or lack effective alignment between image and transcriptomic features. To address these limitations, we propose a unified multimodal learning framework that integrates histological imaging and spatial transcriptomics through a shared latent representation space. Specifically, histological H&E images are encoded by a ResNet50-based convolutional stem and a MobileViT Transformer backbone to extract hierarchical visual representations. Both modalities are projected into a shared latent space via linear-GELU-dropout transformation blocks, enabling cross-modal alignment through a contrastive learning objective that maximizes agreement between the corresponding image and the spot embeddings. Experimental results on the 10x Genomics Visium dataset of human liver tissue demonstrate that MViTGene achieves significantly higher prediction accuracy than existing methods across multiple gene subsets, with improvements of 20%, 33%, and 12% in predicting marker genes, highly expressed genes, and highly variable genes, respectively. The significant improvement in relevance indicates that the model can more accurately capture the true correspondence between tissue morphology and gene expression, therefore enabling more reliable biological interpretation. It provides a computational tool for high-throughput spatial gene expression prediction that balances performance and interpretability.

Humans↗