PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Predictive Learning Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Dynamic self-efficacy and outcome expectancies: prediction of smoking lapse and relapse.

According to social learning models of drug relapse, decreases in abstinence self-efficacy (ASE) and increases in positive smoking outcome expectancies (POEs) should foreshadow lapses and relapse. In this study, the authors examined this hypothesis by using ecological momentary assessment data from 305 smokers who achieved initial abstinence from smoking and monitored their smoking and their ASE and POEs by using palmtop computers. Daily ASE and POEs predicted the occurrence of a 1st lapse on the following day. Following a lapse, variations in daily ASE predicted the onset of relapse, even after controlling for concurrent smoking. ASE and POEs generally neither mediated nor moderated each other's effects. These data emphasize the role of dynamic factors in the relapse process.

Adult↗

ExoShorkie: predicting RNA-seq coverage of exogenous genomes in yeast by transfer learning.

MOTIVATION: Predicting the RNA-seq coverage of native and exogenous sequences is central to many molecular- and synthetic-biology applications. Substantial progress has been made in developing methods to predict the RNA-seq coverage of native genomic sequences, with the recently developed Shorkie achieving state-of-the-art performance in yeast. However, prediction performance of these methods over exogenous DNA is still unknown. Recent studies measured RNA-seq coverage of large exogenous genomes in yeast, providing a unique opportunity to train machine-learning models on a large exogenous sequence space and to improve both prediction performance and our understanding of regulatory mechanisms. RESULTS: We introduce ExoShorkie, a method we developed by extending Shorkie through transfer learning across multiple exogenous RNA-seq datasets. We demonstrate that ExoShorkie significantly improves prediction performance on held-out exogenous genomes and outperforms both a native-genome-trained Shorkie baseline and Yorzoi, the only competing method for predicting exogenous RNA-seq coverage in yeast, in cross-validation and in leave-one-genome-out evaluations. Furthermore, through interpretability analyses we reveal biologically meaningful regulatory motifs and distinct regulatory rules in exogenous genomes in yeast, providing new insights into transcriptional regulation. AVAILABILITY AND IMPLEMENTATION: ExoShorkie is available at https://github.com/OrensteinLab/ExoShorkie.

Genome, Fungal↗

Impaired learning in mice with abnormal short-lived plasticity.

BACKGROUND: Many studies suggest that long term potentiation (LTP) has a role in learning and memory. In contrast, little is known about the function of short-lived plasticity (SLP). Modeling results suggested that SLP could be responsible for temporary memory storage, as in working memory, or that it may be involved in processing information regarding the timing of events. These models predict that abnormalities in SLP should lead to learning deficits. We tested this prediction in four lines of mutant mice with abnormal SLP, but apparently normal LTP-mice heterozygous for a alpha-calcium calmodulin kinase II mutation (alpha CaMKII +/-) have lower paired-pulse facilitation (PPF) and increased post-tetanic potentiation (PTP); mice lacking synapsin II (SyII-/-), and mice defective in both synapsin I and synapsin II (SyI/II-/-), show normal PPF but lower PTP; in contrast, mice just lacking synapsin I (SyI-/-) have increased PPF, but normal PTP. RESULTS: Our behavioral results demonstrate that alpha CaMKII +/-, SyII-/- and SyI/II-/- mutant mice, which have decreased PPF or PTP, have profound impairments in learning tasks. In contrast, behavioral analysis did not reveal learning deficits in SyI-/- mice, which have increased PPF. CONCLUSIONS: Our results are consistent with models that propose a role for SLP in learning, as mice with decreased PPF or PTP, in the absence of known LTP deficits, also show profound learning impairments. Importantly, analysis of the SyI-/- mutants demonstrated that an increase in PPF does not disrupt learning.

Animals↗

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans↗

Splicing-site recognition of rice (Oryza sativa L.) DNA sequences by support vector machines.

MOTIVATION: It was found that high accuracy splicing-site recognition of rice (Oryza sativa L.) DNA sequence is especially difficult. We described a new method for the splicing-site recognition of rice DNA sequences. METHOD: Based on the intron in eukaryotic organisms conforming to the principle of GT-AG, we used support vector machines (SVM) to predict the splicing sites. By machine learning, we built a model and used it to test the effect of the test data set of true and pseudo splicing sites. RESULTS: The prediction accuracy we obtained was 87.53% at the true 5' end splicing site and 87.37% at the true 3' end splicing sites. The results suggested that the SVM approach could achieve higher accuracy than the previous approaches.

Algorithms↗

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics↗

Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction.

MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.

CRISPR-Cas Systems↗

[Transformation of kinematic characteristics of a precise movement after change in a spatial task].

Brain mechanisms of motor programming were studied with the use of the model of learning precise horizontal elbow flexion. To exclude visual control and make learning to be based, predominantly, on proprioception, experiments were carried out in darkness. The target position was not demonstrated beforehand. Subject (S) had to find an adequate angle of flexion during training with a short light-diode flash which marked the moment of target reaching. Ss were asked to perform a precise horizontal elbow flexion as fast as possible. Movement amplitude, velocity and acceleration were on-line recorded. Ten Ss were divided in two groups. The first group was initially trained to make the precise movement with the preset amplitude of 70 degrees and the second group had to perform similar movement with the amplitude of 55 degrees. Each S was trained to the moment of acquisition of a stable skill (within the 5% error of preset flexion amplitude). After that the target position was unex pectedly changed (from 70 for 55 degree or visa verse). This work was a continuation of our earlier search for a mathematical hypothesis most correctly explaining the central mechanism of motor learning. The dynamics of kinematic characteristics of learning in our experiments fitted well to A. Barto and J. Houk's "Cerebellar Model of Timing and Prediction". A comparison of a computer simulation of this model to the learning characteristics allowed us to make some refinements of the model very important for data analysis possible under conditions of noninvasive investigations. The analysis of acceleration dynamics not considered by the authors of the model made it possible to identify this index with the "pulse phase" similar to the period of LTD of Purkinje cells (the key mechanism of the model). We took such an interpretation as principal in our analysis of experimental data. We analyzed integrals of positive and negative acceleration which made it possible to gain a deeper insight into the physiological mechanism of a replacement of one central command by the other as a consequence of change in spatial task conditions.

Adolescent↗

Acquisition and performance of delayed-response tasks: a neural network model.

We study the time evolution of a neural network model as it learns the three stages of a visual delayed-matching-to-sample (DMS) task: identification of the sample, retention during delay, and matching of sample and target, ignoring distractors. We introduce a neurobiologically plausible, uncommitted architecture, comprising an "executive" subnetwork gating connections to and from a "working" layer. The network learns DMS by reinforcement: reward-dependent synaptic plasticity generates task-dependent behaviour. During learning, working layer cells exhibit stimulus specialization and increased tuning of their firing. The emergence of top-down activity is observed, reproducing aspects of prefrontal cortex control on activity in the visual areas of inferior temporal cortex. We observe a lability of neural systems during learning, with a tendency to encode spurious associations. Executive areas are instrumental during learning to prevent such associations; they are also fundamental for the "mature" network to keep passing DMS. In the mature model, the working layer functions as a short-term memory. The mature system is remarkably robust against cell damage and its performance degrades gracefully as damage increases. The model underlines that executive systems, which regulate the flow of information between working memory and sensory areas, are required for passing tests such as DMS. At the behavioural level, the model makes testable predictions about the errors expected from subjects learning the DMS.

Animals↗

The conceptual basis of function learning and extrapolation: comparison of rule-based and associative-based models.

The purpose of this article is to provide a foundation for a more formal, systematic, and integrative approach to function learning that parallels the existing progress in category learning. First, we note limitations of existing formal theories. Next, we develop several potential formal models of function learning, which include expansion of classic rule-based approaches and associative-based models. We specify for the first time psychologically based learning mechanisms for the rule models. We then present new, rigorous tests of these competing models that take into account order of difficulty for learning different function forms and extrapolation performance. Critically, detailed learning performance was also used to conduct the model evaluations. The results favor a hybrid model that combines associative learning of trained input-prediction pairs with a rule-based output response for extrapolation (EXAM).

Animals↗

Human mutations in high-confidence Tourette disorder genes affect sensorimotor behavior, reward learning, and striatal dopamine in mice.

UNLABELLED: Tourette disorder (TD) is poorly understood, despite affecting 1/160 children. A lack of animal models possessing construct, face, and predictive validity hinders progress in the field. We used CRISPR/Cas9 genome editing to generate mice with mutations orthologous to human de novo variants in two high-confidence Tourette genes, CELSR3 and WWC1 . Mice with human mutations in Celsr3 and Wwc1 exhibit cognitive and/or sensorimotor behavioral phenotypes consistent with TD. Sensorimotor gating deficits, as measured by acoustic prepulse inhibition, occur in both male and female Celsr3 TD models. Wwc1 mice show reduced prepulse inhibition only in females. Repetitive motor behaviors, common to Celsr3 mice and more pronounced in females, include vertical rearing and grooming. Sensorimotor gating deficits and rearing are attenuated by aripiprazole, a partial agonist at dopamine type II receptors. Unsupervised machine learning reveals numerous changes to spontaneous motor behavior and less predictable patterns of movement. Continuous fixed-ratio reinforcement shows Celsr3 TD mice have enhanced motor responding and reward learning. Electrically evoked striatal dopamine release, tested in one model, is greater. Brain development is otherwise grossly normal without signs of striatal interneuron loss. Altogether, mice expressing human mutations in high-confidence TD genes exhibit face and predictive validity. Reduced prepulse inhibition and repetitive motor behaviors are core behavioral phenotypes and are responsive to aripiprazole. Enhanced reward learning and motor responding occurs alongside greater evoked dopamine release. Phenotypes can also vary by sex and show stronger affection in females, an unexpected finding considering males are more frequently affected in TD. SIGNIFICANCE STATEMENT: We generated mouse models that express mutations in high-confidence genes linked to Tourette disorder (TD). These models show sensorimotor and cognitive behavioral phenotypes resembling TD-like behaviors. Sensorimotor gating deficits and repetitive motor behaviors are attenuated by drugs that act on dopamine. Reward learning and striatal dopamine is enhanced. Brain development is grossly normal, including cortical layering and patterning of major axon tracts. Further, no signs of striatal interneuron loss are detected. Interestingly, behavioral phenotypes in affected females can be more pronounced than in males, despite male sex bias in the diagnosis of TD. These novel mouse models with construct, face, and predictive validity provide a new resource to study neural substrates that cause tics and related behavioral phenotypes in TD.

Preprint↗

A neural network model based on the analogy with the immune system.

The similarities between the immune system and the central nervous system lead to the formulation of an unorthodox neural network model. The similarities between the two systems are strong at the system level, but do not seem to be so striking at the level of the components. A new model of a neuron is therefore formulated, in order that the analogy can be used. The essential feature of the hypothetical neuron is that it exhibits hysteresis at the single neuron level. A network of N such neurons is modelled by an N-dimensional system of ordinary differential equations, which exhibits almost 2N attractors. The model has a property that resembles free will. A conjecture concerning how the network might learn stimulus-response behaviour is described. According to the conjecture, learning does not involve modifications of the strengths of synaptic connections. Instead, stimuli ("questions") selectively applied to the network by a "teacher" can be used to take the system to a region of the N-dimensional phase space where the network gives the desired stimulus-response behaviour. A key role for sleep in the learning process is suggested. The model for sleep leads to prediction that the variance in the rates of firing of the neurons associated with memory should increase during waking hours, and decrease during sleep.

Artificial Intelligence↗

Machine learning-based analysis of the impact of 5'&#xa0;untranslated region on protein expression.

The 5' untranslated region (5'UTR) plays a crucial regulatory role in messenger RNA (mRNA), with modified 5'UTRs extensively utilized in vaccine production, gene therapy, etc. Nevertheless, manually optimizing 5'UTRs may encounter difficulties in balancing the effects of various cis-elements. Consequently, multiple 5'UTR libraries have been created, and machine learning models have been employed to analyze and predict translation efficiency (TE) and protein expression, providing insights into critical regulatory features. On the one hand, these screening libraries, based on TE and mean ribosome load, struggle to accurately quantify protein expression; on the other hand, a precise method for quantifying 5'UTRs necessitates a significantly costlier library. To resolve this dilemma, we constructed a library utilizing firefly luciferase as the reporter to measure accurate protein expression. In addition, we optimized the library construction method by clustering mRNA sequences to reduce redundant data and minimize the size of the dataset. This dual strategy by increasing accuracy and reducing dataset size was found to be effective in predicting the 5'UTRs from the PC3 cell line.

5' Untranslated Regions↗

A sequential neural network model for diabetes prediction.

This paper presents a neural network (NN) model to evaluate an existing Health Risk Appraisal (HRA) for diabetes prediction over 3 years (1996-1998) based on a simulated learning algorithm on individual prognostic process, using the repeatedly measured HRAs of 6142 participants. The approach uses a sequential multi-layered perceptron (SMLP) with backpropagation learning, and an explicit model of time-varying inputs along with the sequentially obtained prediction probability, which was obtained by embedding a multivariate logistic function for consecutive years. The study captures the time-sensitive feature of associating risk factors as predictors to the occurrence of diabetes in the corresponding period. This approach outperforms the baseline classification and regression models in terms of gains (average profit: 0.18) and sensitivity (86.04%) for a test data. The result enables a time-sensitive disease prevention and management program as a prospective effort.

Diabetes Mellitus↗

Alcohol-related hazardous behavior among college students.

This chapter compares a social learning and deterrence model for DUI among college students. Our assumption is that deviant behavior, or driving under the influence, is a result of social learning that occurs in ongoing interaction with significant others. A deterrence model that is concerned with the threat and fear of death and beliefs about the capacity of the driver to minimize danger when drunk and statutory commands through laws also play a role in such behavior. Using multivariate analysis, specifically discriminant function and multiple regression, we differentiate our sample into those who have "never," only once, and regularly DUI. The major item in the social learning model contributing to DUI is whether the respondent has ever been a passenger with a drunk driver. The deterrence model also has value since the number of "tricks" the respondent feels are useful in counteracting the influence of alcohol contributes to believed risk reduction. Results indicate that alone, the deterrence model fails to explain DUI. The social learning model with its emphasis on the proximate social environment is needed as a supplement to predict DUI among college students. Efforts to modify this hazardous behavior among college students will need to incorporate a social learning model along with the deterrence model.

Adolescent↗

Towards an integrated protein-protein interaction network: a relational Markov network approach.

Protein-protein interactions play a major role in most cellular processes. Thus, the challenge of identifying the full repertoire of interacting proteins in the cell is of great importance and has been addressed both experimentally and computationally. Today, large scale experimental studies of protein interactions, while partial and noisy, allow us to characterize properties of interacting proteins and develop predictive algorithms. Most existing algorithms, however, ignore possible dependencies between interacting pairs and predict them independently of one another. In this study, we present a computational approach that overcomes this drawback by predicting protein-protein interactions simultaneously. In addition, our approach allows us to integrate various protein attributes and explicitly account for uncertainty of assay measurements. Using the language of relational Markov networks, we build a unified probabilistic model that includes all of these elements. We show how we can learn our model properties and then use it to predict all unobserved interactions simultaneously. Our results show that by modeling dependencies between interactions, as well as by taking into account protein attributes and measurement noise, we achieve a more accurate description of the protein interaction network. Furthermore, our approach allows us to gain new insights into the properties of interacting proteins.

Algorithms↗

Prediction of special education placement from birth certificate data.

The overall goal of this research effort was to develop procedures for accurately identifying children at high risk for special education placement, based on information available at the time of birth. A file containing information on all births in New York City between 1976 and 1986 was matched against the 1992 BIOFILE, which contains information on all children enrolled in the New York City public school system in 1992. A matched file containing birth and school information on 471,165 children resulted from this process. Three sets of risk factors were derived from birth certificate data: parental, pregnancy-related, and child-related. Using these risk factors as independent variables, a survival analysis model was developed predicting special education placement for each of three major disability categories: learning disability, emotional disorder, and mental retardation. A model combining all disability categories was also developed. The significant predictors of special education placement were Medicaid payment for birth (a poverty indicator), unmarried status of mother, large family size, low parental education, a mother born in the United States, a low level of prenatal care, male gender, low birthweight, and a low Apgar score. Male gender was the strongest risk factor in all models. Examination of selected survival curves indicated that the predictive power of the models is substantial. The methodology described in this article can be used to identify at-risk children for whom screening and other early interventions, including preschool programs, may be appropriate.(ABSTRACT TRUNCATED AT 250 WORDS)

Apgar Score↗

Two-sample comparison based on prediction error, with applications to candidate gene association studies.

To take advantage of the increasingly available high-density SNP maps across the genome, various tests that compare multilocus genotypes or estimated haplotypes between cases and controls have been developed for candidate gene association studies. Here we view this two-sample testing problem from the perspective of supervised machine learning and propose a new association test. The approach adopts the flexible and easy-to-understand classification tree model as the learning machine, and uses the estimated prediction error of the resulting prediction rule as the test statistic. This procedure not only provides an association test but also generates a prediction rule that can be useful in understanding the mechanisms underlying complex disease. Under the set-up of a haplotype-based transmission/disequilibrium test (TDT) type of analysis, we find through simulation studies that the proposed procedure has the correct type I error rates and is robust to population stratification. The power of the proposed procedure is sensitive to the chosen prediction error estimator. Among commonly used prediction error estimators, the .632+ estimator results in a test that has the best overall performance. We also find that the test using the .632+ estimator is more powerful than the standard single-point TDT analysis, the Pearson's goodness-of-fit test based on estimated haplotype frequencies, and two haplotype-based global tests implemented in the genetic analysis package FBAT. To illustrate the application of the proposed method in population-based association studies, we use the procedure to study the association between non-Hodgkin lymphoma and the IL10 gene.

Adult↗