PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Predictive Learning Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Knowledge acquisition, accessibility, and use in person perception and stereotyping: simulation with a recurrent connectionist network.

Connectionist models contrast in many ways with the symbolic models that have traditionally been applied within social psychology. In this article the authors apply an autoassociative connectionist model originally developed by J. L. McClelland and D. E. Rumelhart (1986) to reproduce several well-replicated and theoretically important phenomena related to person perception and stereotyping. These phenomena are exemplar-based inference, group-based stereotyping, the simultaneous application of several stereotypes to generate emergent characteristics, and the effects of recency and frequency of prior exposures on accessibility (the probability of a representation's use). Though many of these phenomena are explained by current theories in social psychology, the simulation contributes to parsimony and theoretical integration by showing that a single, very simple mechanism can generate them all. The model also predicts a new phenomenon--rapid recovery of accessibility after it has declined to zero.

Computer Simulation↗

On the use of machine learning to identify topological rules in the packing of beta-strands.

The machine learning program GOLEM was applied to discover topological rules in the packing of beta-sheets in alpha/beta-domain proteins. Rules (constraints) were determined for four features of beta-sheet packing: (i) whether a beta-strand is at an edge; (ii) whether two consecutive beta-strands pack parallel or anti-parallel; (iii) whether two beta-strands pack adjacently; and (iv) the winding direction of two consecutive beta-strands. Rules were found with high predictive accuracy and coverage. The errors were generally associated with complications in domain folds, especially in one doubly would domains. Investigation of the rules revealed interesting patterns, some of which were known previously, others that are novel. Novel features include (i) the relationship between pairs of sequential strands is in general one of decreasing size; (ii) more sequential pairs of strands wind in the direction out than in; and (iii) it takes a larger alteration in hydrophobicity to change a strand from winding in the direction out than in. These patterns in the data may be the result of folding pathways in the domains. The rules found are of predictive value and could be used in the combinatorial prediction of protein structure, or as a general test of model structures, e.g. those produced by threading. We conclude that machine learning has a useful role in the analysis of protein structures.

Amino Acid Sequence↗

A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites.

Genome-scale modeling of microbiome metabolism enables the simulation of diet-host-microbiome-disease interactions. However, current genome-scale reconstruction resources are limited in scope by computational challenges. We developed an optimized and highly parallelized reconstruction and analysis pipeline to build a resource of 247,092 microbial genome-scale metabolic reconstructions, deemed APOLLO. APOLLO spans 19 phyla, contains >60% of uncharacterized strains, and accounts for strains from 34 countries, all age groups, and multiple body sites. Using machine learning, we predicted with high accuracy the taxonomic assignment of strains based on the computed metabolic features. We then built 14,451 metagenomic sample-specific microbiome community models to systematically interrogate their community-level metabolic capabilities. We show that sample-specific metabolic pathways accurately stratify microbiomes by body site, age, and disease state. APOLLO is freely available, enables the systematic interrogation of the metabolic capabilities of largely still uncultured and unclassified species, and provides unprecedented opportunities for systems-level modeling of personalized host-microbiome co-metabolism.

Humans↗

Structure-based prediction of bZIP partnering specificity.

Predicting protein interaction specificity from sequence is an important goal in computational biology. We present a model for predicting the interaction preferences of coiled-coil peptides derived from bZIP transcription factors that performs very well when tested against experimental protein microarray data. We used only sequence information to build atomic-resolution structures for 1711 dimeric complexes, and evaluated these with a variety of functions based on physics, learned empirical weights or experimental coupling energies. A purely physical model, similar to those used for protein design studies, gave reasonable performance. The results were improved significantly when helix propensities were used in place of a structurally explicit model to represent the unfolded reference state. Further improvement resulted upon accounting for residue-residue interactions in competing states in a generic way. Purely physical structure-based methods had difficulty capturing core interactions accurately, especially those involving polar residues such as asparagine. When these terms were replaced with weights from a machine-learning approach, the resulting model was able to correctly order the stabilities of over 6000 pairs of complexes with greater than 90% accuracy. The final model is physically interpretable, and suggests specific pairs of residues that are important for bZIP interaction specificity. Our results illustrate the power and potential of structural modeling as a method for predicting protein interactions and highlight obstacles that must be overcome to reach quantitative accuracy using a de novo approach. Our method shows unprecedented performance in predicting protein-protein interaction specificity accurately using structural modeling and suggests that predicting coiled-coil interactions generally may be within reach.

Basic-Leucine Zipper Transcription Factors↗

A probabilistic classification system for predicting the cellular localization sites of proteins.

We have defined a simple model of classification which combines human provided expert knowledge with probabilistic reasoning. We have developed software to implement this model and have applied it to the problem of classifying proteins into their various cellular localization sites based on their amino acid sequences. Since our system requires no hand tuning to learn training data, we can now evaluate the prediction accuracy of protein localization sites by a more objective cross-validation method than earlier studies using production rule type expert systems. 336 E. coli proteins were classified into 8 classes with an accuracy of 81% while 1484 yeast proteins were classified into 10 classes with an accuracy of 55%. Additionally we report empirical results using three different strategies for handling continuously valued variables in our probabilistic reasoning system.

Cells↗

Learning to predict protein-protein interactions from protein sequences.

In order to understand the molecular machinery of the cell, we need to know about the multitude of protein-protein interactions that allow the cell to function. High-throughput technologies provide some data about these interactions, but so far that data is fairly noisy. Therefore, computational techniques for predicting protein-protein interactions could be of significant value. One approach to predicting interactions in silico is to produce from first principles a detailed model of a candidate interaction. We take an alternative approach, employing a relatively simple model that learns dynamically from a large collection of data. In this work, we describe an attraction-repulsion model, in which the interaction between a pair of proteins is represented as the sum of attractive and repulsive forces associated with small, domain- or motif-sized features along the length of each protein. The model is discriminative, learning simultaneously from known interactions and from pairs of proteins that are known (or suspected) not to interact. The model is efficient to compute and scales well to very large collections of data. In a cross-validated comparison using known yeast interactions, the attraction-repulsion method performs better than several competing techniques.

Algorithms↗

Refining sequence-to-expression modelling with chromatin accessibility.

MOTIVATION: Sequence-to-expression models typically do not consider chromatin accessibility, a major factor limiting gene regulation. We hypothesized that supplying accessibility as an input feature would allow a sequence-to-expression model to focus on important open regions of the genome. RESULTS: We found that the performance of such an augmented model was significantly better than that of sequence-only or accessibility-only models with similar architectures. Specifically, its ability to predict the expression of highly variable genes and gene expression in other cell types improved, and higher attribution scores in the input DNA sequences of the augmented model conformed to accessibility, enabling the learning of cell type-specific sequence patterns. Additionally, we show that fine-tuning a pre-trained sequence-only model with both sequence and accessibility can boost performance further and highlight the importance of sequencing depth in sequence-to-expression prediction. AVAILABILITY AND IMPLEMENTATION: Source code is available on GitHub at https://github.com/lapohosorsolya/accessible_seq2exp.

Chromatin↗

Gene selection for classification of cancers using probabilistic model building genetic algorithm.

Recently, DNA microarray-based gene expression profiles have been used to correlate the clinical behavior of cancers with the differential gene expression levels in cancerous and normal tissues. To this end, after selection of some predictive genes based on signal-to-noise (S2N) ratio, unsupervised learning like clustering and supervised learning like k-nearest neighbor (k NN) classifier are widely used. Instead of S2N ratio, adaptive searches like Probabilistic Model Building Genetic Algorithm (PMBGA) can be applied for selection of a smaller size gene subset that would classify patient samples more accurately. In this paper, we propose a new PMBGA-based method for identification of informative genes from microarray data. By applying our proposed method to classification of three microarray data sets of binary and multi-type tumors, we demonstrate that the gene subsets selected with our technique yield better classification accuracy.

Algorithms↗

Quantifying generalization from trial-by-trial behavior of adaptive systems that learn with basis functions: theory and experiments in human motor control.

During reaching movements, the brain's internal models map desired limb motion into predicted forces. When the forces in the task change, these models adapt. Adaptation is guided by generalization: errors in one movement influence prediction in other types of movement. If the mapping is accomplished with population coding, combining basis elements that encode different regions of movement space, then generalization can reveal the encoding of the basis elements. We present a theory that relates encoding to generalization using trial-by-trial changes in behavior during adaptation. We consider adaptation during reaching movements in various velocity-dependent force fields and quantify how errors generalize across direction. We find that the measurement of error is critical to the theory. A typical assumption in motor control is that error is the difference between a current trajectory and a desired trajectory (DJ) that does not change during adaptation. Under this assumption, in all force fields that we examined, including one in which force randomly changes from trial to trial, we found a bimodal generalization pattern, perhaps reflecting basis elements that encode direction bimodally. If the DJ was allowed to vary, bimodality was reduced or eliminated, but the generalization function accounted for nearly twice as much variance. We suggest, therefore, that basis elements representing the internal model of dynamics are sensitive to limb velocity with bimodal tuning; however, it is also possible that during adaptation the error metric itself adapts, which affects the implied shape of the basis elements.

Adaptation, Physiological↗

Hebbian learning of context in recurrent neural networks.

Single electrode recording in the inferotemporal cortex of monkeys during delayed visual memory tasks provide evidence for attractor dynamics in the observed region. The persistent elevated delay activities could be internal representations of features of the learned visual stimuli shown to the monkey during training. When uncorrelated stimuli are presented during training in a fixed sequence, these experiments display significant correlations between the internal representations. Recently a simple model of attractor neural network has reproduced quantitatively the measured correlations. An underlying assumption of the model is that the synaptic matrix formed during the training phase contains in its efficacies information about the contiguity of persistent stimuli in the training sequence. We present here a simple unsupervised learning dynamics that produces such a synaptic matrix if sequences of stimuli are repeatedly presented to the network at fixed order. The resulting matrix is then shown to convert temporal correlations during training into spatial correlations between attractors. The scenario is that, in the presence of selective delay activity, at the presentation of each stimulus, the activity distribution in the neural assembly contain information of both the current stimulus and the previous one (carried by the attractor). Thus the recurrent synaptic matrix can code not only for each of the stimuli presented to the network but also for their context. We combine the idea that for learning to be effective, synaptic modification should be stochastic, with the fact that attractors provide learnable information about two consecutive stimuli. We calculate explicitly the probability distribution of synaptic efficacies as a function of training protocol, that is, the order in which stimuli are presented to the network. We then solve for the dynamics of a network composed of integrate-and-fire excitatory and inhibitory neurons with a matrix of synaptic collaterals resulting from the learning dynamics. The network has a stable spontaneous activity, and stable delay activity develops after a critical learning stage. The availability of a learning dynamics makes possible a number of experimental predictions for the dependence of the delay activity distributions and the correlations between them, on the learning stage and the learning protocol. In particular it makes specific predictions for pair-associates delay experiments.

Animals↗

Statistical model of natural stimuli predicts edge-like pooling of spatial frequency channels in V2.

BACKGROUND: It has been shown that the classical receptive fields of simple and complex cells in the primary visual cortex emerge from the statistical properties of natural images by forcing the cell responses to be maximally sparse or independent. We investigate how to learn features beyond the primary visual cortex from the statistical properties of modelled complex-cell outputs. In previous work, we showed that a new model, non-negative sparse coding, led to the emergence of features which code for contours of a given spatial frequency band. RESULTS: We applied ordinary independent component analysis to modelled outputs of complex cells that span different frequency bands. The analysis led to the emergence of features which pool spatially coherent across-frequency activity in the modelled primary visual cortex. Thus, the statistically optimal way of processing complex-cell outputs abandons separate frequency channels, while preserving and even enhancing orientation tuning and spatial localization. As a technical aside, we found that the non-negativity constraint is not necessary: ordinary independent component analysis produces essentially the same results as our previous work. CONCLUSION: We propose that the pooling that emerges allows the features to code for realistic low-level image features related to step edges. Further, the results prove the viability of statistical modelling of natural images as a framework that produces quantitative predictions of visual processing.

Models, Statistical↗

Predator learning favours mimicry of a less-toxic model in poison frogs.

Batesian mimicry--resemblance of a toxic model by an edible mimic--depends on deceiving predators. Mimetic advantage is considered to be dependent on frequency because an increase in mimic abundance leads to breakdown of the warning signal. Where multiple toxic species are available, batesian polymorphism is predicted--that is, mimics diversify to match sympatric models. Despite the prevalence of batesian mimicry in nature, batesian polymorphism is relatively rare. Here we explore a poison-frog mimicry complex comprising two parapatric models and a geographically dimorphic mimic that shows monomorphism where models co-occur. Contrary to classical predictions, our toxicity assays, field observations and spectral reflectances show that mimics resemble the less-toxic and less-abundant model. We examine "stimulus generalization" as a mechanism for this non-intuitive result with learning experiments using naive avian predators and live poison frogs. We find that predators differ in avoidance generalization depending on toxicity of the model, conferring greater protection to mimics resembling the less-toxic model owing to overlap of generalized avoidance curves. Our work supports a mechanism of toxicity-dependent stimulus generalization, revealing an additional solution for batesian mimicry where multiple models coexist.

Animals↗

Relationship of neuropsychological and MRI measures to age of onset of schizophrenia.

Age of onset of schizophrenia (AOS) may be largely determined by neurobiological factors. We examined in a diverse sample of schizophrenia out-patients the relationships of AOS with neuropsychological abilities and structural brain abnormalities as measured on cerebral magnetic resonance imaging (MRI). A total of 82 out-patients meeting DSM-III-R criteria for schizophrenia were evaluated with a comprehensive neuropsychological battery and semi-automated quantitatively analysed cerebral MRI. Earlier AOS correlated with poorer performance in learning and abstraction/cognitive flexibility, and with larger volumes of caudate and lenticular nuclei, and smaller volume of thalamus on MRI. A model for predicting AOS consisting of abstraction and thalamic and caudate volumes remained significant after controlling for duration of illness, current age and daily neuroleptic dose. In conclusion, AOS may be related to specific rather than general measures of cognitive performance and structural brain abnormalities.

Adult↗

Deep DNA and protein level feature integration for robust clinical variant interpretation using probabilistic gradient boosting.

A major challenge in clinical genomics is to classify genetic variations correctly, since it directly affects disease diagnosis and personal care. The existing methods tend to be based on the combination of different factors, such as protein structure, population frequencies, phenotypic annotations, and sequence conservation. Nevertheless, these methods often cannot be used to achieve the necessary interpretability, quantify uncertainty, and address rare cases. This paper presents a probabilistic gradient boosting model on variant pathogenicity prediction. The suggested framework applies biological characteristics at both level of DNA and protein levels while also scaling the level of uncertainty in clinical decision making. Our machine learning aims to solve the issues of variant interpretation by managing the features and through probability-based pathogenicity prediction. The framework formulation is aimed at generalizing over various datasets and minimizing overfitting. At the same time, it can ensure reasonable performance to facilitate clinical experiments. The model has also been tested on three standard datasets and demonstrated to be more predictive of the pathogenic effect of variants, in comparison with a variety of existing tools. The probabilistic gradient boosting model proposed had ROC AUC values of 0.9293, 0.9610, and 0.9646 on ClinVar variants, GRCh37, and GRCh38 human genome respectively. Furthermore, the dataset was ensured to include both exonic and intronic variants, and Variants of Uncertain Significance were also taken into consideration for Performance Testing. Through this it also aims to provide better clinical significance which will lead to a good interpretable tool for priority of variants for a large variety of disease conditions.

ClinVar↗

The antisaccade task in a sample of 2,006 young males. II. Effects of task parameters.

Antisaccade performance was investigated in a sample of 2,006 young males as part of a large epidemiological study investigating psychosis proneness. This report summarizes the effects of task parameters on performance using a sample of 55,678 antisaccade trials collected from a subpopulation of 947 individuals. Neither the amplitude nor the latency of an error prosaccade in the antisaccade task was correlated with the latency of the ensuing corrective antisaccade that almost always followed an error. However, the latency of the corrective antisaccade decreased with increasing stimulus distance. Concerning the effects of specific task parameters, trials with stimuli closer to the central fixation point and trials preceded by shorter fixation intervals resulted in more errors and longer latencies for the antisaccades. Finally, there were learning and fatigue effects reflected mainly in the error rate, which was greater at the beginning and at the end of the 5-min task. We used a model to predict whether an error or a correct antisaccade would follow a particular trial. All task parameters were significant predictors of the trial outcome but their power was negligible. However, when modeled alone, response latency of the first movement predicted 40% of errors. In particular, the smaller this latency was, the higher the probability of an error. These findings are discussed in light of current hypotheses on antisaccade production mechanisms involving mainly the superior colliculus.

Adolescent↗

Analysis of respiratory pressure-volume curves in intensive care medicine using inductive machine learning.

We present a case study of machine learning and data mining in intensive care medicine. In the study, we compared different methods of measuring pressure-volume curves in artificially ventilated patients suffering from the adult respiratory distress syndrome (ARDS). Our aim was to show that inductive machine learning can be used to gain insights into differences and similarities among these methods. We defined two tasks: the first one was to recognize the measurement method producing a given pressure-volume curve. This was defined as the task of classifying pressure-volume curves (the classes being the measurement methods). The second was to model the curves themselves, that is, to predict the volume given the pressure, the measurement method and the patient data. Clearly, this can be defined as a regression task. For these two tasks, we applied C5.0 and CUBIST, two inductive machine learning tools, respectively. Apart from medical findings regarding the characteristics of the measurement methods, we found some evidence showing the value of an abstract representation for classifying curves: normalization and high-level descriptors from curve fitting played a crucial role in obtaining reasonably accurate models. Another useful feature of algorithms for inductive machine learning is the possibility of incorporating background knowledge. In our study, the incorporation of patient data helped to improve regression results dramatically, which might open the door for the individual respiratory treatment of patients in the future.

Adult↗

Paternally Expressed Gene 10 Promoter Methylation Level as a Predictor of HBeAg Seroconversion in Chronic Hepatitis B Patients.

The management of chronic hepatitis B (CHB) encounters challenges like suboptimal antiviral response and the lack of predictive biomarkers. In this study, the role of paternally expressed gene 10 (PEG10) in hepatitis B e antigen (HBeAg) seroconversion (HBeAg SC) was explored to identify a therapeutic target and predictive model. In total, 349 participants were recruited, and 141 HBeAg-positive patients were followed up after 48 weeks of antiviral therapy. Key genes were screened by machine learning algorithms (BORUTA, RF and LASSO). PEG10 mRNA, promoter methylation and plasma levels were examined. The effect of PEG10 was assessed by logistic regression, and HBeAg SC was predicted by nomograms. HBeAg-positive patients showed markedly elevated PEG10 mRNA expression (p&#x2009;<&#x2009;0.001), which correlated strongly with major virological markers such as HBV DNA (r&#x2009;=&#x2009;0.520, p&#x2009;<&#x2009;0.001), HBeAg (r&#x2009;=&#x2009;0.490, p&#x2009;<&#x2009;0.001) and HBsAg (r&#x2009;=&#x2009;0.400, p&#x2009;<&#x2009;0.001). In addition, HBeAg-positive patients exhibited a significant reduction in PEG10 promoter methylation levels compared with controls (p&#x2009;<&#x2009;0.001). According to logistic regression analysis, PEG10 promoter methylation status was an independent predictor of HBeAg SC. The predictive nomogram incorporating PEG10 promoter methylation ratio (PMR), albumin (ALB), aspartate aminotransferase (AST) and HBeAg demonstrated excellent clinical predictive value (area under curve (AUC)&#x2009;=&#x2009;0.895,95% confidence interval (CI): 0.808&#x2009;~&#x2009;0.963). The methylation status of the PEG10 promoter represents a promising biomarker for the prediction of HBeAg SC in patients with CHB. CLINICAL TRIAL REGISTRATION: Not applicable.

Humans↗

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans↗