PubMed HealthSearch

SEARCH · PubMed Health

Results for “ensemble”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

One chromatin, many structures: From ensemble contact maps to single-cell 3D organization.

Understanding how chromatin folds in three dimensions remains challenging because most experimental assays capture low-dimensional projections of an underlying, highly heterogeneous polymer. Here, we present an ensemble-based interpretive framework built on the previously introduced Self-Returning Excluded Volume (SR-EV) model, a minimal generator of chromatin conformations using a nucleosome-indexed coarse-grained representation based on stochastic return rules and excluded-volume geometry. Despite its simplicity, SR-EV recapitulates key experimental signatures across scales: heterogeneous nanoscale packing domains resembling ChromEMT and ChromSTEM observations, sparse and highly variable single-configuration contact patterns analogous to single-cell chromosome conformation capture (Hi-C), and robust ensemble-level contact enrichment consistent with topologically associating domains (TADs). In this framework, Hi-C loop and TAD signatures are interpreted as ensemble-level statistical enrichments rather than invariant features of single-cell conformations. SR-EV is explicitly designed to generate large ensembles of complete three-dimensional chromatin configurations that can be projected consistently onto two-dimensional contact maps and one-dimensional genomic profiles. By introducing architectural-protein effects only through ensemble selection rather than explicit forces, SR-EV supports a separation between intrinsic polymer geometry and regulatory bias and suggests that TAD-like features can emerge as statistical enrichments rather than deterministic three-dimensional structures. Coordination number and probe-based accessibility computed directly from SR-EV provide a unified link between three-dimensional packing, two-dimensional contact maps, and one-dimensional genomic profiles. The main contribution of this work is to show, within a single coarse-grained framework, how these multimodal observables arise as linked projections of the same heterogeneous chromatin ensemble through averaging and conditional sampling. Together, these results establish SR-EV as a minimal and geometrically grounded mesoscale reference framework for interpreting how heterogeneous chromatin ensembles give rise to multimodal experimental observables while remaining consistent with the fact that chromatin organization is realized in individual cells.

Chromatin

Pioneer in Molecular Biology: Conformational Ensembles in Molecular Recognition, Allostery, and Cell Function.

In 1978, for my PhD, I developed the efficient O(n3) dynamic programming algorithm for the-then open problem of RNA secondary structure prediction. This algorithm, now dubbed the "Nussinov algorithm", "Nussinov plots", and "Nussinov diagrams", is still taught across Europe and the U.S. As sequences started coming out in the 1980s, I started seeking genome-encoded functional signals, later becoming a bioinformatics trend. In the early 1990s I transited to proteins, co-developing a powerful computer vision-based docking algorithm. In the late 1990s, I proposed the foundational role of conformational ensembles in molecular recognition and allostery. At the time, conformational ensembles and free energy landscapes were viewed as physical properties of proteins but were not associated with function. The classical view of molecular recognition and binding was based on only two conformations captured by crystallography: open and closed. I proposed that all conformational states preexist. Proteins always have not one folded form-nor two-but many folded forms. Thus, rather than inducing fit, binding can work by shifting the ensembles between states, and this shifting, or redistributing the ensembles to maintain equilibrium, is the origin of the allosteric effect and protein, thus cell, function. This transformative paradigm impacted community views in allosteric drug design, catalysis, and regulation. Dynamic conformational ensemble shifts are now acknowledged as the origin of recognition, allostery, and signaling, underscoring that conformational ensembles-not proteins-are the workhorses of the cell, pioneering the fundamental idea that dynamic ensembles are the driving force behind cellular processes. Nussinov was recognized as pioneer in molecular biology by JMB.

Molecular Biology

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense

RNA G-quadruplexes emerge from a compacted coil-like ensemble via multiple pathways.

RNA G-quadruplexes (rG4s) are emerging as vital structural elements involved in processes like gene regulation, translation, and genome stability. Found in untranslated regions of messenger RNAs (mRNAs), they influence translation efficiency and mRNA localization. Additionally, rG4s of long noncoding RNAs and telomeric RNA play roles in RNA processing and cellular aging. Despite their significance, the atomic-level folding mechanisms of rG4s remain poorly understood due to their complexity. We studied the folding of the r(GGGA)3GGG and r(GGGUUA)3GGG (TERRA) sequences into parallel-stranded rG4 using all-atom enhanced-sampling molecular dynamics simulations, applying well-tempered metadynamics coupled with solute tempering. The obtained folding pathways suggest that RNA initially adopts a compacted coil-like ensemble characterized by dynamic guanine stacking and pairing. The three-quartet rG4 gradually forms from this compacted coil ensemble via diverse routes involving strand rearrangements and guanine incorporations. While the folding mechanism is multipathway, various two-quartet rG4 structures appear to be a common transitory ensemble along most routes. Thus, the process seems more complex than previously predicted, as G-hairpins or G-triplexes do not act as distinct intermediates, even though some are occasionally sampled. We also discuss the challenges of applying enhanced sampling methodologies to such a multidimensional free-energy surface and address the force-field limitations.

G-Quadruplexes

[Probability neuronal ensembles from the position of the theory of fuzzy sets].

The questions of organization and functioning of the probability neuronal ensembles in the cerebral central analysatory structures were regarded from the standpoints of the washed out multitudes theory. The possibility of formalization of structure of a neuronal ensemble (the washed out multitude of neurons) was shown, examples of determining the function of a neuron belonging to a neuronal ensemble were presented, the questions of transition from washed out multitudes to unwashed those were considered.

Mathematics

Quantifying the peripheral surface information entropy from conformational ensembles of globular protein-peptide complexes.

Predicting favorable protein-peptide binding events remains a central challenge in biophysics, with continued uncertainty surrounding how nonlocal effects shape the global energy landscape. Here, we introduce peripheral surface information entropy, SΨ, a quantitative measure of the statistical variability in apolar and charged non-interacting surface (NIS) proportions across conformational ensembles. Within the Gibbs free-energy relation ΔG = ΔH - TΔS, SΨ is proposed as a computationally tractable entropic proxy rather than a direct thermodynamic observable or stand-alone estimator of binding affinity. Using energy-directed molecular docking via HADDOCK3 and explicit-solvent molecular dynamics simulations, it is demonstrated that favorable binding partners exhibit emergent, low-entropy N-states (discrete macrostates in NIS state space) indicative of preferential apolar/charged surface configurations. Across dozens of peptides and multiple receptor systems (WW, PDZ, and MDM2 domains), dominant N-states persisted under varied docking parameters and initial conditions. A meta-ensemble of 657 complexes from 36 experiments over 15 years confirmed the presence of dominant NIS modes independent of in silico methodology, suggesting an evolutionary selection pressure toward specific NIS fingerprints. These findings establish SΨ as a thermoinformatic descriptor that encodes favorable binding constraints into unique statistical signatures of the NIS.

Entropy

Structural basis of differential gene expression at eQTLs loci from high-resolution ensemble models of 3D single-cell chromatin conformations.

MOTIVATION: Techniques such as high-throughput chromosome conformation capture (Hi-C) have provided a wealth of information on nucleus organization and genome important for understanding gene expression regulation. Genome-Wide Association Studies have identified numerous loci associated with complex traits. Expression quantitative trait loci (eQTL) studies have further linked the genetic variants to alteration in expression levels of associated target genes across individuals. However, the functional roles of many eQTLs in noncoding regions remain unclear. Current joint analyses of Hi-C and eQTLs data lack advanced computational tools, limiting what can be learned from these data. RESULTS: We developed a computational method for simultaneous analysis of Hi-C and eQTL data, capable of identifying a small set of nonrandom interactions from all Hi-C interactions. Using these nonrandom interactions, we reconstructed large ensembles (×105) of high-resolution single-cell 3D chromatin conformations with thorough sampling, accurately replicating Hi-C measurements. Our results revealed many-body interactions in chromatin conformation at the single-cell level within eQTL loci, providing a detailed view of how 3D chromatin structures form the physical foundation for gene regulation, including how genetic variants of eQTLs affect the expression of associated eGenes. Furthermore, our method can deconvolve chromatin heterogeneity and investigate the spatial associations of eQTLs and eGenes at subpopulation level, revealing their regulatory impacts on gene expression. Together, ensemble modeling of thoroughly sampled single-cell chromatin conformations combined with eQTL data, helps decipher how 3D chromatin structures provide the physical basis for gene regulation, expression control, and aid in understanding the overall structure-function relationships of genome organization. AVAILABILITY AND IMPLEMENTATION: It is available at https://github.com/uic-liang-lab/3DChromFolding-eQTL-Loci.

Quantitative Trait Loci

[Propagation of spikes in statistical neuron ensembles. I. Concept of phase transitions].

A system of two coupled integro-differential equations for the propagation of sipkes is presented. A qualitative consideration of the system shows possibility of concentrational phase transition in the neuron ensemble from the state of spontaneous firing to the strong periodic oscilltation activity. Naer the point of the phase transition the neuron ensemble becomes labile, which maintains appropriate conditons for the existence of mosaic structures in the neuron network.

Models, Neurological

Ensemble DNA methylation clock demonstrates Immune-metabolic aging signatures associated with mortality.

Aging is a multifactorial process that is best described in terms of the progressive acquisition of multiple layers of phenotypic changes, such as epigenetic modifications, inflammation, and metabolic dysregulation. DNA methylation clocks have been extensively used to construct epigenetic clocks based on the DNAm profiles that can be used to estimate biological age and predict age-associated outcomes. Nevertheless, the vast majority of clocks constructed so far have been based on linear models, which are unlikely to fully account for the heterogeneity and non-linearity of survival-related DNAm signatures. In this work, we constructed a heterogeneous stacked ensemble survival model based on DNAm data obtained from the Framingham Heart Study. We first identified 190 CpG loci using elastic net Cox regression and subsequently constructed a survival prediction model based on the fusion of five complementary survival models by means of a neural network meta-learner. The prediction power of the survival model was evaluated in an external validation cohort, where we observed strong performance for predicting all-cause mortality that significantly exceeded PhenoAge and was statistically comparable to GrimAge. These performance estimates were derived in cohorts of European ancestry and externally validated in postmenopausal women aged 50-79 years, and should therefore be interpreted as applicable only to demographically similar populations.

Humans

Excitatory synaptic ensemble properties in the visual cortex of the macaque monkey: a current source density analysis of electrically evoked potentials.

The spatio-temporal distributions of excitatory synaptic ensemble activities in A17 and A18 of the visual cortex of the macaque monkey have been investigated. The synaptic activities were elicited by electrical stimulation of the primary efferents and were localized by applying the current source density analysis to the intracortically recorded field potentials. The principal results are as follows: 1. In A17, two groups of activity, evoked by fast and slow afferents, respectively, were distinguishable. 2. The fast afferents induced monosynaptic activity in layer IV C alpha and layer VI, disynaptic activity in layer IV C alpha and in the supragranular layers and trisynaptic activity in layer IV B. 3. The slow efferents induced monosynaptic activity in lower layers IV C beta and layer VI, disynaptic activity via strong connections in upper layer IV C beta, further disynaptic activity in layers III and IV B and trisynaptic activity in layers V A and II. 4. With the exception that the CSD data reveal more polysynaptic activity within layer IV, there is good agreement between the spatio-temporal distribution of synaptic activities and the cortical circuit diagrams proposed in anatomical studies. 5. In A18, activities from the slow and fast conducting afferent systems are revealed in layer IV, both most likely mediated by the monosynaptically activated target cells of A17. These activities are passed on to the supra-and infragranular layers. 6. In the lateral geniculate nucleus the safety factor of transmission is higher for activity conveyed by slow-than by fast-conducting retinal afferents. 7. The spatial distribution of monocularly evoked surface potentials failed to reveal the ocular dominance columns. 8. Comparison with the cat indicates that, with respect to the intracortical circuitry and LHN-transmission, there are more similarities between the fast-group activity in the monkey and the y-system in the cat and between the slow-group activity in the monkey and the x-system in the cat than vice versa.

Animals

Oncogenic DEAD-box ATPase DDX41 establishes transcript ensembles via CLK3-dependent and -independent mechanisms.

Post-transcriptional diversification of RNA transcripts mediated by complex processing machinery, including DEAD-box ATPases, establishes and maintains cellular phenotypes. For example, DDX41 controls RNA splicing, innate immune signaling, and genome stability. Although heterozygous DDX41 germline genetic variation occurs in familial myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML), the DDX41 contributions to splicing globally, biological processes, and pathogenic mechanisms are incompletely defined. Using a genetic rescue system with Ddx41+/- myeloid progenitors, we established global wildtype DDX41 and pathogenic variant mechanisms. Differing from pathogenic variants of other RNA splicing regulators, DDX41 deficiency compromised multiple splicing steps. DDX41-regulated transcripts encoded factors controlling RNA splicing, including Cdc2-like kinase 3 (CLK3). DDX41 regulated Clk3 transcripts, and elevated CLK3 during myeloid differentiation. Loss-of-function analysis revealed DDX41-regulated splicing commonly, but not always, required CLK3. Thus, through a mechanism utilizing a splicing factor kinase that itself is DDX41-regulated, DDX41 establishes transcript ensembles in myeloid progenitors.

DEAD-box RNA Helicases

Characterizing the regulatory logic of transcriptional control at the DNA sequence level by ensembles of thermodynamic models.

MOTIVATION: Understanding how the genome encodes the regulatory logic of transcription is a main challenge of the post-genomic era, and can be overcome with the aid of customized computational tools. RESULTS: We report an automated framework for analyzing an ensemble of fits to data of a thermodynamics-based sequence-level model for transcriptional regulation. The fits are clustered accordingly with their intrinsic regulatory logic. A multiscale analysis enables visualization of quantitative features resulting from the deconvolution of the regulatory profile provided by multiple transcription factors interacting with the locus of a gene. Quantitative experimental data on reporters driven by the whole locus of the even-skipped gene in the blastoderm of Drosophila embryos was used for validating our approach. A few clusters of highly active DNA binding sites within the enhancers collectively modulate even-skipped gene transcription. Analysis of variable enhancers' length shows the importance of bound protein-protein interactions for transcriptional regulation. The interplay between activation and quenching enables function conservation of enhancers despite length variations. AVAILABILITY AND IMPLEMENTATION: The transcription factor level data used for performing the reported study is accessible in the input files in Zenodo and GitHub as well the full code. Additional data from formerly FlyEx database will be available under request.

Thermodynamics

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic

Viral surveillance beyond detection: JMTV and the need for ensemble approaches in emerging virus discovery.

The recent report by T. Murillo, L. E. Enrique Chaves-González, S. Temmam, S. Bermúdez, et al. (Microbiol Spectr 14:e04078-25, 2026, https://doi.org/10.1128/spectrum.04078-25) expands the known geographic and ecological range of Jingmen tick virus (JMTV) by detecting the virus in Amblyomma mixtum ticks collected from horses in Costa Rica. This is an important finding because A. mixtum can feed on wildlife, domestic animals, and humans, creating a possible interface for virus movement across various hosts. The study also places the Costa Rican virus in a wider phylogenetic context, linking it to JMTV diversity reported from other regions. However, the detection of viral RNA in ticks should not be interpreted as proof of local disease, human infection, or active transmission, especially in the absence of supporting results. Instead, it reflects an important signal for careful viral surveillance. Here, I discuss how JMTV illustrates the need for ensemble approaches that combine field sampling, phylogeny, segment-level genome analysis, serology, experimental validation, and data-driven virus discovery tools.

emerging viruses

Interactions among an ensemble of chordotonal organ receptors and motor neurons of the crayfish claw.

1. Action potentials of crayfish propodite-dactyl (PD) chordotonal organ receptors and two claw motor neurons, the opener inhibitor (OI) and slow closer excitor CE) were simultaneously monitored during imposed step and ramp movements of the dactyl or while the dactyl was held at various positions. 2. The activities of the cells during imposed displacements were analyzed using peristimulus time histograms and response and contour planes. The proprioceptive fields (PFs) of individual receptors resemble components of the more complex motor neuron PFs. Some receptors are briefly active after each successive opening step, while others do not respond to steps near the closed position but respond as the joint angle increases, becoming active when the claw is held open. Another type of receptor responds to closing movements. 3. Interactions among the various types of receptors and the two motor neurons were detected and analyzed by various statistical methods and intracellular recording techniques. The results indicate that receptors activated during opening movements and when the dactyl is held at open positions excite OI and CE via divergent functional connections. The efficacies of the connections made by a receptor may differ. Receptors activated by closing movements produce hyperpolarizing synaptic potentials in both efferents, possible directly or via interneurons. 4. It is concluded that several types of chordotonal organ receptors form an ensemble of parallel input channels, which modulates the activities of OI and CE and contributes to the generation of the spatial-temporal nonuniformities of their proprioceptive reflex responses.

Animals

[Spike transmission in statistical neuronal ensembles. IVa. Stating the problem in a diffusion approximation].

An area of tissue in the field CA3 of the Hippocampus has been selected controled by a basket cell. The functioning of the basket cell has been described by a point approximation method forthe transport equation, and that of the ensemble of pyramidal cells by a diffusion approximation one. Furthermore, the area of tissue is broken up into N2 of "cells" wherein a transition is effected from the diffusion equation of the N2 of the point equations involved. If the sizes of "cells" are commeasurable to the diffusional length of the pyramidal cell collaterals each "cells" being relatively independent generator of activity. In the latter case the fundamental interaction type is the mutual negative influence of the "cells" through the basket cell. Should the area be broken up into further, the mutual influence of the "cells" becomes noticeable and cannot be neglected. The behaviour of spatially homogenous diffusion approximation model is equivalent of the point approximation model.

Action Potentials

[Spike transmission in statistical neuron ensembles. III. Phase transition in a model of hippocampal field CA3].

The model consists of two types of neurons i. e. excitatory pyramidal cells and inhibitory basket ones. (the problem was formulated by Drs. O. S. Vinogradova and A. G. Bragin). The analysis of neuron activity has been carried out on the basis of "point approximation" of spikes transport equations. The graphs were obtained by computer. These graphs of the postsynaptic potentials averaged over ensemble are in good agreement with experimental data. The model observed demonstrates the phase transition over parameter characterizing the conduction of excitation between pyramidal cells. For weak pyramidal cells link there takes place the spontaneous activity regime. For strong link there were observed the epileptoid firing of neurons at 3 divided by 5 hertz and 140 divided by 240 msec phases of inhibition between bursts.

Hippocampus

Investigating cross-organism prediction of prokaryotic essential proteins using unsupervised language model and ensemble strategy.

Cross-organism prediction of essential proteins is a critical task for drug discovery and microbial engineering, yet the generalizability of existing machine learning models across diverse species remains a significant challenge. In this study, we propose DeepPEP, a large language model-based framework designed to reliably transfer essential protein annotations between distantly related organisms. Utilizing 66 curated prokaryotic datasets, we systematically evaluated DeepPEP's cross-organism performance under various conditions. Initial pairwise predictions revealed a correlation between performance and evolutionary distance; however, further investigation demonstrated that integrating training data from multiple organisms yields superior predictive power. In a benchmark scenario designed to simulate real-world applications, DeepPEP outperformed the state-of-the-art tool Geptop 2.0, showcasing a robust ability to identify species-specific essential proteins. Finally, a case study on novel genomes confirmed the model's practical effectiveness. Our results suggest that DeepPEP is a powerful strategy for prokaryotic essential protein prediction, and the rigorous evaluation framework established in this study provides a new benchmark for the field.

Large Language Models