PubMed HealthSearch

SEARCH · PubMed Health

Results for “Generative models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis

scSurv: a deep generative model for single-cell survival analysis.

MOTIVATION: Single-cell omics analysis has unveiled the heterogeneity of various cell types within tumors. However, no methodology currently reveals how this heterogeneity influences cancer patient survival at single-cell resolution. Here, we introduce scSurv, combining a Cox proportional hazards model with a deep generative model of single-cell transcriptome, to estimate individual cellular contributions to clinical outcomes. RESULTS: The accuracy of scSurv was validated using both simulated and real datasets. This method identifies cells associated with favorable or adverse prognoses and extracts genes correlated with their contribution levels. In melanoma, scSurv reproduces known prognostic macrophage classifications and facilitates hazard mapping through spatial transcriptomics in renal cell carcinoma. We also identified genes consistently associated with prognosis across multiple cancers and demonstrated the applicability of this method to infectious diseases. scSurv is a novel framework for quantifying the heterogeneity of individual cellular effects on clinical outcomes. AVAILABILITY: The implementation of scSurv is available on GitHub (https://github.com/3254c/scSurv) and Zenodo (https://doi.org/10.5281/zenodo.17793054).

Humans

A probabilistic generative model for quantification of DNA modifications enables analysis of demethylation pathways.

We present a generative model, Lux, to quantify DNA methylation modifications from any combination of bisulfite sequencing approaches, including reduced, oxidative, TET-assisted, chemical-modification assisted, and methylase-assisted bisulfite sequencing data. Lux models all cytosine modifications (C, 5mC, 5hmC, 5fC, and 5caC) simultaneously together with experimental parameters, including bisulfite conversion and oxidation efficiencies, as well as various chemical labeling and protection steps. We show that Lux improves the quantification and comparison of cytosine modification levels and that Lux can process any oxidized methylcytosine sequencing data sets to quantify all cytosine modifications. Analysis of targeted data from Tet2-knockdown embryonic stem cells and T cells during development demonstrates DNA modification quantification at unprecedented detail, quantifies active demethylation pathways and reveals 5hmC localization in putative regulatory regions.

5-Methylcytosine

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Generative model for the first cell fate bifurcation in mammalian development.

The first cell fate bifurcation in mammalian development directs cells toward either the trophectoderm (TE) or inner cell mass (ICM) compartments in pre-implantation embryos. This decision is regulated by the subcellular localization of a transcriptional co-activator YAP and takes place over several progressively asynchronous cleavage divisions. As a result of this asynchrony and variable arrangement of blastomeres, reconstructing the dynamics of the TE/ICM cell specification from fixed embryos is extremely challenging. To address this, we developed a live-imaging approach and applied it to measure pairwise dynamics of nuclear YAP and its direct target genes, CDX2 and SOX2, which are key transcription factors of the TE and ICM, respectively. Using these datasets, we constructed a generative model of the first cell fate bifurcation, which reveals the time-dependent statistics of the TE and ICM cell allocation. In addition to making testable predictions for the joint dynamics of the full YAP/CDX2/SOX2 motif, the model revealed the stochastic nature of the induction timing of the key cell fate determinants and identified the features of YAP dynamics that are necessary or sufficient for this induction. Notably, temporal heterogeneity was particularly prominent for SOX2 expression among ICM cells. As heterogeneities within the ICM have been linked to the initiation of the second cell fate decision in the embryo, understanding the origins of this variability is of key significance. The presented approach reveals the dynamics of the first cell fate choice and lays the groundwork for dissecting the next cell fate decisions in mouse development.

Animals

CoxFormer enables spatial omics inference with multimodal generative modeling.

Gene co-expression maps transcriptome-wide gene-gene relationships, yet high-quality estimates cover less than half the genome. Meanwhile, spatial omics either profiles restricted in situ panels or lacks cellular resolution. Extending co-expression transcriptome-wide could overcome these limitations by inferring unassayed gene expression at subcellular resolution. Here we show that CoxFormer integrates literature-derived gene knowledge with co-expression networks from bulk tissues and large-scale single-cell atlases to learn 512-dimensional representations for 32,016 human genes. These embeddings capture functional gene relationships and serve as a generative prior for spatial inference across platforms and modalities. Without requiring a matched single-cell RNA-sequencing reference, CoxFormer supports four applications beyond measured genes: histology-based expression imputation, gene activity prediction from chromatin accessibility, subcellular super-resolution inference, and pathological region detection. Together, CoxFormer extends gene embedding from gene- and cell-level tasks to whole-transcriptome spatial inference, providing a unified framework for biological analysis beyond the limited gene coverage of current spatial omics technologies.

Humans

Sex-linked genes in age-structured populations.

We study the progress towards equilibrium of the frequencies of sex-linked genes in elementary discrete time models of age-structured, overlapping generation populations. It is found that, if a finite upper age limit is assumed, the difference in the frequencies of an allele in males and females will oscillate as in the familiar non-overlapping generation models, although the oscillations may be irregular. Monotonic convergence of that difference, as found by Nagylaki (1975) in continuous-time overlapping generation models without age-structure, occurs in the models considered here only when there is no upper age limit and when there is "sufficient" overlap of generations.

Age Factors

Generative AI Models in Time-Varying Biomedical Data: Scoping Review.

BACKGROUND: Trajectory modeling is a long-standing challenge in the application of computational methods to health care. In the age of big data, traditional statistical and machine learning methods do not achieve satisfactory results as they often fail to capture the complex underlying distributions of multimodal health data and long-term dependencies throughout medical histories. Recent advances in generative artificial intelligence (AI) have provided powerful tools to represent complex distributions and patterns with minimal underlying assumptions, with major impact in fields such as finance and environmental sciences, prompting researchers to apply these methods for disease modeling in health care. OBJECTIVE: While AI methods have proven powerful, their application in clinical practice remains limited due to their highly complex nature. The proliferation of AI algorithms also poses a significant challenge for nondevelopers to track and incorporate these advances into clinical research and application. In this paper, we introduce basic concepts in generative AI and discuss current algorithms and how they can be applied to health care for practitioners with little background in computer science. METHODS: We surveyed peer-reviewed papers on generative AI models with specific applications to time-series health data. Our search included single- and multimodal generative AI models that operated over structured and unstructured data, physiological waveforms, medical imaging, and multi-omics data. We introduce current generative AI methods, review their applications, and discuss their limitations and future directions in each data modality. RESULTS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and reviewed 155 articles on generative AI applications to time-series health care data across modalities. Furthermore, we offer a systematic framework for clinicians to easily identify suitable AI methods for their data and task at hand. CONCLUSIONS: We reviewed and critiqued existing applications of generative AI to time-series health data with the aim of bridging the gap between computational methods and clinical application. We also identified the shortcomings of existing approaches and highlighted recent advances in generative AI that represent promising directions for health care modeling.

Artificial Intelligence

A nonhuman primate vascular shunt model for thrombus generation.

A nonhuman primate model for thrombus generation was developed. Three different types of test devices constructed of polystyrene or polyethylene-Silastic were exposed to flowing blood in arterioarterial or arteriovenous vascular shunts in rhesus monkeys. The test device design included a simple tube, a vortical flow device, and a turbulent flow device. The amount of thrombus deposited within each individual test device after a 15-minute exposure to blood flowing through the shunt was determined gravimetrically. These studies indicate that a test device designed to include an area of vortical flow generated the greatest amount of thrombus. Test devices fabricated from polystyrene consistently generated larger thrombus deposits than did similar test devices fabricated from polyethylene-Silastic. Arteriovenous shunts proved superior to arterioarterial shunts in that flow was predictable in the former and unpredictable in the latter; venovenous shunts thrombosed quickly. Hematologic studies indicated a progressive fall in platelet count during the 4-hour test interval in all animals, whereas in only a few animals were there a shortening of partial thromboplastin time values, a fall in fibrinogen levels, and the appearance of fibrin degradation products. An optimal model for thrombus generation appears to include vortical flow test devices fabricated of polystyrene exposed to flowing blood in an arteriovenous shunt.

Animals

Computer-generated graphic models of the N2-substituted deoxyguanosine adducts of 2-acetylaminofluorene and benzo[a]pyrene and the O6-substituted deoxyguanosine adduct of 1-naphthylamine in the DNA double helix.

Computer models of three deoxyguanosine-carcinogen adducts in double-helical DNA are presented. The carcinogen moiety is rotated and the best fit within the double helix is evaluated. The 2-acetylaminofluorene (AAF) derivative, 3-(deoxyguanosin-N2-yl)-AAF, is found to be situated within the minor groove, has very little freedom of rotation and causes little helical distortion. The (+)-anti-benzo[a]-pyrene (BP)-diol epoxide-N2 adduct, 10beta-(deoxyguanosin-N2-yl)-7beta, 8alpha,9alpha-trihydroxy-7,8,9,10-tetrahydro-BP, has a similar fit with a greater degree of steric interaction, suggesting that this adduct could cause some local destabilization. The 1-naphthylamine (NA) derivative, N1-(deoxyguanosine-O6-yl)-1-NA, resides within the major groove, does not perturb the helix and has considerable freedom of movement.

1-Naphthylamine

CRISPR/Cpf1-mediated knockout of FLG in human induced pluripotent stem cells generates a model for studying epidermal barrier dysfunction.

Loss of filaggrin (FLG) function impairs skin barrier formation and contributes to common inflammatory skin diseases. In this study, we established a FLG knockout human induced pluripotent stem cell (iPSC) line based on KOLF2.1 J using CRISPR/Cas12a (Cpf1)-mediated genome editing. A guide RNA targeting exon 2 introduced a homozygous mutation, which was confirmed by sequencing. The edited cells maintained typical pluripotent stem cell morphology, expressed key undifferentiated markers, and retained the ability to differentiate into all three germ layers. Karyotype and copy number variation (CNV) analyses confirmed genomic stability and parental origin; the cells were free of mycoplasma. This cell line enables studies of FLG-associated skin biology and pathology.

Humans

Model-directed generation of artificial CRISPR-Cas13a guide RNA sequences improves nucleic acid detection.

CRISPR guide RNA sequences deriving exactly from natural sequences may not perform optimally in every application. Here we implement and evaluate algorithms for designing maximally fit, artificial CRISPR-Cas13a guides with multiple mismatches to natural sequences that are tailored for diagnostic applications. These guides offer more sensitive detection of diverse pathogens and discrimination of pathogen variants compared with guides derived directly from natural sequences and illuminate design principles that broaden Cas13a targeting.

CRISPR-Cas Systems

The effects of consultation style on consultee productivity.

This study examined the effectiveness of three different consultation styles adapted from Bindman's typology. Consultees were nurses on eight wards in a state hospital for the retarded, who were assigned to Expert, Resource, and Process consultation groups plus a no-treatment control. Data on the number of new programs independently initiated by consultees were collected during a 6-week base line, 12-week consultation, and 6-week follow-up period. Results showed a general increase in number of programs initiated during the second half of the consultation period, with trends established there continued through the follow-up. Degree of change was directly related to the style of consultation: the Expert role proved no better than the control condition; the Resource and Process roles generated significant consultee activity, with the Process model generating the most programs in both experimental and follow-up periods.

Behavior Therapy

GENKI: A generative framework for scalable and robust metabolic kinetic modeling.

GENKI (Generative ENsemble KPI-Informed) is a variational autoencoder-based framework for large-scale kinetic modeling of metabolism. Developed for metabolic engineering applications, GENKI is designed to improve the recovery of kinetically feasible models that reproduce experimentally observed phenotypes under genetic and environmental perturbations. The framework is trained on feasible kinetic model ensembles and uses phenotype-based key performance indicators (KPIs), derived from multi-omics and bioprocess data, to label and enrich models according to their agreement with mutant and condition-specific observations. This enables targeted generation of biologically relevant parameter sets with improved predictive performance. Crucially, GENKI recovers kinetic parameter sets that jointly reproduce wild-type and multiple perturbed physiologies within a single model. We apply GENKI to large-scale kinetic models of Escherichia coli and Saccharomyces cerevisiae under enzyme perturbations and oxygen shifts. In both systems, GENKI enriches kinetic ensembles with models that more accurately reproduce experimentally observed physiologies across multiple perturbations and conditions. GENKI therefore provides a practical framework for perturbation-aware kinetic model refinement within iterative Design-Build-Test-Learn workflows.

DBTL