PubMed HealthSearch

SEARCH · PubMed Health

Results for “Deep generative model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Towards mechanistic models of mutational effects: Deep learning on Alzheimer's Aβ peptide.

Deep Mutational Scanning (DMS) has enabled multiplexed measurement of mutational effects on protein properties, including kinematics and self-organization, with unprecedented resolution. However, potential bottlenecks of DMS characterization include experimental design, data quality, and depth of mutational coverage. Here, we apply deep learning to comprehensively model the mutational effect of the Alzheimer's Disease associated peptide Aβ42 on aggregation-related biochemical traits from DMS measurements. Among tested neural network architectures, Convolutional Neural Networks and Recurrent Neural Networks are found to be the most cost-effective models with high performance even under insufficiently-sampled DMS studies. While sequence features are essential for satisfactory prediction from neural networks, geometric-structural features further enhance the prediction performance. Notably, we demonstrate how mechanistic insights into phenotype may be extracted from the neural networks themselves suitably designed. This methodological benefit is particularly relevant for biochemical systems displaying a strong coupling between structure and phenotype such as the conformation of Aβ42 aggregate and nucleation, as shown here using a Graph Convolutional Neural Network (GCN) developed from the protein atomic structure input. In addition to accurate imputation of missing values (which here ranged up to 55% of all phenotype values at key residues), the mutationally-defined nucleation phenotype generated from a GCN shows improved resolution for identifying known disease-causing mutations relative to the original DMS phenotype. Our study suggests that neural network derived sequence-phenotype mapping can be exploited not only to provide direct support for protein engineering or genome editing but also to facilitate therapeutic design with the gained perspectives from biological modeling.

Alzheimer's disease

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (Ψ) represents one of the most abundant and conserved RNA modifications. Ψ provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of Ψ sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel Ψ site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA Ψ-site prediction. The Ψ modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA Ψ-site prediction. Meta-PseU offers a new framework for robust Ψ-site identification by using long sequences.

Pseudouridine

Gold-195m: a steady-state imaging agent for venography that gives blood velocity measurement.

Gold-195m has found applications in first-pass studies for investigating both right and left ventricular activity as well as lung transit. Owing to its reasonably short half-life of 30 sec we have found it particularly useful for imaging leg veins up to and including the inferior vena cava. Its short half-life prevents recirculation activity from appearing, so continuous perfusion into a superficial foot vein and application of ankle tourniquets yield a steady-state image of the deep veins, with particularly good resolution. Its decay pattern along a vessel is very sensitive to blood velocity, so measurement of activity at various points on a vein in a computer static image can give velocity values that reveal abnormalities due to partial or complete thrombosis. The radiation dosimetry of 195mAu used in this way is lower than contrast and technetium-99m macroaggregated albumin [( 99mTc]MAA) venography, making it particularly useful for investigating deep vein thrombosis (DVT) in pregnancy.

Abdomen

Unveiling novel antimicrobial peptides from the ruminant gastrointestinal microbiomes: A deep learning-driven approach yields an anti-MRSA candidate.

INTRODUCTION: Antimicrobial peptides (AMPs) present a promising avenue to combat the growing threat of antibiotic resistance. The ruminant gastrointestinal microbiome serves as a unique ecosystem that offers untapped potential for AMP discovery. OBJECTIVES: The aims of this study are to develop an effective methodology for the identification of novel AMPs from ruminant gastrointestinal microbiomes, followed by evaluating their antimicrobial efficacy and elucidating the mechanisms underlying their activity. METHODS: We developed a deep learning-based model to identify AMP candidates from a dataset comprising 120 metagenomes and 10,373 metagenome-assembled genomes derived from the ruminant gastrointestinal tract. Both in vivo and in vitro experiments were performed to examine and validate the antimicrobial activities of the AMP candidates that were selected through bioinformatic analysis and subsequently synthesized chemically. Additionally, molecular dynamics simulations were conducted to explore the action mechanism of the most potent AMP candidate. RESULTS: The deep learning model identified 27,192 potential secretory AMP candidates. Following bioinformatic analysis, 39 candidates were synthesized and tested. Remarkably, all synthesized peptides demonstrated antimicrobial activity against Staphylococcus aureus, with 79.5% showing effectiveness against multiple pathogens. Notably, Peptide 4, which exhibited the highest antimicrobial activity against methicillin-resistant Staphylococcus aureus (MRSA), confirmed this effect in a mouse model with wound infection, exhibiting a low propensity for resistance development and minimal cytotoxicity and hemolysis towards mammalian cells. Molecular dynamics simulations provided insights into the mechanism of Peptide 4, primarily its ability to disrupt bacterial cell membranes, leading to cell death. CONCLUSION: This study highlights the power of combining deep learning with microbiome research to uncover novel therapeutic candidates, paving the way for the development of next-generation antimicrobials like Peptide 4 to combat the growing threat of MRSA would infections. It also underscores the value of utilizing ruminant microbial resources.

Animals

A computational model for heat generation in a radially layered tissue inside a 'coaxial TEM' applicator.

The electromagnetic heat dissipation in a radially layered biological tissue inside a circular cylinder has been investigated theoretically. The theory is based on a three-dimensional model and the electromagnetic field is assumed to be generated by a prescribed electric field along a ring-shaped aperture. The method of computation employs the spatial Fourier transform of all field quantities with respect to the axial coordinate, after which the field equations are solved in the spectral domain. Subsequently, an inverse Fourier transform is carried out to compute the quantities that are of interest to the clinical deep-body hyperthermia system at hand. For a number of representative configurations numerical results at 70 MHz are given.

Body Temperature

Single-organ proteomics in Drosophila melanogaster larva.

The combination of genetic accessibility, organ complexity, evolutionary conservation, and cost-efficiency makes Drosophila melanogaster (Dm) a well-known model system for biomedical and fundamental biological research. Proteomic analysis of single organs enables the identification and quantification of proteins expressed in specific organs. This will help to uncover specific biological functions and unique protein profiles that are not detectable in whole-organism analyses. In this study we have isolated single organs form Dm larvae, and we have performed a deep proteomics mapping by following a minimal manipulation preparation procedure. The combined dataset across all organs comprised 9132 identified proteins. As anticipated, principal component analysis (PCA) revealed clear separation between the proteomes of most organs, confirming distinct protein profiles. These findings demonstrate the applicability of the sample preparation strategy for high-resolution proteomic characterization of individual organs in Drosophila. Given the extensive genetic tools available for this model organism, our approach has the potential to open new avenues for proteomic studies in Drosophila melanogaster and any other biological systems where the sample amount is limiting. SIGNIFICANCE STATEMENT: Drosophila melanogaster is a well-known model system for biomedical and fundamental biological research that serves as a valuable in vivo model organism due to its high degree of evolutionary conservation with higher vertebrates, tractable genetics, and logistical efficiency. However, the proteome of Drosophila at single organ level has been elusive to date, due to several factors like low sensitivity of previous generation mass spectrometers and sample preparation procedures, difficult isolation of some organs. In this study we have applied a compilation of advanced methods including minimal sample manipulation together with simple, straightforward and efficient protein extraction and digestion methods. Obtained peptides were minimally handled to be analyzed by applying specific and sensitive nLC methods coupled on-line to state-of-the-art MS/MS system. Altogether, the applied strategy allowed us to get the first single organ study to date for this animal. These datasets represent a significative resource for future genomic, transcriptomic and proteomic studies in Drosophila, as multi-omic integration requires deep proteomics to translate data into functional biochemistry, and serves as a critical bridge and an indispensable standalone resource across the genomic, transcriptomic, and proteomic landscapes.

Animals

Flow through a venous valve and its implication for thrombus formation.

To elucidate the possible connection between the flow patterns in the pockets of venous valves and thrombus formation, detailed studies of the behavior of model particles and red cells flowing through a venous valve have been carried out using isolated transparent dog saphenous veins containing two-leaflet valves, and cinemicrographic techniques. It was found that large paired vortices, located symmetrically on both sides of the bisector plane of the valve leaflets, were present in each valve pocket under physiological flow conditions. Particles continually entered the valve pockets from the mainstream, spending long periods of time describing a series of spiral orbits of decreasing diameter, while moving away from the bisector plane, and eventually left the vortex, rejoining the mainstream. With concentrated suspensions of red cells, it was found that another smaller counter-rotating secondary vortex, driven by the large primary vortex existed deep in each valve pocket. The concentration of red cells in this secondary vortex remained appreciably lower than that in the mainstream. In such regions, fluid circulated with extremely low velocities, thus creating a very low shear field which allowed red cells to form aggregates. The results suggest that in some pathological states, the valve-pocket vortices could act as automatic traps and generators of thrombi in a fashion similar to that previously demonstrated in an annular vortex formed downstream from a sudden tubular expansion.

Animals

Translating functional molecular knowledge into crop-breeding success.

Historical plant breeding, which optimizes phenotypes through selective crossing guided by phenotypic evaluation and molecular markers, is limited by evolutionary constraints that hinder rapid crop improvement. A new paradigm, precision breeding, circumvents these limitations by targeting genetic variants through functional molecular knowledge. To generate this knowledge at scale, sequence-based deep learning leverages high-quality genome sequence data to predict variant effects at base-pair resolution. When linked to agronomically important traits, these predictions enable breeders to prioritize variants for precision selection or editing. Although it is still in the early stages of development, we foresee three key applications for this approach: introgressing genes from distant breeding pools, purging deleterious mutations and designing new plant ideotypes. Looking ahead, refined computational models will facilitate targeted editing and the systematic redesign of complex physiological processes to address emerging breeding goals under shifting environmental conditions.

Crops, Agricultural

Flexibility plot of proteins.

The flexibility plot of a protein lies on the observation that amino acid residues with the highest turn potential, i.e. located in highly mobile regions of protein surface, also possess the smallest volumes as well as the lowest hydrophobicities. The plot is generated by shifting a five residue window along the protein sequence and calculating the value of the hydrophobicity-volume product for consecutive quintuplets of amino acid residues. The concomitant occurrence of small volumes and low hydrophobicities results in very deep minima. A threshold value has also been introduced in order to discriminate significant minima. To substantiate the interpretation that the selected minima actually indicate very flexible segments of a protein (loops, turns, etc.), we have compared plots obtained for model proteins (lysozyme, myoglobin, ribonuclease, trypsin, thermolysin and T4 lysozyme) with X-ray thermal factors profiles available for the same proteins. When compared to thermal profiles, the majority of flexible segments evidenced by our plots have been found to be in agreement with regions characterized by high thermal factors. Results have also been discussed in the light of local organization possessed by examined proteins.

Amino Acids

Evaluation of electrofulguration in control of bleeding of experimental gastric ulcers.

The safety and efficacy of electrofulguration for control of bleeding from standard canine experimental gastric ulcers was studied. At settings of 2, 5, and 8 on a Valleylab SSE-3 generator, 0.5-sec applications provided effective hemostasis. However, a setting of 2 required an excessive number of applications. Settings of 5 and 8 showed deep injury to the muscularis externa when examined histologically. In an attempt to reduce the depth of injury, a more easily ionizable gas mixture of 50% argon gas and 50% CO2 was compared to CO2 alone. At a generator setting of 5 with 0.5-sec applications the argon-CO2 mixture produced slightly less deep injury than CO2 alone, but the difference was not significant. Although electrofulguration was effective in stopping bleeding in these experiments, the tissue injury was unpredictable and deep.

Animals

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

Optimization of the intensity gain of multiple-focus phased-array heating patterns.

A new technique for enhancing the intensity gain at the focal points in multiple-focus patterns is introduced. The new technique is shown to be effective in reducing the interference typically associated with multiple-focus patterns. This reduction in interference patterns allows multiple-focus scanning to generate highly localized heating. Simulation results indicate that multiple-focus scanning not only provides an alternative to single-focus scanning, but also achieves better localization in the heating pattern. The maximization of intensity gain of multiple-focus heating patterns significantly reduces the pre-focal-depth high-temperature regions that can be caused by single-focus scanning. This is shown by computer simulation of a two-dimensional cylindrical-section array (CSA2D) as a heating applicator. Two series of simulations are presented in which different scan trajectories were used to therapeutically heat a small deep-seated target volume. In every case the heating pattern was generated using single-focus scanning and multiple-focus scanning (with and without intensity gain maximization). Multiple-focus scanning with gain maximization offers the best localization of heating to the target volume of the three methods.

Biophysical Phenomena

Profiler: an open web platform for multi-omics analysis.

MOTIVATION: High-throughput multi-omics technologies produce increasingly large and heterogeneous datasets that are difficult to analyze without advanced computational expertise. Existing bioinformatics tools are often fragmented or limited to specific omics types, hindering reproducibility and accessibility. There is a critical need for an integrated, user-friendly, and scalable platform capable of supporting multi-omics analyses across different data modalities. RESULTS: We present Profiler, an open-source, modular platform that unifies data import, quality control, preprocessing, statistical testing, machine and deep learning, biomarker discovery, pathway and drug-target enrichment, and survival modeling within a single reproducible environment. Built in Python with Streamlit, Profiler is available as both a web-based platform deployed on high-performance computing and a desktop version for local execution, enabling flexible usage across computational infrastructures. Profiler supports diverse omics modalities, including proteomics, transcriptomics, lipidomics, and electroencephalogram data. Through applications to glioblastoma proteomic, pancancer, and multi-omics datasets, Profiler reproduced known molecular subtypes, revealed potential therapeutic targets, and generated fully traceable analysis reports within minutes. By integrating advanced analytics behind an intuitive interface, Profiler democratizes multi-omics analysis and provides a robust, scalable foundation for systems biology and precision medicine research. AVAILABILITY AND IMPLEMENTATION: Profiler is open-source and freely available via its web platform (https://prism-profiler.univ-lille.fr) and GitHub (web version: https://github.com/yanisZirem/Profiler_v1_requests_datatests, desktop version: https://github.com/yanisZirem/prism-profiler), and archived on Zenodo (DOI: https://doi.org/10.5281/zenodo.17478158).

Software

Source modelling of the rolandic focus.

In rolandic epilepsy, consideration of the stereotyped ictal symptomatology suggests that the epileptic zone is likely to be in the same cortical structure in different patients. Routine EEG tracings of the interictal spike activity suggests a deep Sylvian fissure location. On the basis of the predominantly tangential potential field at the peak spike negativity seen in this group of patients, the inferior bank of the Sylvian fissure appears to be a good candidate. Without invasive studies, little refinement to this rather imprecise localization can be made as there is neither neurologic deficit nor lesion to provide a marker on radiological imaging. However, the application of source modelling technique using a simple single-dipole spherical head model has resulted in improved understanding of the generator behaviour, and facilitated the generation of new ways of analyzing spikes (e.g., stability index). Review of newer quantitative approaches including matrix and singular value decomposition of the dataset, spatial-temporal constrained source estimates etc. suggest other fruitful approaches. At least in some patients with partial epilepsy, the source characteristics of interictal scalp spikes appear to contain information of the ictal generator. Under certain circumstances, such derived information which is not otherwise available from routine electrophysiology may influence clinical management and prognosis. This is an additional bonus to the primary objectives of quantification and data reduction.

Brain

Treemble: a graphical tool to generate Newick strings from phylogenetic tree images.

SUMMARY: Phylogenetic trees are ubiquitous and central to biology, but most published trees are available only as visual diagrams and not in the machine-readable Newick format. There are, thus, thousands of published trees in the scientific literature that are unavailable for follow-up analyses, comparisons, and supertree construction. Experts can easily read such diagrams, but the manual construction of a Newick string from a diagram is laborious, error-prone, and time-consuming. Previous attempts to semi-automate the reading of tree images relied on image processing techniques. These often encounter difficulties as typical published tree diagrams contain various graphical elements and annotations that overlap the branches, such as error bars on internal nodes. Here we introduce Treemble, a user-friendly desktop application for generating Newick strings from tree images. The user simply clicks to mark node locations, assisted by a deep learning-based node detection tool, and Treemble algorithmically assembles the tree from the node coordinates alone. Treemble also facilitates the automatic reading of tip name labels and can be used for both rectangular and circular trees. AVAILABILITY AND IMPLEMENTATION: Treemble is a native desktop application for macOS and Windows and is freely available, with documentation, at treemble.org. Source code is available at github.com/John-Allard/Treemble. The trained node detection model is available at huggingface.co/John-Allard/treemble-1.

Phylogeny

Comparative theoretical performance for two types of regional hyperthermia systems.

Regional hyperthermia systems have drawn attention because of their potential for depositing power noninvasively in deep-seated tumors. Two such systems that have received clinical attention because of their ability to deposit significant amounts of power in tissue are magnetic induction devices and annular phased array applicators. In this paper, theoretical calculations for the specific absorption rate (SAR) and the resulting temperature distributions for these systems are compared. The finite element method is used in the formulation of both the electromagnetic and thermal boundary value problems. Six detailed patient models based on CT-scan data from the pelvic, visceral, and thoracic regions are generated to simulate a variety of tumor locations. In general, the annular phased array deposited more power within the tumor and produced better temperature distributions than the magnetic induction device. However, the ratio of the maximum power absorbed by the tumor to the maximum power absorbed in normal tissue does not appear to be high enough for either device to heat significant portions of perfused tumors to therapeutic temperatures under a wide range of physiological conditions. The results contained herein should aid the physician in comparative treatment planning with existing regional hyperthermia systems.

Electromagnetic Fields

Sequence optimization targeting mRNA stability enhances monoclonal antibody titers in CHO cells.

This study presents a DNA sequence optimization approach that integrates mRNA stability as a tunable design parameter to enhance monoclonal antibody expression in Chinese hamster ovary (CHO) cells. A comprehensive combinatorial library of synonymous coding-sequence variants of an IgG1 light chain was integrated as single copies at a defined genomic locus in CHO cells with identical regulatory elements. Steady-state mRNA abundance, quantified by deep sequencing of gDNA and mRNA, served as a proxy for mRNA stability. These data were used to train a machine learning model that predicts mRNA abundance from coding sequence using embeddings from a pre-trained nucleotide transformer. This abundance predictor, together with established translational metrics, was incorporated into a genetic algorithm for multi-objective codon optimization. As proof-of-concept, we optimized sequences encoding Trastuzumab to either maximize or minimize the abundance criterion and obtained benchmark sequences from two commercial providers. Using targeted integration, we generated CHO cell lines and measured protein titer and cell-specific productivity. Sequences optimized for high abundance significantly increased intracellular mRNA levels (+41%), protein titer (+59%), and cell-specific productivity (+85%) relative to low-abundance designs, while viable cell densities remained comparable. Compared to commercial benchmarks, high-abundance sequences achieved significantly higher titer (+70%) and cell-specific productivity (+98%). These findings establish mRNA stability as a practical and complementary design parameter for codon optimization in monoclonal antibody production, with potential applicability to other proteins and expression systems.

CHO

Modelling coronary thrombosis from nonanticoagulated human blood in vitro.

The prevalence of ruptured atheromatous plaques underlying the adherent thrombus in the infarct-related coronary arteries, is well documented. In the thrombotic process associated with plaque rupture, hemodynamic forces and the interaction of platelets with exposed collagen fibers play the decisive roles. The shear-induced hemostasis from a nonanticoagulated human blood sample, perfused through polyethylene tubing, was used to simulate rheological changes in the coronary circulation due to plaque disruption. When the hemodynamic conditions of a plaque fissure were mimicked, the sequence of events corresponded to that in vivo: hemostasis (i.e., platelet plug formation in the wall) resulted in the formation of an occlusive thrombus in the lumen of the tubing. Further, the thrombus growth on a collagen fiber, mounted in the lumen of polyethylene tubing through which nonanticoagulated human blood was perfused, was used to mimick the exposure of thrombogenic elements during deep vessel wall injury and the formation of thrombus superimposed on plaque disruption. Morphology of both types of thrombi revealed large numbers of neutrophils and monocytes associated with the platelet mass. The mechanisms of thrombotic reactions were characterized by antagonists and monoclonal antibodies against the platelet activation pathways. Generation of thrombin at an early stage was shown to be the key event and determinant of the final outcome of both thrombotic reactions. It is suggested that the simultaneous measurements of shear-induced hemostasis, clotting, and platelet-collagen interaction from nonanticoagulated human blood provide close experimental approximations to the pathological process of acute coronary syndromes, namely thrombus formation at the site of a disrupted atherosclerotic plaque.

Adolescent