PubMed HealthSearch

SEARCH · PubMed Health

Results for “deep transfer model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Integration of single cell multiomics data by deep transfer hypergraph neural network.

Multi-omics characterization of individual cells offers remarkable potential for analyzing the dynamics and relationships of gene regulatory states across millions of cells. How to integrate multimodal data is an open problem, existing integration methods struggle with accuracy and modality-specific biological variation retention. In this paper, we present scHyper (scalable, interpretable machine learning for single cell integration), a low-code and data-efficient deep transfer model designed for integrating paired and unpaired single-cell multimodal data. We benchmark scHyper against datasets from different multimodal data. ScHyper learns a low-dimensional representation and aligns the covariance matrices of the measured modalities, achieving high accuracy even with large scale atlas-level datasets with low memory and computational time across different cell lines, shedding light on regulatory relationships between different types of omics. Altogether, we show that scHyper is a versatile and robust tool for cell-type label transfer and integration from multimodal single-cell datasets.

Single-Cell Analysis

Hepatocyte proteome destabilization and novel targets for PFASs unveiled through combined thermal proteome profiling and deep transfer learning.

Identifying protein targets for per- and polyfluoroalkyl substances (PFASs) is essential to understand their toxicity and health risks. However, knowledge about their interacting proteins is limited since reliable identification methods are lacking. We developed an integrated approach combining thermal proteome profiling (TPP) and deep transfer learning (DTL) modeling to efficiently identify cellular targets of PFAS. TPP measured PFAS binding proteins and the affinities by nanospray liquid chromatography tandem mass spectrometry, while DTL models were constructed to predict PFAS-protein affinities using neural network algorithms. TPP results revealed that PFASs uniquely destabilized the proteome of HepG2 cells, unlike the stabilizing effects by other xenobiotics. Key protein targets for three representative PFASs (PFOA, GenX and Novec 649) were identified, which exhibited weak binding affinities (median EC50 ≈ 30 μM). The number of protein targets increased with molecular weights among the three PFASs. The DTL model achieved a higher Pearson correlation coefficient of 0.89, and reduced mean squared errors by 54 % over previous models for drug-protein interactions. Notably, TPP and DTL jointly pinpointed ribosomal proteins as novel targets of GenX, potentially linking it to cell apoptosis through disrupted protein synthesis. Biolayer interferometry validated GenX binding to RPL4 protein, driven by electrostatic interactions and halogen bonds. This integrated approach effectively uncovers novel PFASs targets, advancing insights into their adverse health effects.

Humans

Crystal structure of a complex between lumiflavin and 2,6-diamino-9-ethylpurine: a flavin adenine dinucleotide model exhibiting charge-transfer interactions.

The x-ray structure of the deep red crystalline complex lumiflavin-2,6-diamino-9-ethylpurine has been determined. The flavin and adenine derivatives form hydrogen-bonded base pairs of the Watson-Crick type. The molecules in the crystal also associate via extensively overlapped flavin/adenine and flavin/flavin stacking interactions in which there are several contacts that are closer than van der Waals distances. This, together with the red color of the crystals, is indicative of the formation of a charge-transfer complex.

Adenine

EvoSNR-Prom: Predicting promoters at single-nucleotide resolution with label-aware transfer learning of the pretrained EVO model.

The precise identification of promoters is crucial for understanding gene regulation. Deep learning methods have achieved considerable success in promoter prediction, yet most operate at the sequence level with coarse-grained labels. This means they label an entire DNA segment as either a "promoter" or "non-promoter," which results in a lack of the nucleotide-level resolution in prediction. In this study, we propose EvoSNR-Prom, a model designed for promoter prediction at single-nucleotide resolution. EvoSNR-Prom is built on the Evo foundation model and formulates promoter identification as a token-level sequence labeling problem, analogous to named entity recognition in natural language processing. To address the limited contextual information available in single-nucleotide tokenization, we introduce a lexicon-enhanced embedding strategy that incorporates biologically meaningful DNA lexicons, enriching contextual representations and improving the model's ability to capture complex sequence motifs. Furthermore, to enhance predictive performance on small size datasets, we integrate a label-aware transfer learning framework to leverage knowledge from well-annotated source species to a target organism. The results across various prokaryotic datasets show that EvoSNR-Prom achieves excellent performance. This work provides a valuable computational framework for the high-precision analysis of gene regulatory elements, contributing to the advancement of promoter prediction at single-nucleotide resolution.

Promoter Regions, Genetic

Electronic properties of sulfhydryl- and imidazole-containing peptide-cobalt(II) complexes: their relationship to cobalt(II)-substituted "blue" copper proteins.

The electronic properties of 2:1 sulfhydryl- and imidazole-containing peptide-Co(II) complexes have been investigated and compared with those of Co(II)-substituted "blue" copper proteins. The Co(II) complexes of N-mercaptoacetyl-L-histidine and 3-mercaptopropionyl-L-histidine gave the ligand field parameters of deltat = 4110 and B = 756 cm(-1), and of deltat = 4120 and B = 724 cm(-1), respectively. These values correspond well to those (deltat = 4900 and B = 730 cm(-1)) of Co(II)-substituted "blue" copper proteins. The energy differences between S leads to M(II) charge transfer bands of Co(II)-Cu(II) couples were about 14,000 cm(-1) in both the proteins and the model complexes. The spectral results suggest that "blue" copper site has a pseudotetrahedral geometry and a deep absorption near 600 nm atributes to S leads to Cu(II) charge transfer.

Cobalt

Predictive design of tissue-specific mammalian enhancers that function in the mouse embryo.

Enhancers control tissue-specific gene expression across animals1. Although deep learning2,3 has enabled enhancer prediction and design in mammalian cell lines and non-mammalian model organisms4-10 (reviewed in a previous publication11), it remains unclear whether such approaches can operate within the regulatory complexity of mammalian genomes and tissues in vivo. Here we present a general strategy for designing tissue-specific enhancers that function reliably in mice. We use deep learning to train compact convolutional neural networks on curated chromatin accessibility data and fine-tune them by transfer learning on validated human and mouse enhancers. Guided by these models, we design 15 synthetic enhancers for the heart, limb and central nervous system in mouse embryos, all of which are active in their intended target tissue. These results demonstrate that mammalian enhancer function can be reliably inferred from DNA sequence alone, enabling the predictive de novo design of tissue-specific synthetic enhancers from modest training sets. This work establishes a generalizable framework for programmable control of mammalian gene expression in vivo, opening new avenues in functional genomics, synthetic biology and gene therapy.

Animals

Space-filling models of kinase clefts and conformation changes.

Space-filling models of yeast hexokinase, adenylate kinase, and phosphoglycerate kinase drawn by computer clearly portray the bilobal character of these phosphoryl transfer enzymes, and the deep cleft which is formed between the lobes. A dramatic conformational change occurs in hexokinase as glucose binds to the bottom of the cleft, which causes the two lobes of hexokinase to come together. A substrate-induced closing of the active site cleft is postulated to occur in other kinases as well. This change may provide a mechanism by which some of these enzymes reduce their inherent adenosine triphosphatase activity and could be a general requirement of the kinase reaction.

Adenylate Kinase

DeepWheat: predicting the effects of genomic variants on gene expression and regulatory activities across tissues and varieties in wheat using deep learning.

Spatiotemporal gene expression shapes key agronomic traits, yet tissue-specific prediction remains challenging in complex crops. We present DeepWheat, a broadly applicable deep learning framework comprising DeepEXP and DeepEPI, for accurate, tissue-specific gene expression prediction. DeepEXP integrates sequence and epigenomic features to predict gene expression (PCC 0.82-0.88), while DeepEPI predicts epigenomic maps from DNA sequence to support model transfer across varieties. Validations in five wheat cultivars confirm robustness and accuracy. DeepWheat also identifies regulatory variants with strong expression effects, enabling targeted cis-regulatory elements editing and offering a powerful tool for crop functional genomics and breeding.

Triticum

Blood mitochondrial heteroplasmic variants and cognitive performance in late midlife: REGARDS study.

BACKGROUND: Studies linking mitochondrial DNA (mtDNA) variants to cognition yielded inconsistent findings, and the underlying mechanisms remain unclear. We investigated whether mtDNA heteroplasmic variants were associated with cognitive outcomes, including the Montreal Cognitive Assessment (MoCA), in 197 late midlife adults from the Reasons for Geographic and Racial Differences in Stroke (REGARDS) cohort with complete data. METHODS: MtDNA was sequenced from blood using targeted deep sequencing. Adjusted linear and mixed-effects models examined the associations by functional regions, genes, total variant burden, nonsynonymous variants, and control regions. RESULTS: Heteroplasmic variants in the control region (β = -0.44, 95% CI: -0.83, -0.05, p = 0.027) and transfer RNA (tRNA) genes (β = -1.34, 95% CI: -2.58, -0.11, p = 0.034) were associated with MoCA baseline scores. Individual variants in cytochrome c oxidase subunit 1 (CO1) (β = -1.51, 95% CI: -2.54, -0.47, p = 0.005), NADH dehydrogenase subunit 1 (ND1) (β = -2.63, 95% CI: -4.56, -0.70, p = 0.008), and Displacement Loop (D-LOOP2) (β = -2.25, 95% CI: -4.20, -0.30, p = 0.025) was associated with reduced baseline MoCA scores. The ND6 (β = −1.23, 95% CI: −2.09, − 0.37, p = 0.006), ND4 (β = −1.11, 95% CI: −2.02, − 0.20, p = 0.018), ATP Synthase Membrane Subunit 8 (ATP8; β = −1.38, 95% CI: −2.63, − 0.13, p = 0.031), and D-LOOP1 (β = −0.61, 95% CI: −1.20, − 0.01, p = 0.045) genes suggested a potential association with executive function. Longitudinal Animal Fluency Test (AFT) scores were inversely associated with heteroplasmic variants in coding regions (β = -0.10, 95% CI: -0.19, -0.006, p = 0.049), the total number of variants (β = -0.06, 95% CI: -0.11, -0.003, p = 0.037) and total nonsynonymous variants (β = -0.11, 95% CI: -0.21, -0.01, p = 0.040). Variants in the control region were associated with the greatest decline in verbal fluency (β = −0.20, 95% CI: −0.39 to − 0.002, p = 0.049). No associations were observed between mitochondrial variants and verbal memory performance or the MoCA composite scores. CONCLUSIONS: Our study indicates that mitochondrial variants measured in blood may provide insight into cognitive function during midlife. However, additional studies are needed to validate these associations and to address potential power limitations in our study.

Humans

Automated CEAP Classification of Venous Duplex Reports Using Multimodal Artificial Intelligence.

OBJECTIVE: To develop and internally validate a prototype multimodal artificial intelligence system for automated CEAP (Clinical, Etiological, Anatomical and Pathophysiological) classification of venous duplex ultrasound (VDUS) reports, integrating natural language processing of free-text components with computer vision analysis of hand-drawn anatomical diagrams. METHODS: Single centre retrospective observational study using routinely collected clinical data. One thousand consecutive venous duplex ultrasound reports from Cambridge University Hospitals NHS Foundation Trust, UK (July 2024 - May 2025) were labelled according to the CEAP classification, excluding the Etiological component, which could not be reliably determined from duplex reports alone. Transfer learning was applied using ClinicalBERT for text and MobileNetV3 for diagrammatic data. Clinical classes were predicted from request line text. Text- and image-based pathophysiological models were developed for four anatomical territories (Great Saphenous Vein, Small Saphenous Vein, Deep system, Perforators), combined using late fusion with probability averaging. RESULTS: The clinical CEAP model achieved accuracy of 0.91, macro-F1 of 0.82, and macro-AUC of 0.98. Pathophysiological prediction varied, with text models broadly outperforming image models. Fusion yielded heterogeneous benefits, improving SSV performance but reducing Deep system accuracy. The performance of the final pathophysiological CEAP fusion models varied across anatomical territories: accuracy ranged from 0.70-0.92 and macro-AUC from 0.80-0.92. CONCLUSION: This study demonstrates the feasibility of automated CEAP classification from VDUS reports. Despite class imbalance affecting minority class predictions, the strong discriminatory performance validates this multimodal ML model for extracting clinically meaningful information from real-world data. This approach offers potential, pending external validation, to streamline vascular services through automated triage and guideline-compliant decision making.

Artificial intelligence

Immunologic responses to Candida albicans. III. Effects of passive transfer of lymphoid cells or serum on murine candidiasis.

Passive transfer of immune serum gave a significant degree of protection against deep seated candidiasis in mice. Repeated attempts to transfer resistance by the transfer of sensitized lymphoid cells gave negative results, even though cutaneous delayed hypersensitivity was transferred by the cells. The results suggest that cell-mediated immunity is not of primary importance in this model of murine candidiasis, and that humoral immunity contributes to protection.

Animals

Investigation of epileptic structure properties by transfer function.

The paper deals with the possibility of applying linear control theory and the theory of statistical dynamics to the complex and quantitative analysis of electroencephalographic records. The aim is to specify signals in the brain environment, in particular sources, speed and direction of signals, and the determination from results gained of the model of signal transfer, The possibilities of applying Fourier analysis and correlation analysis are examined. The methodical approach of the transient analysis is suggested. Conditions are discussed for approximate application of linear theory. Experimental results, as well as the calculation of transfer parameters of signals are evaluated. The method of magnetic EEG records, the scanning and processing of data on analogue and digital computers are introduced. The method described of correlation and transient analysis arises from the necessity to detect epileptic structures from an evaluation of SEEG records of the deep brain structure.

Electroencephalography

Transfer learning with multiomics integration and deep neural networks reveals drug resistance mechanisms in cancer.

Drug resistance remains one of the primary challenges in effective cancer therapy. In this study, we employed a deep neural network (DNN)-based transfer learning (TL) approach to predict drug response and uncover drug resistance mechanisms. We integrated gene expression, somatic mutation, and copy number aberration (CNA) data with drug response profiles using multi-omics integration (MI). We used the Genomics of Drug Sensitivity in Cancer (GDSC) data for training and incorporated drugs with same pathways into the training models. We then evaluated drug response predictions on independent in-vivo PDX Encyclopedia (PDX) and ex-vivo the Cancer Genome Atlas (TCGA) datasets. In addition, we conducted pathway enrichment analyses to elucidate the mechanisms underlying drug resistance for paclitaxel, 5-fluorouracil (5-FU), gemcitabine, and cetuximab. We also applied Fisher's exact test (FET) to assess potential associations between drug resistance and the presence of mutations or CNAs. Our pan-drug models outperformed other methods based on the area under the precision-recall curve (AUCPR). Our pathway enrichment analyses revealed LDHB-mediated pyruvate metabolism and FYN-mediated focal adhesion might have pivotal roles in paclitaxel resistance, while PINK1-mediated mitophagy might be critical in 5-FU resistance. In addition to transcriptional activation, FET suggested that CNAs in LDHB and PINK1 may also be associated with resistance to paclitaxel and 5-FU, respectively. Furthermore, enrichment results for paclitaxel and cetuximab indicated shared resistance mechanisms between the two drugs. Importantly, our findings are consistent with prior experimental studies, providing literature-based validation of our results. Overall, our DNN-based TL approach achieved strong predictive performance across PDX & TCGA datasets and enrichment analyses provided valuable biological insights into drug resistance mechanisms.

Humans

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural

Decoding TnsC Filament Assembly in CRISPR-Associated Transposons Using Interpretable Deep Learning and Molecular Simulations.

CRISPR-associated transposons (CASTs) enable programmable DNA integration, yet how the TnsC regulator forms processive filaments on DNA to coordinate RNA-guided transposition in type V-K CAST systems remains unknown. Here, we integrate large-scale molecular simulations, interpretable deep learning using graph attention networks (GATs), and causal inference analyses to define the molecular determinants of TnsC filament nucleation and elongation. We show that TnsC nucleates by inducing localized DNA deformation that propagates along extended filaments, with Granger causality revealing that TnsC motions precede and predict DNA deformation. Interpretable GAT models demonstrate that elongation is determined during early recognition between incoming and DNA-bound subunits, followed by structural reorganization that regenerates the recruitment interface and enables processive assembly. These results elucidate the molecular mechanism of processive TnsC filament assembly and explain why isolated TnsC filaments preferentially elongate in the 5' → 3' direction, while accessory transposition factors can reshape the interaction landscape and alter filament growth polarity. Together, these findings advance our understanding of CAST function and inform the engineering of programmable DNA integration platforms. Beyond CAST systems, this work introduces an interpretable GAT approach as a general and transferable deep learning strategy for uncovering molecular mechanisms in biological systems, while demonstrating the power of causal inference for dissecting directional relationships in molecular dynamics.

Deep Learning

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross‐taxon transferability

Pretraining improves prediction of genomic datasets across species.

MOTIVATION: Recent studies suggest that deep neural network models trained on thousands of human genomic datasets can accurately predict genomic features, including gene expression and chromatin accessibility. However, training these models is computation- and time-intensive, and datasets of comparable size do not exist for most other organisms. RESULTS: Here, we identify modifications to an existing state-of-the-art model that improve model accuracy while reducing training time and computational cost. Using this streamlined model architecture, we investigate the ability of models pretrained on human genomic datasets to transfer performance to a variety of different tasks. Models pretrained on human data but fine-tuned on genomic datasets from diverse tissues and species achieved significantly higher prediction accuracy while significantly reducing training time compared to models trained from scratch, with Pearson correlation coefficients between experimental results and predictions as high as 0.8. Further, we found that including excessive training tasks decreased model performance and that this decrease could be partially but not completely rescued by fine-tuning. Thus, simplifying model architecture, applying pretrained models, and carefully considering the number of training tasks may be effective and economical techniques for building new models across data types, tissues, and species. AVAILABILITY AND IMPLEMENTATION: Code is available on GitHub and Figshare: https://github.com/optimizedlearning/genomicsML, https://doi.org/10.6084/m9.figshare.31796116.

Genomics