PubMed HealthSearch

SEARCH · PubMed Health

Results for “transfer machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

32 records · Page 2Linked to original sources

Metagenomic polymorphic toxin effector and immunity profiling predicts microbiome development and disease-related dysbiosis.

Bacteria use antagonistic interbacterial weapons, such as polymorphic toxin secretion systems (TSS), to compete for niches in the human gut microbiome. We hypothesized that TSS influence gut microbiome development and disease-related dysbiosis. We developed a bioinformatic marker gene approach (PolyProf) to quantify TSS including ~200 effector and immunity genes and applied it to ~15,000 publicly available human metagenomes. PolyProf alpha and beta diversity readily distinguished 12 different human disease states and enabled the construction of highly accurate linear regression classifier machine learning models. Elastic net machine learning models integrating bacterial taxonomy with PolyProf had strong predictive value for 12 disease states, outperforming models utilizing taxonomy alone. During microbiome development in the first year of life, PolyProf alpha diversity increases, and beta diversity becomes increasingly like the maternal microbiome, influenced by vertical transfer, delivery mode, and breastfeeding. PolyProf is related to strain sharing among adults through social interactions. In summary, TSS genes strongly correlate with microbiome development and interpersonal strain sharing, suggesting roles for interbacterial antagonism. Since PolyProf distinguishes diverse adult disease statuses, these dynamics may contribute to non-genetic inheritance.IMPORTANCEPrevious research has demonstrated that bacteria compete within the gut microbiome using toxin secretion systems (TSS). How TSS contribute to human microbiome development and the microbiome alterations observed in human diseases is not known. This study develops a new bioinformatic tool for profiling TSS-related genes in metagenomic data. Application of this approach to large-scale human fecal metagenomic data demonstrates the dynamic association of TSS during microbiome development, including the exchange of strains among social contacts. TSS gene abundance patterns are highly predictive of 12 disease states. This study advances the field by enabling TSS profiling in metagenomes and by identifying disease and microbiome development biomarkers that provide hypotheses for future mechanistic studies and may be useful for disease diagnosis.

Dysbiosis

Seeing and Feeling DNA Methylation: Single-Molecule Biophysics Meets Machine Learning.

DNA methylation at 5-methylcytosine (5mC) is crucial for embryonic development and cellular function, while aberrant patterns strongly drive disease onset and progression. Its reversible nature offers substantial therapeutic potential, emphasizing the need for precise, context-specific genome wide 5mC mapping. Conventional techniques such as bisulfite sequencing and ensemble biosensor assays are hindered by DNA degradation, amplification bias, high cost, and inability to resolve single-molecule structural and mechanical effects of methylation. This review examines advances in single-molecule biophysical methods (nanopore sensing, smFRET, optical/magnetic tweezers, and AFM) that provide direct, label-free/minimally invasive 5mC detection, along with quantitative insights into DNA conformation, mechanics, and protein-DNA interactions. These techniques complement traditional methylome mapping by linking genomic localization to molecular mechanisms. Emerging machine-learning approaches are revolutionizing analysis, particularly in nanopore sensing, while promising applications in smFRET, tweezers, and AFM address throughput and reproducibility challenges. Their convergence promises scalable, high-resolution epigenetic profiling, advancing precision epigenomics toward clinical application.

DNA Methylation

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease

OpenSpliceAI: An efficient, modular implementation of SpliceAI enabling easy retraining on non-human species.

The SpliceAI deep learning system is currently one of the most accurate methods for identifying splicing signals directly from DNA sequences. However, its utility is limited by its reliance on older software frameworks and human-centric training data. Here we introduce OpenSpliceAI, a trainable, open-source version of SpliceAI implemented in PyTorch to address these challenges. OpenSpliceAI supports both training from scratch and transfer learning, enabling seamless retraining on species-specific datasets and mitigating human-centric biases. Our experiments show that it achieves faster processing speeds and lower memory usage than the original SpliceAI code, allowing large-scale analyses of extensive genomic regions on a single GPU. Additionally, OpenSpliceAI's flexible architecture makes for easier integration with established machine learning ecosystems, simplifying the development of custom splicing models for different species and applications. We demonstrate that OpenSpliceAI's output is highly concordant with SpliceAI. In silico mutagenesis (ISM) analyses confirm that both models rely on similar sequence features, and calibration experiments demonstrate similar score probability estimates.

Journal Article

A robust transfer learning approach for high-dimensional linear regression to support integration of multi-source gene expression data.

Transfer learning aims to integrate useful information from multi-source datasets to improve the learning performance of target data. This can be effectively applied in genomics when we learn the gene associations in a target tissue, and data from other tissues can be integrated. However, heavy-tail distribution and outliers are common in genomics data, which poses challenges to the effectiveness of current transfer learning approaches. In this paper, we study the transfer learning problem under high-dimensional linear models with t-distributed error (Trans-PtLR), which aims to improve the estimation and prediction of target data by borrowing information from useful source data and offering robustness to accommodate complex data with heavy tails and outliers. In the oracle case with known transferable source datasets, a transfer learning algorithm based on penalized maximum likelihood and expectation-maximization algorithm is established. To avoid including non-informative sources, we propose to select the transferable sources based on cross-validation. Extensive simulation experiments as well as an application demonstrate that Trans-PtLR demonstrates robustness and better performance of estimation and prediction when heavy-tail and outliers exist compared to transfer learning for linear regression model with normal error distribution. Data integration, Variable selection, T distribution, Expectation maximization algorithm, Genotype-Tissue Expression, Cross validation.

Linear Models

Can host genetics transform the sustainable control of tropical theileriosis? Insights from the Tick-Theileria interface.

Tropical theileriosis, caused by the tick-transmitted apicomplexan parasite Theileria annulata, remains a major constraint on cattle production across North Africa, the Mediterranean basin, the Middle East and South Asia. Current control depends on acaricides, the theilericidal drug buparvaquone and live attenuated schizont vaccines, but acaricide resistance, buparvaquone-resistance mutations and the logistical demands of vaccination are eroding the sustainability of these tools. Host genetics offers a complementary and durable alternative. Indigenous Bos indicus breeds are consistently more resistant to ticks and tolerate T. annulata infection better than exotic Bos taurus cattle, and this advantage has a measurable heritable component. Unlike previous reviews, which treat tick resistance, T. annulata immunobiology and livestock genomic selection as separate subjects, we integrate all three and assess host genetics specifically against the failure modes of current control. We review the tick, parasite and host interface, the evidence for natural resistance, and the genetic and immunological mechanisms involved, including signal-regulatory protein, bovine major histocompatibility complex class II and inflammatory pathway genes. We then assess whether genomic selection, multi-omics, machine learning and gene editing can translate these mechanisms into resistant cattle, and we weigh the biological, economic and infrastructural barriers to implementation. The evidence indicates that host genetics will not replace existing control but could reduce reliance on acaricides and chemotherapy. That contribution remains prospective rather than demonstrated: no resistance marker for T. annulata has yet been validated, prediction accuracies are moderate and transfer poorly between breeds, and no endemic production system has implemented selection for resistance.

Animals

Water source, latrine type, and rainfall are associated with detection of non-optimal and enteric bacteria in the vaginal microbiome: a prospective observational cohort study nested within a cluster randomized controlled trial.

BACKGROUND: Less than one-third of sub-Saharan Africans have access to improved water sources. In US, Indian, and African studies, Bacterial vaginosis (BV) is increased among women with poor water, sanitation, and hygiene (WASH). We examined water source, sanitation (latrine type), and rainfall in relation to the vaginal microbiome (VMB). METHODS: In a cluster randomized controlled trial of menstrual cups and cash transfer, we measured the impact of cups on VMB via 16S rRNA gene amplicon sequencing in a subset of 436 adolescent girls. We analyzed how self-reported water source and latrine type at home related to VMB over 18-months, examining community state type I (CST-I, L. crispatus dominant) vs. other CST; alpha diversity; targeted taxa (coliform and other water-related pathogens); and non-targeted taxa via machine learning approaches. Mixed effects multivariable longitudinal models were adjusted for intervention arm, age, socioeconomic status, sexual activity, and cluster-level school WASH and rainfall (in millimeters). RESULTS: Adjusting for all covariates in all models: (1) the odds of CST-I were increased among participants with piped water (vs. pond), and decreased with traditional pit latrine vs. flush toilet. (2) Alpha diversity varied by water source and latrine type without consistent trends. (3) Coliform bacteria relative abundance (RA) was higher among participants with traditional pit or ventilated improved pit latrines vs. flush toilet, and higher among participants relying on stream vs. pond water. Streptococcus agalactiae RA was higher among participants with non-flush toilets, while Bacteroides fragilis RA was lower with non-flush toilets. (4) Key taxa from non-targeted analyses associated with water source and latrine type included typical vaginal bacteria, opportunistic pathogens, and urinary tract pathobionts. (6) Increased rainfall was associated with decreased odds of CST-I. TRIAL REGISTRATION: ClinicalTrials.gov NCT03051789, February 14, 2017.

Adolescent

Crosstalk between cysteine and lysine modifications: Integrating redox and metabolic regulation.

Protein post-translational modifications (PTMs) on amino acid residues enable dynamic cellular responses to changes in metabolic and redox state. Cysteine and lysine are among the most extensively modified amino acid residues, with both undergoing a diversity of acylation and oxidative modifications. Indeed, proximal (<10&#x202f;&#xc5;) cysteine and lysine residues may form integration nodes for crosstalk between metabolism and redox homeostasis pathways. This review highlights the interaction of proximal Cys-Lys residues, including influence on residue pKa by local electrostatics, cysteine-to-lysine transfer of PTM moieties, and covalent crosslinking. We discuss candidate Cys-Lys regulatory pairs in proteins involved in redox regulation, proteostasis, metabolic adaptation and inflammation. We further utilize computational modeling to identify proximity between cysteine and lysine residues in proteins known to be regulated by acylation and oxidative PTMs, and to demonstrate changes in these distances and local electrostatic potential due to lysine acetylation. Finally, we review how mass spectrometry-based proteomics and machine-learning PTM predictive tools can enable the identification, validation, and interpretation of proximal Cys-Lys interactions that regulate cellular responses to oxidative challenge and metabolic flux.

Cysteine

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural

Nanocarrier-Based Gene Delivery Systems: Mechanisms, Clinical Translation, and Future Perspectives.

Gene therapy holds revolutionary potential for managing genetic disorders, cancers and infectious illnesses. However, one of the biggest challenges is delivering DNA or RNA into targeted cells and in the safe and effective way. In this review, nano carrier-based approaches for gene delivery are critically examined, focusing on both viral and non-viral systems. The advancement of CRISPR-Cas genome editing, machine learning-assisted nanocarrier optimization, and biologically inspired delivery systems is being quickly pushed forward in this area. In this review, a comparative analysis of gene delivery systems is being provided, and the key challenges to clinical translation are being pointed out. In addition, expert opinions on future research directions are being offered, with a heavy focus on the development of multifunctional, precisely targeted, and easily scalable delivery systems that can be integrated with next-generation therapeutic technologies.

Humans

What might echography learn from image science?

We review the current state of knowledge of the processes by which the information content of ultrasonic pulse-echo images is transferred to an observer, to the point of contributing to diagnostic judgments. As systematic knowledge in this specific field is rather sparse, we present relevant information and techniques derived from other areas of image science, both medical and otherwise. Quantitative measures both of the information content of ultrasonic and other images and of their characteristic noise content are first considered. An account is then given of the relevant aspects of human visual psychophysics, with particular reference to perception of contrast and detail, image texture, movement and colour, again with emphasis on documenting quantitative aspects of such behaviour. Against this background, we consider the efficiency, in current practice, of image information transfer to a human observer, how and to what extent this could be improved by changes in practice and, in particular, in what situations substantial innovations in machine processing of image data would be expected to improve human performance. It is suggested that several problems in the field may provide a worthwhile and challenging scope for future research.

Humans

Integrative multi-omics and single-cell analysis identifies EGFR pathway activation and metabolic reprogramming as potential synthetic lethal vulnerabilities in resistance to the FGFR inhibitor AZD4547.

BACKGROUND: Although fibroblast growth factor receptor (FGFR) inhibitors (FGFRi) have demonstrated clinical promise, the inevitable emergence of acquired resistance remains a critical bottleneck, severely compromising their long-term clinical efficacy. The pan-cancer molecular landscape and heterogeneous mechanisms driving this resistance, ranging from genetic alterations to dynamic network rewiring, remain poorly understood. METHODS: We integrated large-scale pharmacogenomic profiling of the FGFR inhibitor AZD4547 from the GDSC2 and PRISM databases with single-cell RNA sequencing to dissect the multi-omics landscape of FGFRi resistance across 312 cell lines from 8 cancer types. This multi-omics framework was further extended by machine learning modeling and systematic synthetic lethality screening to uncover actionable therapeutic targets. In vitro viability assays and western blot analysis were subsequently conducted to experimentally evaluate the predicted FGFR-EGFR synthetic lethality. RESULTS: Our dual-database analysis unveiled a multi-dimensional atlas of FGFRi resistance. We identified cancer-specific genomic drivers, such as ELF4 amplification in glioblastoma, alongside key transcriptomic markers including UCP2 and FSCN1, highlighting a shift towards metabolic reprogramming and epithelial-mesenchymal transition (EMT). Single-cell analysis unveiled that resistance is linked to the heterogeneous enrichment of baseline subpopulations characterized by distinct metaprograms, including cell-cycle dysregulation. Furthermore, a random forest model built on a LASSO-derived transcriptomic signature was constructed, demonstrating promising predictive capability for AZD4547 sensitivity (mean test-set AUC&#x2009;=&#x2009;0.73, 95% CI [0.63, 0.80]); the signature generalized well to erdafitinib but showed limited transferability to some other FGFR inhibitors (e.g. pemigatinib, BGJ398). Most notably, our synthetic lethal screening revealed a convergent reliance on compensatory RTK signaling (specifically EGFR pathway enrichment) and downstream MAPK/PI3K cascades in resistant phenotypes, providing converging computational evidence for EGFR pathway activation as an adaptive bypass mechanism. This predicted synthetic lethality was experimentally supported in two FGFR-dependent cell line models (RT112 and CCLP1), in which combined FGFR-EGFR inhibition produced marked synergistic antiproliferative effects. CONCLUSIONS: This study establishes a comprehensive multi-omics atlas of resistance to the FGFR inhibitor AZD4547, delineating convergent mechanisms of metabolic reprogramming and EGFR-mediated bypass signaling. Our findings characterize the resistance as a dynamic network rewiring and nominate rational combination strategies to overcome this therapeutic bottleneck. While FGFR-EGFR co-inhibition is experimentally supported, metabolic co-targeting remains a computationally derived, hypothesis-generating strategy.

Benzamides

Common to rare transfer learning (CORAL) enables inference and prediction for a quarter million rare Malagasy arthropods.

DNA-based biodiversity surveys result in massive-scale data, including up to millions of species-of which, most are rare. Making the most of such data for inference and prediction requires modeling approaches that can relate species occurrences to environmental and spatial predictors, while incorporating information about their taxonomic or phylogenetic placement. Even if the scalability of joint species distribution models to large communities has greatly advanced, incorporating hundreds of thousands of species has not been feasible to date, leading to compromised analyses. Here we present a 'common to rare transfer learning' (CORAL) approach, based on borrowing information from the common species to enable statistically and computationally efficient modeling of both common and rare species. We illustrate that CORAL leads to much improved prediction and inference in the context of DNA metabarcoding data from Madagascar, comprising 255,188 arthropod species detected in 2,874 samples.

Animals

[Difficulties of describing human relationships with the help of psychoanalytic descriptive models: a critique of ego-psychology].

Progress of psychotherapy and of related behaviour sciences makes evident the importance of a better understanding of human relations. But psychoanalysis finds it hard to describe interpersonal processes without transference. In order to remain within the conceptional frame of metapsychology it has to see interaction between individuals as the oral, aggressive or sexual cathexis of an object or as satisfaction or denial of the subject by the object. The structure of the "ego", which--in analogy to medical thinking--is conceived as an organ with its functions, is considered to have no interpersonal activities. The "ego" of the classic psychoanalytic theory is chiefly occupied with itself. It has to care for its egoistical interests and to guarantee its self-preservation. As an auxiliary and meanwhile popular concept the "self" has been introduced to describe object-relations. This concept is not sharply defined. Due to its metapsychological implications it produces additional theoretical difficulties. Linguistic studies show that every inventory of words implies a certain insight into reality. For this reason the metapsychological machine-like concept of psychic structures does not permit new ideas about interpersonal relations. If we leave metapsychology and base on colloquial speech we see that the experience of "I" is much more related to persons than the rather autistic concept of the "ego" shows. Further we learn that self-preservation cannot be an egoistical interest; it depends on the attachment to others. All feelings of self-esteem depend much more on interpersonal relations than on "narcissistic regulations". From these experiences three conclusions are derived: a) One of the main qualities of the ego is the relatedness to persons. b) The concept of narcissistic regulation as a successor of primary narcissism is no longer useful. Narcissistic traits develop as the secundary compensations if the individual failed to build up satisfactory interpersonal relations. c) The revision of (a) ego-psychology and (b) theory of narcissism asks for modifications of the therapeutic technique, where now the interest is especially concentrated on interpersonal problems instead on the pathology of the ego.

Drive