PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Deep learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

PAT: An Image Analysis Tool for Automated Scoring of Pollen in Alexander-Stained Anthers.

Quantitative pollen viability analysis is a critical but labor-intensive step in plant reproductive biology. Existing deep-learning Segment Anything Models (SAM) fail to reliably segment viable pollen in Alexander-stained anthers. To address this, we fine-tuned an existing Cellpose-SAM model for pollen segmentation. We integrated it into PAT (Pollen Analysis Tool), a cross-platform desktop application. PAT features instance segmentation with interactive quality control, an in-app model retraining module, and publication-ready statistical outputs. We deployed PAT in an EMS suppressor screen of semi-sterile Arabidopsis smg7-6 mutants, enabling efficient candidate prioritization for whole-genome sequencing and mapping of the candidate mutation. This screen led to the identification of a point mutation in CAP-D2 (capd2-2), a Condensin I subunit, that rescues the smg7-6 meiotic phenotype. Notably, mutation in a Condensin II subunits (CAP-D3 and CAP-H2) does not confer rescue. Further characterization suggests the capd2-2 allele is hypomorphic, showing no defects in vegetative growth, chromocenter compaction, or transposable element silencing. Collectively, we demonstrate that accessible AI tools have the potential to bridge gaps in plant phenotyping and accelerate the pace of biological discovery.

Alexander staining↗

H&E to recurrence score: A step forward, but not yet a substitute for genomic testing.

Shamai and colleagues developed a multimodal deep-learning model that predicts Oncotype DX recurrence scores from routine H&E slides and clinicopathological variables in hormone receptor‑positive, HER2‑negative early breast cancer. Validated across the TAILORx trial and six external cohorts (over 5000 patients), the model achieved an AUC of 0.898 for identifying recurrence score ≥26 and recapitulated genomic assay patterns of chemotherapy benefit. Notably, 31% of clinically high-risk postmenopausal women were downgraded to low risk by AI, suggesting potential to reduce overtreatment. However, several limitations preclude immediate clinical substitution for genomic testing. First, intratumoural heterogeneity leads to discordant predictions with unclear management guidance. Second, the model's chemotherapy benefit estimates rely on TAILORx's age-based menopausal surrogates, which may not reflect real-world hormonal status or LHRH agonist use. Third, predictive value in node-positive disease remains untested in randomised datasets such as RxPONDER. Additionally, calibration uncertainty near risk thresholds and global scalability issues (including IHC requirements and digital pathology infrastructure) persist. While this represents a landmark step toward democratising precision oncology, the AI tool should currently serve as a complementary decision aid, with genomic testing remaining the gold standard for intermediate, borderline, or discordant cases.

Breast cancer↗

Histology-Based Virtual RNA Inference Identifies Pathways Associated With Metastasis Risk in Colorectal Cancer.

Colorectal cancer (CRC) remains a major health concern, with >150,000 new diagnoses and >50,000 deaths annually in the United States, underscoring an urgent need for improved screening, prognostication, disease management, and therapeutic approaches. The tumor microenvironment (TME)-comprising cancerous and immune cells interacting within the tumor's spatial architecture-plays a critical role in disease progression and treatment outcomes, reinforcing its importance as a prognostic marker for metastasis and recurrence risk. However, traditional methods for TME characterization, such as bulk transcriptomics and multiplex protein assays, lack sufficient spatial resolution. Although spatial transcriptomics (ST) allows for the high-resolution mapping of whole transcriptomes at near-cellular resolution, current ST technologies (eg, Visium and Xenium) are limited by high costs, low throughput, and issues with reproducibility, preventing their widespread application in large-scale molecular epidemiology studies. In this study, we refined and implemented virtual RNA inference (VRI) to derive ST-level molecular information directly from hematoxylin and eosin (H&E)-stained tissue images. Our VRI models were trained on the largest matched CRC ST data set to date, comprising 45 patients and >300,000 Visium spots from primary tumors. Using state-of-the-art deep learning models (UNI, ResNet-50, Vision Transformer, and Vision Mamba), we achieved a median Spearman's correlation coefficient of 0.546 between predicted and measured spot-level expression. As validation, VRI-derived gene signatures linked to specific tissue regions (tumor, interface, submucosa, stroma, serosa, muscularis, and inflammation) showed strong concordance with signatures generated via direct ST, and VRI performed accurately in estimating cell-type proportions spatially from H&E slides. In an expanded CRC cohort controlling for tumor invasiveness and clinical factors, we further identified VRI-derived gene signatures significantly associated with key prognostic outcomes, including metastasis status. Although certain tumor-related pathways are not fully captured by histology alone, our findings highlight the ability of VRI to infer a wide range of "histology-associated" biological pathways at near-cellular resolution without requiring ST profiling. Future efforts will extend this framework to expand TME phenotyping from standard H&E tissue images, with the potential to accelerate translational CRC research at scale.

Humans↗

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans↗

GBFN: A gated bimodal fusion network leveraging foundation model embeddings for cancer drug sensitivity prediction.

Despite recent progress in deep learning for cancer drug sensitivity prediction, many existing models still rely on task-specific representation learning or relatively simple multimodal fusion, which may limit their ability to capture complex drug-cell interactions. To address this issue, we developed GBFN, a gated bimodal fusion network for continuous IC50 prediction that integrates pretrained drug and cell-line representations. Specifically, drug embeddings were obtained from SMI-TED, whereas cell-line embeddings were derived from transcriptomic profiles using BulkFormer. These two modalities were then combined through a dimension-wise gated fusion module and used to predict IC50 values in matched drug-cell line pairs. On the CCLE-based benchmark, GBFN outperformed representative neural baselines, including GraphDRP, TGSA, and TransEDRP, and achieved the best overall performance, with an R² of 0.8714 and an RMSE of 0.8938. Moreover, ablation analysis showed that the model using drug features and cell-line expression data with gated fusion performed better than the corresponding model using direct concatenation, indicating that the improvement was associated with the fusion strategy rather than with the input modalities alone. In addition, cell-line expression data were more informative than mutation data in the present setting, and adding mutation data to the model using drug features and expression data did not further improve performance. Across major cancer types, GBFN maintained generally high cell-line-level predictive performance, and perturbation-based attribution identified biologically relevant transcriptomic programs in selected drug-cell line settings. Together, these findings support GBFN as a compact and effective framework for continuous drug response prediction.

Humans↗

GICPIdb: an archival repository of multimodal data focusing on pathological images for gastrointestinal cancers.

INTRODUCTION: Deep learning (DL) shows great potential for predicting biomarkers from routine histopathological slides of gastrointestinal (GI) cancers. Yet most existing models are validated on limited patient cohorts, while pathological image annotation and molecular marker standardization demand substantial professional expertise. To address these gaps, we constructed the Gastrointestinal Cancer Pathological Image Archive (GICPIdb, gicpidb.shubuzuo.top), a dedicated database and web platform covering seven major GI cancer types. METHODS: High-quality hematoxylin and eosin (H&E)-stained whole-slide images were collected from multiple sources and uniformly processed. Image annotations were performed by board-certified pathologists following standardized protocols. GICPIdb offers five interactive web modules for data uploading, quality control, feature extraction, online annotation and AI-based prediction. Its intuitive interface supports data browsing, retrieval, visualization and downloading. RESULTS: The database houses 2,863 pathologist-annotated, uniformly processed, high-quality H&E stained images collected from 2,655 patients. Of these, 1,699 patients were sourced from The Cancer Genome Atlas (TCGA), 182 from the Clinical Proteomic Tumor Analysis Consortium (CPTAC), and 424 from China-Japan Friendship Hospital and 350 from Chifeng Municipal Hospital in Inner Mongolia, China. It also integrates data on over 50 key molecular markers (e.g., MSI, TMB) and prognostic labels related to survival, recurrence and metastasis. DISCUSSION: GICPIdb aims to promote the development of DL-driven AI tools for cancer research and clinical translation. The multi-institutional data collection and standardized annotation pipeline are expected to enhance the generalizability and reproducibility of AI-based prediction models across diverse patient populations.

deep learning↗

CNV-Finder: Streamlining Copy Number Variation Discovery.

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods-Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)-which may be relevant to neurological diseases such as Alzheimer's disease (AD), Parkinson's disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinson's Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finder's interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

Copy Number Variation (CNV)↗

A fast learning algorithm for deep belief nets.

We show how to use "complementary priors" to eliminate the explaining-away effects that make inference difficult in densely connected belief nets that have many hidden layers. Using complementary priors, we derive a fast, greedy algorithm that can learn deep, directed belief networks one layer at a time, provided the top two layers form an undirected associative memory. The fast, greedy algorithm is used to initialize a slower learning procedure that fine-tunes the weights using a contrastive version of the wake-sleep algorithm. After fine-tuning, a network with three hidden layers forms a very good generative model of the joint distribution of handwritten digit images and their labels. This generative model gives better digit classification than the best discriminative learning algorithms. The low-dimensional manifolds on which the digits lie are modeled by long ravines in the free-energy landscape of the top-level associative memory, and it is easy to explore these ravines by using the directed connections to display what the associative memory has in mind.

Algorithms↗

Student use and perceptions of different learning aids in a problem-based learning (PBL) dentistry course.

First-year dental students in a new problem-based learning (PBL) course, the Bachelor of Dentistry (BDent) Program at the University of Sydney, Australia, completed the Study Process Questionnaire and two other questionnaires in this study. The study aimed to identify student perceptions of a written formative assessment and the helpfulness of various learning aids used to prepare for this assessment and preparing to be a dental clinician. Correlations between approach to learning and perceptions of assessment and learning aids showed theoretically expected associations. Surface learning was associated with students' concerns regarding whether assessment items reflected curriculum content, a valuing of lectures as a learning aid, and low scores for theme sessions. Deep learning was associated with a perception that the assessment tested application of basic and clinical sciences and a valuing of both independent study groups and learning topics as learning aids. An achievement orientation to learning was associated with a valuing of formative assessment as a learning aid and an intention to modify study habits as a result of participating in formative assessment. The findings provide insight into student learning in a PBL context that will help teachers and curriculum developers better understand the value of teaching aids provided in the program and the impact assessment has on study styles.

Achievement↗

The conceptions of the nature of learning of first-year physiotherapy students and their relationship to students' learning outcomes.

Research shows that students' conceptions of what 'learning' is influences both their approaches to learning and their learning outcomes. Lower level conceptions of learning are associated with surface approaches to learning, encouraging poor quality learning outcomes. However, more advanced conceptions are associated with deep learning approaches and high-quality learning outcomes. The Structure of Observed Learning Outcomes (SOLO) taxonomy was used to analyse the conceptions of learning of a new intake of physiotherapy students. The presence of a relationship between these and assessment outcomes was also investigated. As in other studies, a majority of students had lower level conceptions of learning than desired in higher education. However, a larger proportion of students had higher levels of conception than has been found in other research. A direct relationship between conceptions of learning and learning outcomes was also identified. The implications of these findings for achievement of high-quality learning are discussed.

Journal Article↗

Multiple goals, motivation and academic learning.

BACKGROUND: The type of academic goals pursued by students is one of the most important variables in motivational research in educational contexts. Although motivational theory and research have emphasised the somewhat exclusive nature of two types of goal orientation (learning goals versus performance goals), some studies (Meece, 1994; Seifert, 1995, 1996) have shown that the two kinds of goals are relatively complementary and that it is possible for students to have multiple goals simultaneously, which guarantees some flexibility to adapt more efficaciously to various contexts and learning situations. AIM: The principal aim of this study is to determine the academic goals pursued by university students and to analyse the differences in several very significant variables related to motivation and academic learning. SAMPLE: Participants were 609 university students (74% women and 26% men) who filled in several questionnaires about the variables under study. METHOD: We used cluster analysis ('quick cluster analysis' method) to establish the different groups or clusters of individuals as a function of the three types of goals (learning goals, performance goals, and social reinforcement goals). By means of MANOVA, we determined whether the groups or clusters identified were significantly different in the variables that are relevant to motivation and academic learning. Lastly, we performed ANOVA on the variables that revealed significant effects in the previous analysis. RESULTS: Using cluster analysis, three groups of students with different motivational orientations were identified: a group with predominance of performance goals (Group PG: n = 230), a group with predominance of multiple goals (Group MG: n = 238), and a group with predominance of learning goals (Group LG: n = 141). CONCLUSIONS: Groups MG and LG attributed their success more to ability, they had higher perceived ability, they took task characteristics into account when planning which strategies to use in the learning process, they showed higher persistence, and used more deep learning strategies than did the students with predominance of performance goals (Group PG). On the other hand, Groups MG and PG took the evaluation criteria more into account when deciding which strategies to use in order to learn, and they attributed their failures more to luck than did Group LG. Students from Group MG attributed their success more to effort than did the other two groups and they attained higher achievement than Group PG. Group LG tended to attribute their failures more to lack of effort than did the other two groups.

Achievement↗

Mechanistic insights into Claudin-14 dysfunction implicated in veins of Galen malformation.

Claudin-14 (CLDN14) is a key component of tight junctions (TJs) critical for maintaining paracellular barrier function. Variants of CLDN14 have been linked to Vein of Galen malformations (VOGMs), a rare cerebrovascular disorder; however, the molecular mechanisms underlying their pathogenicity remain unknown. Here, we investigate the mechanistic effects of two VOGM-associated mutations, A113P and V143M, using reinforcement-learning driven enhanced sampling molecular dynamics simulations combined with DiffNets-based deep learning and independent trajectory-wide structural analyses. Our analysis reveals that A113P induces broader structural disruption of CLDN14, perturbing paracellular sealing, pore symmetry, and inter-protomer communication, whereas V143M induces structural rearrangements centred around TM3 and the TM3-ECL2 region. In both cases, mutation-specific alterations are observed in structural stability and interfacial organization across oligomeric assemblies. Notably, these effects are qualitatively consistent across different modelled architectures, despite variability in local responses. In the absence of experimentally resolved structures, the structural perturbations reported here provide a mechanistic understanding of how VOGM-associated variants may influence CLDN14 structure and dynamics.

Aneurysm↗

Cancer of unknown primary: the evolution of tissue of origin identification in the artificial intelligence era.

Cancer of Unknown Primary (CUP) presents substantial diagnostic and therapeutic challenges owing to its heterogeneous nature and the absence of an identifiable primary tumor site. This review provides a structured search of the pathogenesis, epidemiological characteristics, and limitations of traditional diagnostic and therapeutic approaches for CUP, with an emphasis on the evolution of Tissue of Origin (TOO) identification techniques. Recent advances in precision medicine have accelerated the development of machine learning-based TOO identification tools, representing a paradigm shift in CUP diagnostics. Deep learning (DL) algorithms that integrate multi-omics data (such as genomics and transcriptomics) with clinical features have markedly enhanced the accuracy of tracing tumor origin, and artificial intelligence (AI) driven TOO models are increasingly being incorporated into clinical practice, offering new insights for pathological diagnosis, treatment selection, and prognostic evaluation. Nevertheless, several challenges remain, including issues of data standardization, model generalizability, and interpretability. Ethical considerations related to data privacy, algorithmic fairness, and clinical implementation also warrant careful attention. Future research should focus on establishing standardized multi-center databases, developing more interpretable AI models, and fostering multidisciplinary collaborative strategies for CUP management. Through continued refinement of technical solutions and regulatory guidelines, TOO identification is anticipated to progress from research to routine clinical application, ultimately supporting precise and personalized care for patients with CUP.

Artificial intelligence↗

Development and Validation of Machine Learning Models for Predicting Early Cognitive Decline Using Home Sensor-Derived Behavioral Data: Sensors in-Home for Elder Wellbeing (SINEW) Cohort Study.

BACKGROUND: As the global population continues to age, the prevalence of geriatric conditions, including dementia and frailty, is also increasing. Early identification of individuals at an elevated risk of these conditions, such as those presenting with mild cognitive impairment (MCI) or prefrailty, can provide a critical window for prompt intervention aimed at preventing or reversing disease progression. To promote such early identification, there is a burgeoning interest in the use of digital sensor technology and predictive modeling. OBJECTIVE: This study aimed to use a continuous, home-based monitoring sensor system for older adults to distinguish those exhibiting normal aging from those with MCI, early dementia, prefrailty, or frailty, and to predict their transition from normal aging to one of these conditions. METHODS: This longitudinal cohort study will recruit 200 community-dwelling adults aged ≥65 years with normal cognition or MCI at baseline. A multi-sensor system will be installed in participants' homes, including passive infrared motion sensors, door contact sensors, bed sensors, medication box sensors, wearable activity bands, and Bluetooth proximity beacons. These devices will continuously capture spatiotemporal activity patterns, mobility indicators, sleep behaviors, and medication-taking routines. Annual assessments will include standardized cognitive tests (eg, Montreal Cognitive Assessment, Mini-Mental State Examination, Rey Auditory-Verbal Learning Test, digit span, Color Trails Test, semantic fluency, Stroop), frailty measures (modified Fried phenotype, gait speed, grip strength), mental health scales, sleep quality, and psychosocial indicators. Sensor-derived features-such as gait variability, activity regularity, sleep fragmentation, and medication adherence patterns-will be integrated with clinical data to develop supervised machine learning models. Planned approaches include logistic regression, random forests, gradient boosting, and deep learning. Model performance will be evaluated using cross-validation and independent test sets. Primary metrics will include area under the receiver operating characteristic curve, sensitivity, specificity, precision, recall, and F1-score. Models will be benchmarked against gold-standard clinical diagnoses and validated using temporal subsets of the dataset. RESULTS: Enrollment for this study started in November 2019 and will continue until March 2030. As of June 2025, we have enrolled 138 participants. Full data analysis has yet to begin. CONCLUSIONS: We aim to develop a reliable and effective sensor system for in-home use that will facilitate the early detection of cognitive and physical decline. In so doing, it will add to our current understanding of digital biomarkers. It is common for older adults to seek clinical intervention only when their cognitive impairment has already reached an advanced stage. The implementation of readily deployable sensor systems within community settings presents us with opportunities for prompt intervention, which holds the potential for delaying or reversing disease progression and allowing for a greater number of functional and meaningful years.

Humans↗

A relational view of learning: implications for nurse education.

For about two decades, student learning in the context of professional and higher education, more generally, has been investigated by groups of researchers in countries such as Sweden, the UK, Australia and South Africa. The focus in many of these studies has been on the experience of students as they undertake complex, realistic tasks such as reading academic articles, listening to lectures, writing essays, solving problems and learning subject matter concepts. Central to this research is the adoption of a relational and holistic model of learning in which the relationships among all elements in the learning situation, including the student, the learning task, the teaching methods and assessment practices have been investigated. In particular, researchers from this line of inquiry have studied the relationship between the ways students go about learning in natural educational settings, referred to as approaches to learning, and what they learn. They have established that a full understanding of subject matter is reliant on students employing deep learning approaches which, in turn, depend on the perceived learning environment encouraging such learning. In this paper, the work of the relational theorists is described and the substantial and practical findings arising from this school of research are reported. Specific implications for the education of undergraduate nurses are also outlined.

Clinical Competence↗

Facilitation of students' discussion in problem-based learning tutorials to create mechanisms: the use of five key questions.

Without the appropriate facilitation of discussion in a problem-based learning (PBL) course and the use of specific educational tools that enhance cognitive skills, students might deprive themselves of achieving the deep learning experience expected to take place in a PBL course. One of the educational tasks in PBL is the creation of mechanisms for hypotheses made by the students, based on their knowledge of the basic sciences and the psychosocial issues raised in a particular case scenario. The whole task is student-constructed and should enhance their ability to explain the scientific basis of the symptoms and clinical signs of the patient enlisted in the case. Because students usually discuss the case without enough prior related knowledge, they might find it difficult to address different aspects of their mechanisms. These gaps in knowledge may be considered part of their "learning issues". In tutorial 2 (a PBL case is usually discussed in 2 or 3 tutorials at the maximum; each tutorial is 2 hours long), students should be able to build a comprehensive mechanism reflecting their deep understanding of the problem. However, students might not be able to integrate information learnt and their mechanisms might show a number of shortcuts and/or lack integration of information, and the flow of the pathophysiological changes may not be logical. This manuscript describes 5 key open-ended questions in PBL tutorials to facilitate students' discussions as they create their mechanisms.

Education, Medical, Undergraduate↗

Attitudes to concept maps as a teaching/learning activity in undergraduate health professional education: influence of preferred learning style.

Concept maps that integrate and relate concepts in a nonlinear fashion are widely accepted as an educational tool that can underpin meaningful learning in medical education. However, student take-up may be affected by a number of cognitive and non-cognitive influences. In the present study, student attitudes to pre-prepared concept maps introduced in Stage 2 conjoint MPharm and BSc Pharmacology lectures were examined in relation to preferred learning styles according to the Felder-Silverman model. There was no statistically significant influence of dichotomous learning style dimension (sensing/intuitive; visual/verbal; active/reflector; sequential/global) on the self-reported utility of such concept maps to learning. However, when strength of preference was analysed within each dimension, moderate/strong verbal learners were found to be significantly less likely to self-report concept maps as useful relative to mild verbal learners. With this important exception, these data now suggest that student attitudes to concept maps are broadly not influenced by preferred learning styles and furthermore highlight the potential of concept maps to address a variety of different learning styles and thereby facilitate 'teaching to all types'. Concept maps could therefore potentially assist motivation, engagement and deep learning in medical and biomedical science education when used as a supplement to more traditional teaching/learning activities.

Concept Formation↗

ProMeta: a meta-learning framework for robust disease diagnosis and prediction from plasma proteomics.

MOTIVATION: The plasma proteome offers a dynamic window of human health, capturing the real-time intersections between genetics and physiology. However, the application of deep learning to proteomics is currently hindered by a reliance on large-scale labeled datasets, rendering standard models ineffective for rare or novel diseases where patient samples are inherently scarce. RESULTS: Here, we present ProMeta, a meta-learning framework designed to enable robust disease modeling under extreme data restrictions. By integrating knowledge-guided pathway encoding with bi-level meta-optimization, ProMeta projects unstructured proteomic profiles into biologically interpretable functional tokens. This architecture allows the model to learn a global initialization containing transferable biological priors from biobank-scale data, facilitating rapid adaptation to novel tasks. Through comprehensive benchmark experiments, ProMeta consistently outperformed transfer learning and traditional machine learning baselines in both disease diagnosis and prediction tasks. In the most challenging 4-shot scenarios (utilizing only 2 cases and 2 controls), the model achieved robust generalization with an average AUROC of ∼0.69, representing a 24.6% relative improvement over the best-performing baseline methods. Mechanistic investigation revealed that ProMeta disentangles cases from controls in the latent space prior to task-specific adaptation, confirming the acquisition of universal biological rules rather than rote memorization. Furthermore, gradient-based interpretation identified disease-specific protein biomarkers and functional pathways consistent with known pathophysiology. Collectively, ProMeta overcomes the data-scarcity bottleneck in precision medicine, providing a scalable, interpretable framework for characterizing the full spectrum of human diseases, particularly for rare conditions lacking extensive clinical cohorts. AVAILABILITY AND IMPLEMENTATION: The source code of ProMeta is available at GitHub (https://github.com/lihan97/ProMeta).

Proteomics↗