PubMed HealthSearch

SEARCH · PubMed Health

Results for “Large Language Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning

GBFN: A gated bimodal fusion network leveraging foundation model embeddings for cancer drug sensitivity prediction.

Despite recent progress in deep learning for cancer drug sensitivity prediction, many existing models still rely on task-specific representation learning or relatively simple multimodal fusion, which may limit their ability to capture complex drug-cell interactions. To address this issue, we developed GBFN, a gated bimodal fusion network for continuous IC50 prediction that integrates pretrained drug and cell-line representations. Specifically, drug embeddings were obtained from SMI-TED, whereas cell-line embeddings were derived from transcriptomic profiles using BulkFormer. These two modalities were then combined through a dimension-wise gated fusion module and used to predict IC50 values in matched drug-cell line pairs. On the CCLE-based benchmark, GBFN outperformed representative neural baselines, including GraphDRP, TGSA, and TransEDRP, and achieved the best overall performance, with an R² of 0.8714 and an RMSE of 0.8938. Moreover, ablation analysis showed that the model using drug features and cell-line expression data with gated fusion performed better than the corresponding model using direct concatenation, indicating that the improvement was associated with the fusion strategy rather than with the input modalities alone. In addition, cell-line expression data were more informative than mutation data in the present setting, and adding mutation data to the model using drug features and expression data did not further improve performance. Across major cancer types, GBFN maintained generally high cell-line-level predictive performance, and perturbation-based attribution identified biologically relevant transcriptomic programs in selected drug-cell line settings. Together, these findings support GBFN as a compact and effective framework for continuous drug response prediction.

Humans

Deep learning-based annotation of plant abiotic stress resistance genes for crops.

The declining costs of DNA sequencing have expanded genomic data, crucial for understanding plant abiotic stress responses and crop improvement. However, accurate gene annotation remains challenging. To address this limitation, we propose the PASRGA, a deep learning approach that leverages transfer learning and contrastive learning to annotate genes related to drought, salt, cold, and UV resistance. PASRGA achieves high F1-scores, area under the receiver operating characteristic (AUROC), area under the precision-recall curve (AUPRC), and Matthews correlation coefficient (MCC) in annotating stress resistance genes, significantly outperforming the general protein annotation model CLEAN, the plant phosphatase gene annotation model PF-NET, the top-ranked model in the CAFA5 challenge NetGO 4.0, and four traditional machine learning methods. Its effectiveness was further validated with a salt stress treatment experiment in Eutrema salsugineum. To facilitate crop breeding practices, we utilized PASRGA to annotate the genomes of 17 major crops. To improve accessibility and utility, we incorporated both manually curated and PASRGA-predicted gene data, together with the PASRGA tool, into the PlantASRG database (https://bioinfor.nefu.edu.cn/PlantASRG/). This comprehensive resource aims to support crop breeding initiatives and ensure food security.

Crops, Agricultural

mamp-ml: A deep learning approach to epitope immunogenicity in plants.

Eukaryotes detect biomolecules through surface-localized receptors, key signaling components. A subset of receptors survey for pathogens, induce immunity, and restrict pathogen growth. Comparative genomics of both hosts and pathogens has unveiled vast sequence variation in receptors and potential ligands, creating an experimental bottleneck. We have developed mamp-ml, a machine learning framework for predicting plant receptor-ligand interactions. We leveraged existing functional data from over two decades of foundational research, together with the large protein language model ESM-2, to build a pipeline and model that predicts immunogenic outcomes using a combination of receptor-ligand features. Our model achieves 73% prediction accuracy on a held-out test set, even when an experimental structure is lacking. Our approach enables high-throughput screening of LRR receptor-ligand combinations and provides a computational framework for engineering plant immune systems.

Journal Article

Agentic AI for Spatial Omics.

This highlight summarises recent advances in agentic artificial intelligence (AI) systems for spatial omics analysis. These systems are compared along two central tensions: autonomy versus accountability, and adaptability versus reproducibility. We argue that progress will depend not on maximising automation, but on defining where autonomy is appropriate.

Artificial Intelligence

Hard to Halt: Automation Bias in Agent-Driven Sequencing Prior Authorization Workflows.

PURPOSE: Prior authorization (PA) for exome or genome sequencing is a time-consuming process that impedes timely rare disease diagnosis. Large language model-based browser agents offer potential for automating these workflows, but their clinical reliability remain uncharacterized. METHODS: We developed a sandbox compromising a simulated ES/GS PA submission payer portal and a synthetic EHR containing 836 patient records spanning compliant profiles and deficient profiles with different types of issues. Gemini 3 Pro, Gemini 3 Flash, and Claude Opus 4.5 were evaluated on task completion rate, form completion accuracy, and appropriate withholding for deficient profiles. RESULTS: Larger models achieved much higher task completion rates (Gemini 3 Pro 95.45%, Claude Opus 4.5 93.67%) compared to Gemini 3 Flash (56.05%), but nearly universally failed to withhold submission for deficient profiles whereas Gemini 3 Flash ironically demonstrated superior withholding performance (17.33%). In a non-agentic setting, Gemini 3 Pro correctly identified 91% of the issues in deficient profiles, indicating that withholding failure is attributable to the browser interaction rather than the model's reasoning limitations. CONCLUSION: Current LLM-based browser agents exhibit a systematic bias towards form submission that poses risks in PA workflows. A modular, multi-agent architecture with human supervision is necessary for a safe clinical deployment.

Journal Article

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2-4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms-Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans

Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models.

Genomic language models (gLMs) have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences. However, standard gLMs adapted from natural language processing often require extremely large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks. Here, we introduce GPN-Star (Genomic Pretrained Network with Species Tree and Alignment Representation), a biologically grounded gLM featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammalian, and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales reveal task-dependent advantages of modeling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms prior methods in prioritizing pathogenic and fine-mapped GWAS variants; yields unprecedented enrichments of complex trait heritability; and improves power in rare variant association testing. Extending beyond humans, we train GPN-Star for five model organisms - Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana - demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful, and flexible new tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Protein Language Model Decoys for Target Decoy Competition in Proteomics: Quality Assessment and Benchmarks.

Large-scale proteomics relies heavily on target-decoy competition for false discovery rate estimation in peptide identification, and the performance of this strategy depends strongly on the design of the decoy database. Classical generators such as reversal and shuffling remain widely used. Here, we introduce the first protein language model-based (PLM) decoy generation for peptide identification and benchmark it against classical strategies. We evaluate these approaches using three complementary quality-control layers: sequence-based separability, search-engine-agnostic spectral-space diagnostics, and end-to-end mass spectrometry benchmarks, including pipelines with rescoring. Across these analyses, PLM-based decoys are harder for sequence-only neural networks to distinguish than most classical generators, suggesting fewer obvious sequence-level artifacts. However, this signal is only weakly informative for search performance. Spectral diagnostics further show that short peptides occupy a particularly crowded target-decoy space and are therefore especially prone to local collisions across all generators. In full search pipelines, reverse decoys remain a strong baseline, and current PLM-based generators do not yet provide a clear overall advantage. We therefore view PLM-based decoys not as universal replacements for reverse decoys but as tunable tools for benchmarking, diagnostics, stress testing, and future adaptive decoy optimization, with increasing value as search models become more expressive.

Proteomics

Predicting the First Onset of Suicidal Thoughts and Behaviors in Adolescents Using Multimodal Risk Factors: A 4-Year Longitudinal Study.

OBJECTIVE: Suicide is one of the leading causes of death among youth worldwide, yet existing studies that aimed to predict the first onset of suicidal thoughts and behaviors (STB) included a limited number of data modalities and/or focused on adult populations. This study aimed to prospectively predict first-onset STB across 4-year follow-ups in adolescents using an existing STB history classification model that was previously applied to baseline data and a new machine learning model with 195 biopsychosocial features. METHOD: Participants were 7,503 unrelated adolescents (54.5% female, ages 9-11 years at baseline) from the multisite, longitudinal Adolescent Brain Cognitive Development (ABCD) Study. An existing baseline STB history classification model was applied to predict longitudinal first-onset STB in adolescents compared with healthy controls and clinical controls (individuals with a mental health disorder but no STB). A new elastic net logistic regression model with 195 features was trained on data from 14 sites (n = 5,220), and the resulting top 15 features were validated at 7 independent sites (n = 2,283). RESULTS: The previously developed model to classify STB lifetime history also prospectively predicted first-onset STB in adolescents with an area under the curve (AUC) [95% CI] of 0.73 [0.70, 0.75], p < .001, compared with healthy controls and AUC [95% CI] of 0.63 [0.60, 0.66], p < .001, compared with clinical controls. The newly trained model with top 15 features performed similarly with AUC [95% CI] of 0.73 [0.71, 0.76], p < .001, and AUC [95% CI] of 0.64 [0.60, 0.66], p < .001, for the same comparison groups. The most consistent predictors across models included female sex, sleep disturbances, and maladaptive home and school environments. CONCLUSION: The models predicted first-onset STB in adolescents with moderate accuracy. This study also confirmed the roles of well-established psychological risk factors for STB and identified several novel neurocognitive and brain imaging risk factors. Future studies should validate these models in large-scale diverse samples before clinical translation. PLAIN LANGUAGE SUMMARY: This study followed over 7,500 adolescents for 4 years and tested 2 machine learning models using psychological, social, and brain data to identify those at risk of experiencing suicidal thoughts or behaviors. Both models predicted first-time suicidal thoughts or behaviors with moderate accuracy. Key risk factors that were identified included being female, experiencing sleep problems, and negative home and school environments. DIVERSITY & INCLUSION STATEMENT: We worked to ensure sex and gender balance in the recruitment of human participants. We worked to ensure race, ethnic, and/or other types of diversity in the recruitment of human participants. We worked to ensure that the study questionnaires were prepared in an inclusive way. Diverse cell lines and/or genomic datasets were not available. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented racial and/or ethnic groups in science. One or more of the authors of this paper self-identifies as a member of one or more historically underrepresented sexual and/or gender groups in science. We actively worked to promote sex and gender balance in our author group. One or more of the authors of this paper received support from a program designed to increase minority representation in science. We actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our author group. While citing references scientifically relevant for this work, we also actively worked to promote sex and gender balance in our reference list. While citing references scientifically relevant for this work, we also actively worked to promote inclusion of historically underrepresented racial and/or ethnic groups in science in our reference list. The author list of this paper includes contributors from the location and/or community where the research was conducted who participated in the data collection, design, analysis, and/or interpretation of the work.

Adolescent

Feature expressions: creating and manipulating sequence datasets.

Annotation of features, such as introns, exons and protein coding regions in GenBank/EMBL/DDBJ entries is now standardized through use of the Features Table (FT) language. The essence of the FT language is described by the relation 'expression-->sequence', meaning that each FT expression evaluates to a sequence. For example, the expression M74750:1..50 evaluates to the first 50 bases of the sequence with accession number M74750. Because FT is intrinsic to the database definition, it can serve as a software- and platform-independent lingua franca for sequence manipulation. The XYLEM package makes it possible to create and manipulate sequence datasets using FT expressions. FEATURES is a program that resolves FT expressions into their corresponding sequences. Annotated features can be retrieved either by feature key or by expression. Even unannotated portions of a sequence can be retrieved by user-generated FT expressions. Applications of the FT language include retrieval of subsequences from large sequence entries, generation of chromosome models or artificial DNA constructs, and representation of restriction maps or mutants.

Base Sequence

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article

Models of natural language understanding.

This paper surveys some of the fundamental problems in natural language (NL) understanding (syntax, semantics, pragmatics, and discourse) and the current approaches to solving them. Some recent developments in NL processing include increased emphasis on corpus-based rather than example- or intuition-based work, attempts to measure the coverage and effectiveness of NL systems, dealing with discourse and dialogue phenomena, and attempts to use both analytic and stochastic knowledge. Critical areas for the future include grammars that are appropriate to processing large amounts of real language; automatic (or at least semi-automatic) methods for deriving models of syntax, semantics, and pragmatics; self-adapting systems; and integration with speech processing. Of particular importance are techniques that can be tuned to such requirements as full versus partial understanding and spoken language versus text. Portability (the ease with which one can configure an NL system for a particular application) is one of the largest barriers to application of this technology.

Cognition

Answering the connectionist challenge: a symbolic model of learning the past tenses of English verbs.

Supporters of eliminative connectionism have argued for a pattern association-based explanation of language learning and language processing. They deny that explicit rules and symbolic representations play any role in language processing and cognition in general. Their argument is based to a large extent on two artificial neural network (ANN) models that are claimed to be able to learn the past tenses of English verbs (Rumelhart & McClelland, 1986, Parallel distributed processing, Vol. 2, Cambridge, MA: MIT Press; MacWhinney & Leinbach, 1991, Cognition, 40, 121-157). In this article we critically review Rumelhart and McClelland's as well as MacWhinney and Leinbach's ANN models and conclude that they do not succeed in the assigned task of learning the past tenses of English verbs. In order to answer their challenge to the symbolic processing approach, we present our symbolic pattern associator (SPA)-a general-purpose pattern associator that can learn to associate arbitrary discrete patterns. We carried out several experiments with the SPA using the same set of verbs that was used in MacWhinney and Leinbach's simulation with more realistic training and testing procedures. The SPA outperformed the connectionist models by a wide margin in the accuracy of learning, and successful inductive generalizations to unseen verbs. Our SPA has very natural and psychologically realistic explanations to many psychological effects such as U-shaped learning curve, and is much closer to human subjects in predicting past tense of the pseudo-verbs. In contrast to ANNs, whose internal representations are entirely opaque, the SPA can represent the acquired knowledge in the form of production rules that allow for further higher-level processing and integration, resulting in linguistically realistic associative templates for irregular verbs and production rules for regular verbs. In the light of these findings, we conclude that eliminative connectionists' vision of cognition as simple pattern association and pattern recognition without symbolic representation is inadequate. Pattern association as such does not imply rule-less or cue-based models of language acquisition or of human learning in general.

Cognition

Image library of biological macromolecules.

An Image Library of Biological Macromolecules is described, which contains image and text files related to structures of biological macromolecules. Currently, the Library has approximately 3000 image files of approximately 300 structures of biological macromolecules whose coordinates are available in the Protein Data Bank and in the Nucleic Acid Database. The entries include all RNA structures, approximately 70 DNA structures, 150 proteins and a few carbohydrates. The Library contains further images of amino acids, of standard and modified nucleotides and of nucleic acid model structures. Each entry consists of an annotation file with bibliographic and sequence information and possibly comments, of a color-coded distance plot and of structure images. Almost all of the images are available both in a mono and in a stereo representation. Standard procedures for generating these images were strictly avoided. Therefore, mixed rendering, coloring and labeling techniques were used extensively. Since May 1995 the Library has a growing division of images in the new Virtual Reality Modeling Language (VRML) format. The Image Library of Biological Macromolecules can be accessed via the World-Wide Web (http://www.imb-jena.de/IMAGE.html). There is a large number of structures determined by experimental and/or modeling techniques which are not intended to be included into the Protein Data Bank or Nucleic Acid Database for some reason. The Image Library could be a repository of these structures and of images of these and other structures of biological macromolecules including structures which are not known at atomic detail. Authors who are willing to make available images or coordinates to the scientific community via the Image Library of Biological Macromolecules are requested to contact the author.

Computer Communication Networks