PubMed HealthSearch

SEARCH · PubMed Health

Results for “Protein Language Model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

GENPRO: automatic generation of Prolog clause files for knowledge-based systems in the biomedical sciences.

With the increasing interest in using knowledge-based approaches for protein structure prediction and modelling, there is a requirement for general techniques to convert molecular biological data into structures that can be interpreted by artificial intelligence programming languages (e.g. Prolog). We describe here an interactive program that generates files in Prolog clausal form from the most commonly distributed protein structural data collections. The program is flexible and enables a variety of clause structures to be defined by the user through a general schema definition system. Our method can be extended to include other types of molecular biological database or those containing non-structural information, thus providing a uniform framework for handling the increasing volume of data available to knowledge-based systems in biomedicine.

Database Management Systems

Plasma von Willebrand Factor and ADAMTS13 Interact With APOE-ε4 in Predicting Longitudinal Brain Atrophy and Cognitive Decline Over a 9-Year Follow-Up.

BACKGROUND: Von Willebrand factor (VWF) and ADAMTS13 (a disintegrin and metalloproteinase with thrombospondin type 1 motif, 13) are linked to dementia risk, and limited evidence suggests apolipoprotein E (APOE)-&#x3b5;4 alters VWF release. This study assessed whether baseline VWF and ADAMTS13 levels predict neurodegeneration and cognitive decline and evaluated effect modification by APOE-&#x3b5;4 carriership. METHODS: Vanderbilt Memory and Aging Project cohort participants (n=332, 73&#xb1;7&#x2009;years, 59% male) completed serial blood draw, neuropsychological assessment, and brain magnetic resonance imaging over 6.4&#x2009;years (range 1.4-9.7&#x2009;years). Baseline plasma VWF and ADAMTS13 levels were quantified using mass spectrometry and Olink. Fully adjusted linear mixed-effects models related protein&#xd7;time and protein&#xd7;APOE-&#x3b5;4&#xd7;time interaction terms to longitudinal brain magnetic resonance imaging and neuropsychological outcomes. RESULTS: Lower baseline ADAMTS13 predicted faster declines in language (&#x3b2;=0.11, P=0.01), information processing speed (&#x3b2;=0.27, P=0.001), executive function (&#x3b2;=0.01, P=0.03), episodic memory (&#x3b2;=0.01, P=0.03), and visuospatial ability (&#x3b2;=0.11, P=0.001) and faster increases in global (&#x3b2;=-0.29, P=0.01) and frontal (&#x3b2;=-0.17, P=0.01) white matter hyperintensity volumes. Associations between ADAMTS13 and faster rates of cognitive decline and white matter injury were driven by APOE-&#x3b5;4 carriers. Models relating VWF to longitudinal outcomes were null. APOE-&#x3b5;4 interacted with VWF on longitudinal gray matter volumetric outcomes, such that faster rates of global gray matter atrophy were observed with higher baseline VWF levels among APOE-&#x3b5;4 noncarriers only (&#x3b2;=-1530.5, P<0.001). CONCLUSIONS: ADAMTS13 shows promise as a potential plasma biomarker for brain aging outcomes, but additional research is warranted to understand the performance of VWF in the presence versus absence of an APOE-&#x3b5;4 allele.

Humans

Post-translational modification of proteins in the human testis development pathway.

BACKGROUND: The foetal testes produce the androgens necessary to masculinise the developing embryo and support the maturation of germ cells, that will eventually develop into sperm, thus ensuring future reproductive capacity. The testes develop from the bi-potential gonads in a highly orchestrated process resulting in the differentiation of a complex tissue with multiple cellular lineages. While recent transcriptomic and chromatin-based analyses of human foetal testes have provided an unprecedented level of insight into signalling pathways activated during this process, proteomic studies of the human foetal gonads remain limited. Proteins are active molecules and post-translational modification (PTM) of proteins influences protein activity, stability and localisation. Studies have shown that PTMs regulate critical proteins in testis development, and their disruptions are implicated in congenital disorders including differences of sex development (DSD), in which sex development is atypical. Despite this, the role and regulation of protein PTM during human testis development remains poorly understood due to limited access to human foetal gonadal tissue, a paucity of large-scale proteomics studies, and a lack of robust of human gonad in vitro models. OBJECTIVE AND RATIONALE: This review aims to provide a comprehensive analysis of validated PTMs affecting proteins critical for testicular development. We discuss PTMs with evidence for a role in normal testis development, and highlight those disrupted in DSD. We review emerging techniques, including proteomic technologies and organ modelling systems that may advance our understanding of PTMs in foetal testis development. We discuss challenges that have restricted the application of these technologies and how overcoming these will significantly improve our understanding of testis development and disease, diagnostics and patient outcomes. SEARCH METHODS: We searched PubMed and the University of Melbourne library for peer-reviewed English-language studies using keywords such as phosphorylation, SUMOylation, acetylation, ubiquitination alongside each protein of interest. PTM sites in proteins involved in testis development were identified using the PhosphoSitePlus database focusing those confirmed in in vitro or animal model studies. ClinVar and the Human Gene Mutation Database were used to identify patient variants that may disrupt PTM sites. OUTCOMES: Our review finds that proteins required for human foetal testis development are subject to extensive PTM. Several PTM sites and PTM-mediated pathways [e.g. MAPK (mitogen-activated protein kinase) pathway] are disrupted in patients with DSD or related conditions. While recent advances in proteomics technologies hold considerable promise, their application to human foetal gonads has been constrained by technical, ethical, and logistical challenges. Encouragingly, emerging high-sensitivity and low-input technologies, alongside stem cell-based approaches, offer viable pathways to overcoming these barriers. WIDER IMPLICATIONS: The relationship between gene regulation, protein expression, and cellular outcome is inherently non-linear, shaped by additional regulatory layers-most notably PTMs. The contribution of PTMs to human testis development in both typical and atypical contexts is a major knowledge gap. Addressing this gap has broad clinical and biological relevance: it may help improve genetic diagnosis or shed light on how proteins or pathways critical for testis development respond to environmental signals-an increasingly pressing question as declining global fertility rates bring testicular function under greater scrutiny. REGISTRATION NUMBER: N/A.

Humans

The expert system approach to predicting protein structure.

Prediction of protein structure is an open-ended problem. Since an approach from first principles cannot be taken in reasonable computer time, short-cuts using further data are necessary. Such data include information about the specific protein in question, and information in databases which are about proteins in general. Is it possible to write a general, flexible 'superalgorithm' which would suit most circumstances? If so, it would seem likely to overcome one of the most understated but nonetheless greatest difficulties associated with molecular modelling and computer-aided drug design--reproducibility. To this end, a 'polymorphic programming environment' has been developed which represents both an expert system and a high-level language for theoretical chemists and molecular biologists. This language is GLOBAL (Ball et al., 1990). In a series of earlier studies, and more recently by means of GLOBAL itself, the nature of reproducibility and its rather surprising limits have been explored, and in general the current status and future potential of protein modelling have been examined.

Expert Systems

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348&#xa0;handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Grammatical model of the regulation of gene expression.

Based on a formal proof that justifies the search for generative grammars in the study of gene regulation, a linguistic formalization of an exhaustive data base of Escherichia coli sigma 70 promoters and their regulatory binding sites has been initiated. The grammar presented here generates all the arrays of the collection plus those that are predicted as consistent with the principles of regulation of sigma 70 promoters. "Systems of regulation," sets of regulatory sites that collaborate in a mechanism of regulation, are represented by means of syntactic categories. A small set of phrase structure rules restricted by an X-bar principle and by a hierarchical, c-command relation generates a representation of arrays of sites of regulation where the selection of the protein(s) identifying the system(s) of regulation occurs. Based on the features of the proteins, optional duplicated proximal and remote sites are generated by means of transformational rules. Consistency with the data, the predictions that the grammar generates, and important similarities and differences with some aspects of the generative theory of natural language are discussed.

Gene Expression Regulation, Bacterial

Out-of-the-box bioinformatics capabilities of large language models (LLMs).

Large Language Models (LLMs), AI agents and co-scientists promise to accelerate scientific discovery across fields ranging from chemistry to biology. Bioinformatics- the analysis of DNA, RNA and protein sequences plays a crucial role in biological research and is especially amenable to AI-driven automation given its computational nature. Here, we assess the bioinformatics capabilities of three popular general-purpose LLMs on a set of tasks covering basic analytical questions that include code writing and multi-step reasoning in the domain. Utilizing questions from Rosalind, a bioinformatics educational platform, we compare the performance of the LLMs vs. humans on 104 questions undertaken by 110 to 68,760 individuals globally. GPT-3.5 provided correct answers for 59/104 (58%) questions, while Llama-3-70B and GPT-4o answered 49/104 (47%) correctly. GPT-3.5 was the best performing in most categories, followed by Llama-3-70B and then GPT-4o. 71% of the questions were correctly answered by at least one LLM. The best performing categories included DNA analysis, while the worst performing were sequence alignment/comparative genomics and genome assembly. Overall, LLMs performance mirrored that of humans with lower performance in tasks in which humans had low performance and vice versa. However, LLMs also failed in some instances where most humans were correct and, in a few cases, LLMs excelled where most humans failed. To the best of our knowledge, this presents the first assessment of general purpose LLMs on basic bioinformatics tasks in distinct areas relative to the performance of hundreds to thousands of humans. LLMs provide correct answers to several questions that require use of biological knowledge, reasoning, statistical analysis and computer code.

Journal Article

Turn prediction in proteins using a pattern-matching approach.

We extend the use of amino acid sequence patterns [Cohen, F.E., Abarbanel, R. M., Kuntz, I. D., & Fletterick, R. J. (1983) Biochemistry 22, 4894-4904] to the identification of turns in globular proteins. The approach uses a conservative strategy, combined with a hierarchical search (strongest patterns first) and length-dependent masking, to achieve high accuracy (95%) on a test set of proteins of known structure. Applying the same procedure to homologous families gives a 90% success rate. Straightforward changes are suggested to improve the predictive power. The computer program, written in Lisp, provides a general pattern-recognition language well suited for a number of investigations of protein and nucleic acid sequences.

Amino Acid Sequence

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery

CLASPP: A unified model for predicting post-translational modifications.

Post-Translational Modifications (PTMs) are a fundamental mechanism for regulating cellular pathways and increasing the functional diversity of the proteome. Accurately predicting the PTM types that are likely to occur at a given site in the primary sequence is a key challenge in functional proteomics. Existing PTM prediction models predominantly focus on either single PTM types or employ ensemble methods that combine multiple models to predict different PTM types. This fragmentation is largely driven by the vast imbalance in data availability across PTM types, making it difficult to predict multiple PTM types with a single model. To address this limitation, we present the Contrastively Learned Attention-based Stratified PTM Predictor (CLASPP), a unified PTM prediction model. CLASPP addresses imbalance challenges by leveraging unsupervised clustering-based undersampling and a novel contrastive learning framework tailored to PTM data. Additionally, our hierarchical data organization and curation are shown to improve CLASPP's performance by balancing the representation of individual PTM types and provides a standardized dataset to train and validate future model designs. Drawing inspiration from advancements in image and natural language processing, the CLASPP model employs a multi-stage training strategy and a high-quality, curated training dataset to improve PTM prediction performance. To uncover what is learned during the contrastive learning stage, the CLASPP model is shown to distinguish known protein kinase substrate specificity profiles as a form of explainability. Finally, we evaluate the application of CLASPP in predicting PTMs in different model organisms and experimentally validated ubiquitination sites in the understudied DCLK3 kinase. Overall, CLASPP represents a unified model for PTM prediction that addresses key bottlenecks in data imbalance and offers new strategies for biological data curation, thereby improving PTM-type prediction performance across diverse organisms.

Protein Processing, Post-Translational

Biallelic variants in ZNF142 lead to a syndromic neurodevelopmental disorder.

Biallelic variants of the gene encoding for the zinc-finger protein 142 (ZNF142) have recently been associated with intellectual disability (ID), speech impairment, seizures, and movement disorders in nine individuals from five families. In this study, we obtained phenotype and genotype information of 26 further individuals from 16 families. Among the 27 different ZNF142 variants identified in the total of 35 individuals only four were missense. Missense variants may give a milder phenotype by changing the local structure of ZF motifs as suggested by protein modeling; but this correlation should be validated in larger cohorts and pathogenicity of the missense variants should be investigated with functional studies. Clinical features of the 35 individuals suggest that biallelic ZNF142 variants lead to a syndromic neurodevelopmental disorder with mild to moderate ID, varying degrees of delay in language and gross motor development, early onset seizures, hypotonia, behavioral features, movement disorders, and facial dysmorphism. The differences in symptom frequencies observed in the unpublished individuals compared to those of published, and recognition of previously underemphasized facial features are likely to be due to the small sizes of the previous cohorts, which underlines the importance of larger cohorts for the phenotype descriptions of rare genetic disorders.

Humans

Computer-assisted evaluation of polydisperse two-dimensional gel patterns of polysaccharide-protein conjugate preparations with regard to size and net charge.

Native Hemophilus influenzae polysaccharide-protein conjugate particles were analyzed by a two-dimensional agarose electrophoresis procedure. In view of their preparation by random chemical crosslinking, the conjugates necessarily exhibit a polydisperse two-dimensional gel pattern which varies depending on the conditions of the particular preparation. The polydisperse patterns were interpreted with regard to the size and surface net charge density of the conjugate on the basis of the extended Ogston model. Data processing was performed by a new program, designated ZWEIDI.DO, written in the language of M-LAB (modeling laboratory). The program computes particle and gel fiber specific parameters from the positions of standards and unknown(s) on the two-dimensional gel using a simultaneous linear least-square curve fitting routine. Based on these calculations, the program serves to compute a nomogram of iso-size and iso-free-mobility profiles. Superimposing these profiles on the gel patterns, the size and free mobility range of the polydisperse conjugate mixtures is obtained. Potentially, the procedure could serve as a tool for quality control in the production of conjugates as vaccines and for the physical characterization of polydisperse subcellular particles and vesicles.

Bacterial Proteins

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Large Language Models

Agentomics: an agentic system that autonomously develops novel state-of-the-art solutions for biomedical machine learning tasks.

MOTIVATION: Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack flexibility, while large language models (LLMs) struggle to consistently deliver reproducible machine learning codebases, and existing LLM Agent-powered solutions lag behind human-engineered ML models. RESULTS: Here, we introduce Agentomics, an autonomous LLM-powered agentic system for end-to-end ML experimentation. Given a biomedical dataset, Agentomics implements various ML modeling strategies, and produces a ready-to-use ML model. Agentomics introduces strict validation checkpoints for standard ML development steps, allowing gradual development on top of working code with defined interfaces and validated artifacts. Further, it offers native support for biomedical foundation models that can be leveraged during experimentation. The generic nature of Agentomics allows the user to create ML solutions for a large variety of datasets and use various LLMs. We evaluate Agentomics across 20 datasets from the domains of Protein Engineering, Drug Discovery, and Regulatory Genomics. When benchmarked against other agentic systems, Agentomics outperformed them in all tested domains. When benchmarked against human expert solutions, Agentomics generated novel state-of-the-art models for 11/20 established benchmark datasets. AVAILABILITY AND IMPLEMENTATION: Agentomics is implemented in Python. Source code and documentation are freely available at: https://github.com/BioGeMT/Agentomics-ML.

Machine Learning

Experimental and clinical applications of molecular cell biology in nutrition and metabolism.

Rapid advances in molecular biology have yielded important new techniques for understanding the cellular mechanisms of normal homeostasis and disease. In particular, molecular laboratory methodologies have become important investigative tools for nutritional studies. Detection techniques for specific DNAs, RNAs, and proteins allow direct examination of cellular regulation of protein expression in health and illness. Construction of transgeneic models by recent techniques of inserting foreign genes into experimental animals has provided novel models for studies of cellular metabolism. In addition, molecular biology has had impact on clinical nutrition and therapy. Molecular techniques not only allow for early diagnosis of many inborn genetic errors of metabolism, recombinant technology has also provided for large-scale production of proteins and hormones of potential therapeutic value. The possibility for direct gene therapies is also nearing reality. Hence, understanding the language of molecular biology and the recent developments in this field is not only of research interest, but is also of clinical relevance.

Animals

Design of highly functional genome editors by modelling CRISPR-Cas sequences.

Gene editing has the potential to solve fundamental challenges in agriculture, biotechnology and human health. CRISPR-based gene editors derived from microorganisms, although powerful, often show notable functional tradeoffs when ported into non-native environments, such as human cells1. Artificial-intelligence-enabled design provides a powerful alternative with the potential to bypass evolutionary constraints and generate editors with optimal properties. Here, using large language models2 trained on biological diversity at scale, we demonstrate successful precision editing of the human genome with a programmable gene editor designed with artificial intelligence. To achieve this goal, we curated a dataset of more than 1&#x2009;million CRISPR operons through systematic mining of 26 terabases of assembled genomes and metagenomes. We demonstrate the capacity of our models by generating 4.8&#xd7; the number of protein clusters across CRISPR-Cas families found in nature and tailoring single-guide RNA sequences for Cas9-like effector proteins. Several of the generated gene editors show comparable or improved activity and specificity relative to SpCas9, the prototypical gene editing effector, while being 400 mutations away in sequence. Finally, we demonstrate that an artificial-intelligence-generated gene editor, denoted as OpenCRISPR-1, exhibits compatibility with base editing. We release OpenCRISPR-1 to facilitate broad, ethical use across research and commercial applications.

CRISPR-Cas Systems

Harnessing the Power of Large Language Models for Drug Discovery: A Systematic Review of Current Applications and Future Directions.

INTRODUCTION: The demand for inventive approaches to drug discovery has increased due to the rising costs, time, and failure rates in pharmaceutical research. Large Language Models (LLMs), with their sophisticated natural language processing and generative capabilities, have become potent instruments that have the potential to revolutionize biomedical research. The function of LLMs in different phases of drug development is methodically examined in this article. METHODS: The PRISMA 2020 principles were adhered to in this systematic study. A thorough search for research published between 2018 and 2025 was done using PubMed, Scopus, Web of Science, and Google Scholar. The search terms "large language model," "transformer," "drug discovery," and important sub-domains (such as "de-novo design" and "ADMET") were merged, and two reviewers independently screened the results. Predetermined inclusion and exclusion criteria were used to filter studies for relevance. 98 studies out of the 1,285 records that were initially retrieved met the requirements for the final qualitative synthesis. RESULTS: 98 studies that demonstrated the use of LLMs in various drug discovery domains were found during the review. These covered molecular generation, genomics, protein-ligand modeling, ADME/T and toxicity profiling, drug-target interaction and DTI prediction, and biomedical text mining. 42 different LLM-based tools were mapped, including BioBERT, SciSpacy, Drug- LLM, DNA-BERT, GPT-4, and ChatGPT. Predictive accuracy, hypothesis creation, target prioritization, and multi-modal data integration all showed notable gains with these techniques. DISCUSSION: By providing scalable, precise, and effective solutions for data-driven drug discovery, LLMs are revolutionizing the pharmaceutical industry. They allow for the creation of hypotheses and individualized insights across multi-modal biological data, and they perform better than conventional approaches in a number of subdomains. Improvements in performance were task-dependent; the most consistent gains occurred for biomedical text mining, disease-genedrug relationship mapping and drug-target interaction prediction tasks. Yet most evidence for clinical applications is still derived from retrospective studies and benchmark datasets, suggesting a higher need for prospective validation. CONCLUSION: There is revolutionary potential in incorporating LLMs into drug discovery processes. Clinical translation and regulatory uptake will depend heavily on collaborative validation, ethical deployment, and standardization as models become more multimodal and interpretable. Before normal use, extensive prospective benchmarking and head-to-head comparisons with established chemoinformatics pipelines are necessary.

De novo design