PubMed HealthSearch

SEARCH · PubMed Health

Results for “annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

argNorm: normalization of antibiotic resistance gene annotations to the Antibiotic Resistance Ontology (ARO).

SUMMARY: Currently available and frequently used tools for annotating antibiotic resistance genes (ARGs) in genomes and metagenomes provide results using inconsistent nomenclature. This makes the comparison of different ARG annotation outputs challenging. The comparability of ARG annotation outputs can be improved by mapping gene names and their categories to a common controlled vocabulary such as the Antibiotic Resistance Ontology (ARO). We developed argNorm, a command line tool and Python library, to normalize all detected genes across six ARG annotation tools (eight databases) to the ARO. argNorm also adds information to the outputs using the same ARG categorization so that they are comparable across tools. AVAILABILITY AND IMPLEMENTATION: argNorm is available as an open-source tool at: https://github.com/BigDataBiology/argNorm. It can also be downloaded as a PyPI package and is available on Bioconda and as an nf-core module.

Molecular Sequence Annotation

VDJ-Insights: simplifying the annotation of genomic immunoglobulin and T cell receptor regions.

MOTIVATION: Accurate annotation of germline immunoglobulin (IG) and T cell receptor (TCR) loci is critical for understanding adaptive immunity. RESULTS: VDJ-Insights provides a user-friendly software package for characterizing these complex immune regions. In addition, it assesses gene segment functionality, identifies recombination signal sequences, and annotates complementarity-determining regions 1 and 2. VDJ-Insights achieved over 99% concordance with curated annotations from multiple species, outperforming existing annotation tools. When applied to 95 haplotypes from the Human Pangenome Reference Consortium, VDJ-Insights identified 652 and 275 novel IG and TCR alleles, respectively, highlighting its scalability for large immunogenetic studies. AVAILABILITY AND IMPLEMENTATION: Datasets and software package are available in the VDJ-insights repository, https://github.com/BPRC-Bioinfo and https://doi.org/10.5281/zenodo.17588835. Additional intermediate datasets used and analyzed during the current study are available from the corresponding authors upon reasonable request.

Software

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software

Integrative evidence-knowledge marker selection enhances LLM-based cell type annotation in single-cell RNA-seq analysis.

BACKGROUND: Cell type annotation is essential for gaining biological insight from single-cell RNA sequencing data, yet manual labeling remains time-consuming and difficult to reproduce. Various computational approaches have been developed to automate this process, and recent studies suggest that large language models can infer cell types with promising accuracy in single-cell analysis. However, most workflows still rely on cluster-specific markers derived from gene expression alone or manual curation. As a result, marker selection can be sensitive to statistical criteria and dataset-dependent bias, which may lead to the selection of less informative genes or missing important markers, while providing limited biological context. RESULTS: To address this limitation, we introduce CELLIA, an LLM-based workflow for automated and robust cell type annotation. CELLIA employs an integrative evidence-knowledge marker selection strategy that combines statistical differential expression criteria with curated tissue-specific marker resources to identify informative marker genes. In benchmarking analyses of 102 cell types, this approach improved agreement with manual annotations. In addition, CELLIA achieved higher agreement in subtype-level analyses of closely related immune populations and was further evaluated in a non-immune stromal subtype setting, covering 25 cell types in total. CONCLUSION: By integrating evidence-knowledge from gene expression with curated biological prior knowledge, CELLIA provides a more stable marker selection and improves the reliability of LLM-cell type annotation.

Cell type annotation

Large language models improve annotation of prokaryotic viral proteins.

Viral genomes are poorly annotated in metagenomic samples, representing an obstacle to understanding viral diversity and function. Current annotation approaches rely on alignment-based sequence homology methods, which are limited by the paucity of characterized viral proteins and divergence among viral sequences. Here we show that protein language models can capture prokaryotic viral protein function, enabling new portions of viral sequence space to be assigned biologically meaningful labels. When applied to global ocean virome data, our classifier expanded the annotated fraction of viral protein families by 29%. Among previously unannotated sequences, we highlight the identification of an integrase defining a mobile element in marine picocyanobacteria and a capsid protein that anchors globally widespread viral elements. Furthermore, improved high-level functional annotation provides a means to characterize similarities in genomic organization among diverse viral sequences. Protein language models thus enhance remote homology detection of viral proteins, serving as a useful complement to existing approaches.

Viral Proteins

Annotated genome of the Atlantic dog whelk, Nucella lapillus.

Nucella lapillus is an important player in rocky shore food chains and has been a focal organism of ecological and evolutionary studies for decades. Despite poor dispersal, they have a broad geographic range, which makes them an ideal species to examine isolation by distance and selection across environmental gradients. Here we present the fully annotated genome of N. lapillus generated with Oxford Nanopore Techonology sequencing at ∼37× coverage. The genome assembly is 2.32 Gbp and consists of 2,525 contigs, with an N50 length of 2 Mbp. Repeat annotation identified 2,491 families that cover 67.56% of the genome, which is similar to other gastropods. Despite its large size and high proportion of repeats, the genome is of high quality. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis revealed a score of 96.8%. Functional annotation of the genome produced 45,848 protein-coding genes with a 96.6% BUSCO score. Genomic resources for mollusks lag behind that of other phyla, perhaps because many of their innate characteristics complicate DNA extraction, sequencing, and assembly. This new N. lapillus genome will increase our genomic understanding of the second largest phylum (and the most diverse class within said phylum) and serve as a key resource to advance studies on the organismal biology and population genetics of this iconic species as well as the connection between genomic variation and community-level processes.

Animals

Chromosome-level assembly and annotation of the Jaguar (Panthera onca) genome.

OBJECTIVES: The Jaguar (Panthera onca) is a large cat species native to the Americas. Despite being successful predators, jaguar populations have declined due to habitat loss. Genome resources can help in conservation efforts as well as in understanding the interesting biology of these Felids. Beside contiguity, a well annotated reference genome provides contextual information for variants that will benefit the design of appropriate conservation programs. DATA DESCRIPTION: We sequenced material from two individuals using a combination of ONT reads and Illumina PE. The resulting nuclear genome assembly has a larger contig N50 (48.04 Mb) compared with the existing annotated chromosome-level assembly published by the DNA Zoo project. Using public Hi-C data, we obtained an improved chromosome-level assembly of the Jaguar genome (mPanOnc3.5) with larger contigs, 99.85% of the sequence assigned to chromosomes and 25,267 protein coding genes annotated. Overall, this improved assembly provides a better reference to study this threatened species.

Animals

Functional mapping and annotation of genetic associations with FUMA.

A main challenge in genome-wide association studies (GWAS) is to pinpoint possible causal variants. Results from GWAS typically do not directly translate into causal variants because the majority of hits are in non-coding or intergenic regions, and the presence of linkage disequilibrium leads to effects being statistically spread out across multiple variants. Post-GWAS annotation facilitates the selection of most likely causal variant(s). Multiple resources are available for post-GWAS annotation, yet these can be time consuming and do not provide integrated visual aids for data interpretation. We, therefore, develop FUMA: an integrative web-based platform using information from multiple biological resources to facilitate functional annotation of GWAS results, gene prioritization and interactive visualization. FUMA accommodates positional, expression quantitative trait loci (eQTL) and chromatin interaction mappings, and provides gene-based, pathway and tissue enrichment results. FUMA results directly aid in generating hypotheses that are testable in functional experiments aimed at proving causal relations.

Chromatin

Genome-wide annotation of human multi-nucleotide variants reveals widespread functional differences from single nucleotide variants.

Multi-nucleotide variants (MNVs) represent a crucial yet underexplored category of genetic variation. Despite previous studies highlighting the prevalence and potential biological impact of MNVs in populations, comprehensive identification and detailed functional annotation of MNVs remain challenging. Here, we develop MNVAnno, a toolbox for rapid identification and annotation of complex MNVs, and utilize it to identify 3,984,258 MNVs from 700,134 human samples, expanding the human MNV list to 8,199,654. Our analysis reveals that MNVs can not only lead to distinct amino acid changes from their constituent single-nucleotide variants, but also significantly impact the function of non-coding regions. Furthermore, through genome-wide association studies, we identify some MNVs associated with multiple cancers, and establish the Human MNV Database to facilitate MNV research. Our study emphasizes the importance of MNV annotation, broadens the human MNV landscape, and opens avenues for exploring genetic variation in phenotypes and diseases.

Humans

Ori-Finder-Arch: An Updated Web Server for the Annotation and Visualization of Archaeal Replication Origins.

Archaea are promising chassis organisms in biotechnology, and the accurate annotation of their chromosomal replication origins (oriCs) is the key to unlocking their full potential. However, the existing Ori-Finder 2 web server suffers from low accuracy, slow speed, and limited scalability. In this study, we present Ori-Finder-Arch, an updated web server for high-performance oriC prediction in archaea. This pipeline integrates HMMER-based replication initiation protein (RIP) annotation, refined consensus motif recognition, and GC profile-based DNA unwinding element (DUE) detection. On a benchmark set of experimentally validated oriCs, Ori-Finder-Arch achieved a recall of 95.6% and a precision of 86.0%, substantially outperforming Ori-Finder 2 (62.2% and 63.6%, respectively), while running 4.75 times faster and supporting diverse assembly levels. When applied to the available archaeal assemblies, it successfully annotated 17,472 oriCs. Meanwhile, the web server provides interactive visualizations at different levels. In conclusion, Ori-Finder-Arch offers an efficient, accurate, and user-friendly platform for advanced studies of archaeal DNA replication initiation and synthetic biology applications, and is freely available at https://tubic.org/Ori-Finder-Arch/ and https://tubic.tju.edu.cn/Ori-Finder-Arch/.

Archaea

Comprehensive Transcriptome Annotation of Thousands of HIV-1 Genomes.

Alternative splicing in HIV-1 has been a central focus of decades of research, uncovering key mechanisms of viral gene regulation, immune evasion, and therapeutic response - yet, no reference resource has existed to support transcriptome-wide analysis, limiting adoption of modern computational methods. We present HIV Atlas (https://ccb.jhu.edu/HIV_Atlas), the first reference-quality annotation of HIV-1 and SIV transcriptional diversity. We manually curated transcriptomes for HIV-1HXB2 and SIVmac239 and developed Vira, an automated annotation-transfer method specifically designed to address unique challenges of viral genome biology, to generate high-quality annotations for 2,077 complete HIV-1 genomes. Using the resources presented in our work, we evaluated conservation of splice sites, revealing near-perfect preservation of major donors and acceptors. Furthermore, using several public datasets, we demonstrate how HIV Atlas enhances methodology, improves the quality and novelty of results, and opens novel avenues for research, supporting more accurate and comprehensive analyses of bulk, single-cell, and spatial RNA-seq in HIV-1 studies.

Journal Article

Annotated genome assemblies of two temperate North American dung beetles, Canthon chalcites and Phanaeus vindex.

Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.

Animals

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation

Mapping the Immune cell-specific gene regulatory network in bipolar disorder: A framework from scTWMR to exploratory drug-target annotation.

BACKGROUND: Although the involvement of the immune system in the genetic susceptibility of bipolar disorder (BD) is widely acknowledged, the causal relationship between gene expression in specific immune cell subtypes and BD requires systematic elucidation. METHODS: We implemented an analytical framework integrating single-cell transcriptome-wide Mendelian randomization (scTWMR) with colocalization analysis. This approach utilized cis-expression quantitative trait loci (cis-eQTLs) derived from 14 distinct immune cell types as instrumental variables to interrogate BD genome-wide association study (GWAS) summary statistics (comprising 41,917 cases and 371,549 controls). Subsequent investigations encompassed functional enrichment analysis, protein-protein interaction (PPI) network construction, phenome-wide association study (PheWAS), and performed an exploratory drug-target annotation. RESULTS: Our analysis identified 33 gene-immune cell associations. Colocalization analysis provided robust evidence (PPH4 > 90%) for shared causal variants implicating the MAD1L1, APOM, and NFKBIL1 loci. Significantly enriched biological pathways included cell cycle regulation, circadian rhythm entrainment, and neuroinflammation. The PPI network revealed a core regulatory module centered on histone-encoding and immune-related genes. Exploratory drug-target annotation nominated compounds for further investigation for compounds targeting APOM, TMEM258, and NFKBIL1. CONCLUSION: This study systematically delineates a genetically supported regulatory network of immune cell-specific gene expression in BD, predominantly implicating CD8⁺ effector T cells, plasma cells, and B cells. The findings corroborate established pathological pathways while uncovering novel cell type-specific therapeutic targets, thereby providing a genetic framework for prioritizing candidate targets for future investigation.

Bipolar disorder

Conditional Diffusion Model-Based Method for Annotation of Antibiotic Resistance Gene Properties.

The crisis of bacterial antibiotic resistance, which has led to a decline in the effectiveness of antibiotics originally used to combat bacterial infections, has emerged as an urgent challenge for public health. Antibiotic resistance genes (ARGs) are one of the key reasons for bacteria to develop resistance to antibiotics. Therefore, accurately identifying and annotating the critical properties of ARGs is of great importance for addressing the antibiotic resistance emergency. Although existing deep learning models demonstrate remarkable effectiveness in extracting local features from sequence data, they still face limitations in the capacity to further gain the enriched latent representations within the data. To address the critical challenge of extracting higher-quality representations from ARGs sequence data, we propose a novel ARGs properties annotation method based on the conditional diffusion model which is used to learn latent representations through domain-specific knowledge injection. Specifically, during the conditional information integration phase, we systematically incorporate ARGs' domain knowledge to guide the diffusion process in generating high-quality latent representations. To overcome information redundancy caused by direct concatenation of conditional information and intermediate features, we design a cross-attention mechanism that enables feature fusion between heterogeneous information sources, thereby enhancing further the quality of obtained representations. Experimental results on widely used data sets demonstrate the framework's effectiveness in achieving superior prediction performance compared to existing methods.

Anti-Bacterial Agents

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals