PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

I2Cnet medical image annotation service.

I2Cnet (Image Indexing by Content network) aims to provide services related to the content-based management of images in healthcare over the World-Wide Web. Each I2Cnet server maintains an autonomous repository of medical images and related information. The annotation service of I2Cnet allows specialists to interact with the contents of the repository, adding comments or illustrations to medical images of interest. I2Cnet annotations may be communicated to other users via e-mail or posted to I2Cnet for inclusion in its local repositories. This paper discusses the annotation service of I2Cnet and argues that such services pave the way towards the evolution of active digital medical image libraries.

Computer Communication Networks↗

Stents. An annotated bibliography of selected references.

This is an annotated bibliography of selected references on stents. It includes 90 annotated references and 6 others that were not annotated. It has been structured to include pertinent details about each citation, including the site at which the device was used (coronary or peripheral), the type of article in which the research was reported (clinical paper, case report, abstract, review, editorial), the number of subjects and lesions treated, the design of the study (case study, single group, group comparison, or randomized trial), acute and late results, and comments about the study and its findings.

Coronary Disease↗

DAVID: Database for Annotation, Visualization, and Integrated Discovery.

BACKGROUND: Functional annotation of differentially expressed genes is a necessary and critical step in the analysis of microarray data. The distributed nature of biological knowledge frequently requires researchers to navigate through numerous web-accessible databases gathering information one gene at a time. A more judicious approach is to provide query-based access to an integrated database that disseminates biologically rich information across large datasets and displays graphic summaries of functional information. RESULTS: Database for Annotation, Visualization, and Integrated Discovery (DAVID; http://www.david.niaid.nih.gov) addresses this need via four web-based analysis modules: 1) Annotation Tool - rapidly appends descriptive data from several public databases to lists of genes; 2) GoCharts - assigns genes to Gene Ontology functional categories based on user selected classifications and term specificity level; 3) KeggCharts - assigns genes to KEGG metabolic processes and enables users to view genes in the context of biochemical pathway maps; and 4) DomainCharts - groups genes according to PFAM conserved protein domains. CONCLUSIONS: Analysis results and graphical displays remain dynamically linked to primary data and external data repositories, thereby furnishing in-depth as well as broad-based data coverage. The functionality provided by DAVID accelerates the analysis of genome-scale datasets by facilitating the transition from data collection to biological meaning.

Computational Biology↗

A DNA Motif Lexicon: cataloguing and annotating sequences.

The rapid proliferation of genomic DNA sequences has created a significant need for software that can both focus on relatively small areas (such as within genes or promoters) and provide wide-zoom views of patterns across entire genomes. We present our DNA Motif Lexicon that enables users to perform genome-wide searches for motifs of interest and create customizable results pages, where results differ in the degree and extent of annotation. Searching for a particular motif is akin to a word search in a natural language; our motif lexicon speaks to this new time when we will increasingly rely upon DNA dictionaries that offer rich types of annotation. Indeed, the concept of "lexomics", introduced in this paper may be appropriate to the types of meta-analyses appropriate to the deciphering of regulatory information. Currently supporting five genomes, our web-based lexicon allows users to look up motifs of interest and build user-defined result pages to include the following: (1) all base pair locations where a motif is found with links to further search the "neighborhoods" near each of these locations; whether each location of the motif is genic (within) a gene, intergenic, or a bridging sequence (overlapping a gene boundary) (2) NCBI hot-links to nearest upstream and downstream genes for each location (3) statistical information about the query (4) whether the motif is a certain type of repeat (5) links for the reverse, complement and reverse-complement of the motif of interest and (6) hot-links to PubMed abstracts which mention the motif of interest. A software framework facilitates the continual development of new annotation modules. The tool is located at: http://genomics.wheatoncollege.edu/cgi-bin/lexicon.exe.

Animals↗

Processing sequence annotation data using the Lua programming language.

The data processing language in a graphical software tool that manages sequence annotation data from genome databases should provide flexible functions for the tasks in molecular biology research. Among currently available languages we adopted the Lua programming language. It fulfills our requirements to perform computational tasks for sequence map layouts, i.e. the handling of data containers, symbolic reference to data, and a simple programming syntax. Upon importing a foreign file, the original data are first decomposed in the Lua language while maintaining the original data schema. The converted data are parsed by the Lua interpreter and the contents are stored in our data warehouse. Then, portions of annotations are selected and arranged into our catalog format to be depicted on the sequence map. Our sequence visualization program was successfully implemented, embedding the Lua language for processing of annotation data and layout script. The program is available at http://staff.aist.go.jp/yutaka.ueno/guppy/.

Computational Biology↗

Annotated proteome of a human T-cell lymphoma.

As the reliable identification of proteins by tandem mass spectrometry becomes increasingly common, the full characterization of large data sets of proteins remains a difficult challenge. Our goal was to survey the proteome of a human T-cell lymphoma-derived cell line in a single set of experiments and present an automated method for the annotation of lists of proteins. A downstream application of these data includes the identification of novel pathogenetic and candidate diagnostic markers of T-cell lymphoma. Total protein isolated from cytoplasmic, membrane, and nuclear fractions of the SUDHL-1 T-cell lymphoma cell line was resolved by SDS-PAGE, and the entire gel lanes digested and analyzed by tandem mass spectrometry. Acquired data files were searched against the UniProt protein database using the SEQUEST algorithm. Search results for each subcellular fraction were analyzed using INTERACT and ProteinProphet. All protein identifications with an error rate of less than 10% were directly exported into excel and analyzed using GOMiner (NIH/NCI). The Gene ontology molecular function and cell location data were summarized for the identified proteins and results exported as user-interactive directed acyclic graphs. A total of 1105 unique proteins were identified and fully annotated, including numerous proteins that had not been previously characterized in lymphoma, in functional categories such as cell adhesion, migration, signaling, and stress response. This study demonstrates the utility of currently available bioinformatics tools for the robust identification and annotation of large numbers of proteins in a batchwise fashion.

Algorithms↗

Integration of data for gene annotation using the BioMediator system.

Gene annotation requires integration of data from multiple sources in order to functionally classify genes. We are using BioMediator, a general purpose data-integration solution, to develop a gene annotation system to automate the process of collecting data from disparate genomic databases. Integration of annotation data from multiple sources into a single format will facilitate use of analytic tools for the proper functional classification of genes.

Base Sequence↗

Annotated genome assemblies of two temperate North American dung beetles, Canthon chalcites and Phanaeus vindex.

Dung beetles serve as cultivators of their natural habitats, improving soil health and functions in both natural and anthropogenic environments. Despite their ecological importance, whole genome sequences for Scarabaeinae are limited. Here, we present the draft annotated genome assemblies for 2 temperate species of North American dung beetles collected from eastern Tennessee: Canthon chalcites and Phanaeus vindex. Both genome assemblies were generated from PacBio long reads and have high completeness, with BUSCO scores of 98.1% and 98.6% for C. chalcites and P. vindex, respectively. For C. chalcites, the BRAKER3 pipeline predicted 12,799 genes, and the gene set was 93.7% complete. For P. vindex, the BRAKER3 predicted 12,252 genes, and the gene set was 94.9% complete. From the annotated gene sets, orthologous protein sequence analyses among C. chalcites, P. vindex, the dung beetle species Onthophagus taurus, and the more evolutionarily distant beetle Tribolium castaneum indicated that there are 260 unique protein clusters for C. chalcites and 210 unique protein clusters for P. vindex. These 2 draft genomes provide valuable data for comparative genomics, evolution, and phylogenic studies for dung beetle species.

Animals↗

Integrating structure and experimental data annotations with computational modeling framework for predicting micro-nanoplastics toxicities.

The wide use of plastic materials leads to increased emissions of micro-nanoplastics (MNPs) into the environment, raising significant concerns about their impact on human health. Traditional experimental approaches for assessing MNPs toxicity are costly, time-consuming, and there are no experimental protocols that are universally acceptable. Computational modeling using machine learning (ML) approaches provides an efficient alternative to MNP toxicity assessment. However, most modeling studies of MNPs are limited due to the lack of high-quality data and there are few previous modeling studies considering complex structures of MNPs for model training. To address this challenge, we constructed three MNP datasets with popular toxicity endpoints from various resources and used nanostructure annotation techniques to create virtual MNPs (vMNPs) for all MNP structures. The MNP structures were digitalized from annotated vMNPs, and geometrical descriptors were calculated using the Delaunay Tessellation approach. Moreover, important experimental information, such as concentrations and cell lines, were transformed into extra training variables. Partial least squares regression (PLSR) models were built using both experimental and geometrical descriptors and validated through a leave-one-out cross validation procedure. The resulting models showed reasonable performance in predicting toxicity potentials of MNPs for the three endpoints in the present datasets. Moreover, an additional library of vMNPs with their predicted properties and bioactivities was constructed, directing further research of new MNPs. This study provides three novel ML models for MNPs by integrating geometrical and experimental descriptors, which have the potential to assess new MNPs for their toxicity. The modeling strategy developed in this study can be easily expanded to model other MNP toxicity endpoints and create promising new models for MNP toxicity assessments.

Data annotation↗

Mapping the Immune cell-specific gene regulatory network in bipolar disorder: A framework from scTWMR to exploratory drug-target annotation.

BACKGROUND: Although the involvement of the immune system in the genetic susceptibility of bipolar disorder (BD) is widely acknowledged, the causal relationship between gene expression in specific immune cell subtypes and BD requires systematic elucidation. METHODS: We implemented an analytical framework integrating single-cell transcriptome-wide Mendelian randomization (scTWMR) with colocalization analysis. This approach utilized cis-expression quantitative trait loci (cis-eQTLs) derived from 14 distinct immune cell types as instrumental variables to interrogate BD genome-wide association study (GWAS) summary statistics (comprising 41,917 cases and 371,549 controls). Subsequent investigations encompassed functional enrichment analysis, protein-protein interaction (PPI) network construction, phenome-wide association study (PheWAS), and performed an exploratory drug-target annotation. RESULTS: Our analysis identified 33 gene-immune cell associations. Colocalization analysis provided robust evidence (PPH4 > 90%) for shared causal variants implicating the MAD1L1, APOM, and NFKBIL1 loci. Significantly enriched biological pathways included cell cycle regulation, circadian rhythm entrainment, and neuroinflammation. The PPI network revealed a core regulatory module centered on histone-encoding and immune-related genes. Exploratory drug-target annotation nominated compounds for further investigation for compounds targeting APOM, TMEM258, and NFKBIL1. CONCLUSION: This study systematically delineates a genetically supported regulatory network of immune cell-specific gene expression in BD, predominantly implicating CD8⁺ effector T cells, plasma cells, and B cells. The findings corroborate established pathological pathways while uncovering novel cell type-specific therapeutic targets, thereby providing a genetic framework for prioritizing candidate targets for future investigation.

Bipolar disorder↗

Conditional Diffusion Model-Based Method for Annotation of Antibiotic Resistance Gene Properties.

The crisis of bacterial antibiotic resistance, which has led to a decline in the effectiveness of antibiotics originally used to combat bacterial infections, has emerged as an urgent challenge for public health. Antibiotic resistance genes (ARGs) are one of the key reasons for bacteria to develop resistance to antibiotics. Therefore, accurately identifying and annotating the critical properties of ARGs is of great importance for addressing the antibiotic resistance emergency. Although existing deep learning models demonstrate remarkable effectiveness in extracting local features from sequence data, they still face limitations in the capacity to further gain the enriched latent representations within the data. To address the critical challenge of extracting higher-quality representations from ARGs sequence data, we propose a novel ARGs properties annotation method based on the conditional diffusion model which is used to learn latent representations through domain-specific knowledge injection. Specifically, during the conditional information integration phase, we systematically incorporate ARGs' domain knowledge to guide the diffusion process in generating high-quality latent representations. To overcome information redundancy caused by direct concatenation of conditional information and intermediate features, we design a cross-attention mechanism that enables feature fusion between heterogeneous information sources, thereby enhancing further the quality of obtained representations. Experimental results on widely used data sets demonstrate the framework's effectiveness in achieving superior prediction performance compared to existing methods.

Anti-Bacterial Agents↗

PubMind: literature-based genetic variant extraction and functional annotation using large language models.

Biomedical literature contains extensive functional knowledge on genetic variants, but much remains inaccessible in unstructured text. Existing resources such as ClinVar and HGMD remain limited by coverage, submission bias, update frequency, and sparse annotation. We develop PubMind, an artificial intelligence (AI) framework that uses large language models (LLMs) to triage and extract variant-function-disease associations and supporting evidence from biomedical text. PubMind captures single-nucleotide, copy-number, structural, and gene-fusion variants, and normalizes records to genomic and transcriptomic coordinates. Benchmarking shows >90% accuracy for variant recognition and 99% precision for disease extraction. Applied to >41 million PubMed abstracts and >5 million full-text articles, PubMind generates PubMind-DB, a database of ~1.3 million unique variants with contextual annotations, accessible via web interface and API. Only ~10% of PubMind variants overlap with ClinVar, and >80% of them show concordant pathogenicity labels. PubMind transforms unstructured biomedical text into structured genomic knowledge, advancing variant interpretation for precision medicine.

Large Language Models↗

Chromosome-level genome assembly and annotation of Pterygoplichthys pardalis.

Suckermouth catfishes, with their evolved powerful features, have become notorious invasive species, causing significant damage to aquatic ecosystems. However, the lack of high-quality genomes severely restricts research on this group within the field. In this study, we de novo assembled the chromosome-level genome assembly of Pterygoplichthys pardalis using multiple platforms of sequencing data, including Illumina short reads, Nanopore long reads, and Hi-C sequencing reads, resulting in a 1.51 Gb genome assembly. Multiple evaluations, including read mapping ratio (98.52%), transcript mapping ratio (99.61%), conserved BUSCO gene set (98.8%), and N50 score (49.47 Mb), indicated the high continuity and accuracy of the genome assembly we generated. Genome annotation found that 0.97 Gb of genome sequences are repetitive sequences, accounting for 64.47% of the genome assembly. Further, 23,859 protein-coding genes were successfully predicted, 92.92% of which could be annotated in functional databases. This high-quality genome assembly of P. pardalis provides a valuable resource for understanding the genetic underpinnings of P. pardalis's invasive success and offers critical data for future fisheries research and management.

Animals↗

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144 hours). RNA from collected samples was isolated and sequenced producing 60.88 M 100 bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome↗

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76 Gb in length with a contig N50 of 33.53 Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals↗

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

Missense variants pathogenicity annotation from homologous proteins.

MOTIVATION: High-throughput DNA sequencing has revealed millions of single nucleotide variants (SNVs) in the human genome, with a small fraction linked to disease. The effect of missense variants, which alter the protein sequence, is particularly challenging to interpret due to the scarcity of clinical annotations and experimental information. While using conservation and structural information, current prediction tools still struggle to predict variant pathogenicity. In this study, we explored the pathogenicity of homologous missense variants-variants in equivalent positions across homologous proteins-focusing on proteins involved in autosomal dominant diseases. RESULTS: Our analysis of 2976 pathogenic and 17&#xa0;555 non-pathogenic homologous variants demonstrated that pathogenicity can be extrapolated with 95% accuracy within a family, or up to 98% for closer homologs. Remarkably, the evaluation of 27 commonly used mutation predictor methods revealed that they were not fully capturing this biological feature. To facilitate the exploration of homologous variants, we created HomolVar, a web server that computationally predicts the pathogenesis of missense variants using annotations from homologous variants, freely available at https://rarevariants.org/HomolVar. Overall, these findings and the accompanying tool offer a robust method for predicting the pathogenicity of unannotated variants, enhancing genotype-phenotype correlations, and contributing to diagnosing rare genetic disorders. AVAILABILITY AND IMPLEMENTATION: HomolVar is freely available at https://rarevariants.org/HomolVar.

Mutation, Missense↗