PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Genes linked by fusion events are generally of the same functional category: a systematic analysis of 30 microbial genomes.

Recent work in computational genomics has shown that a functional association between two genes can be derived from the existence of a fusion of the two as one continuous sequence in another genome. For each of 30 completely sequenced microbial genomes, we established all such fusion links among its genes and determined the distribution of links within and among 15 broad functional categories. We found that 72% of all fusion links related genes of the same functional category. A comparison of the distribution of links to simulations on the basis of a random model further confirmed the significance of intracategory fusion links. Where a gene of annotated function is linked to an unclassified gene, the fusion link suggests that the two genes belong to the same functional category. The predictions based on fusion links are shown here for Methanobacterium thermoautotrophicum, and another 661 predictions are available at http://fusion.bu.edu.

Gene Expression Regulation, Bacterial↗

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Genetic overlap between depression and C-reactive protein levels: Evidence from a cross-trait analysis.

Inflammation and depression have been consistently associated, with elevated C-reactive protein (CRP) levels observed in a significant subset of affected individuals. However, the genetic mechanisms underlying this association remain poorly understood. We integrated results from large-scale genome-wide association studies (GWAS) of depression and CRP levels in a cross-trait analysis specifically focusing on identifying horizontally pleiotropic loci. Identified variants were stratified as concordant versus discordant based on their direction of effects on the two traits and followed up using functional annotation, gene set enrichment, and colocalization analyses. We also explored causal relationships using Mendelian Randomization (MR) analysis with extensive sensitivity analyses, including adjustment for body mass index (BMI). We identified 9 novel loci. Functional analyses revealed that concordant loci were enriched in genes linked to immune and inflammatory processes, while discordant loci mostly mapped to metabolic pathways, including lipid regulation. MR provided strong evidence for body mass index driving a causal relationship between the genetic liability of depression on CRP levels. Our findings suggest that the association between depression and CRP levels is partly driven by shared genetic influences, pointing to different biological pathways depending on whether genetic effects are concordant or discordant. These results underscore the importance of considering effect direction when assessing the genetic overlap between depression and inflammatory processes. In addition, they highlight BMI as a key factor in the causal relationship between depression and systemic inflammation.

C-Reactive Protein↗

A protein-protein interaction map of the Caenorhabditis elegans 26S proteasome.

The ubiquitin-proteasome proteolytic pathway is pivotal in most biological processes. Despite a great level of information available for the eukaryotic 26S proteasome-the protease responsible for the degradation of ubiquitylated proteins-several structural and functional questions remain unanswered. To gain more insight into the assembly and function of the metazoan 26S proteasome, a two-hybrid-based protein interaction map was generated using 30 Caenorhabditis elegans proteasome subunits. The results recapitulate interactions reported for other organisms and reveal new potential interactions both within the 19S regulatory complex and between the 19S and 20S subcomplexes. Moreover, novel potential proteasome interactors were identified, including an E3 ubiquitin ligase, transcription factors, chaperone proteins and other proteins not yet functionally annotated. By providing a wealth of novel biological hypotheses, this interaction map constitutes a framework for further analysis of the ubiquitin-proteasome pathway in a multicellular organism amenable to both classical genetics and functional genomics.

Animals↗

RAREsim2: flexible simulation of rare variant genetic data using real haplotypes.

MOTIVATION: Realistic simulated data is critical for advancing methodological development and optimizing study design in genetics research. However, many genetic simulation tools are unable to replicate the distribution of rare variants or incorporate key genetic information, such as functional annotations and linkage disequilibrium. RAREsim, an accurate rare variant simulation algorithm that uses real genetic haplotypes, was developed to address these limitations. Here, we introduce RAREsim2, an update that provides both streamlined software and new functionalities for simulating individual-level differences (e.g., case-control status, technological or batch effects) and variant-level differences to represent a variety of causal models. RESULTS: We demonstrate RAREsim2's utility with three rare variant association methods (Burden, SKAT, and SKAT-O) across several simulation scenarios, including various genetic ancestries, gene sizes, strengths of association, and proportions of risk variants. Type I Error was maintained and the test with the highest power matched previously known patterns. Importantly, real genetic regions can be simulated to include known variant functions and disease associations. Ultimately, RAREsim2 offers additional flexibility and ease in simulating a multitude of realistic genetic scenarios. AVAILABILITY AND IMPLEMENTATION: The RAREsim2 Python package is publicly available on Github (https://github.com/Hendricks-Research-Team/RAREsim2), PyPI (https://pypi.org/project/raresim/), and Zenodo (https://doi.org/10.5281/zenodo.19442523). Code for the example demonstration can be found at https://github.com/JessMurphy/RAREsim2-demo.

Software↗

Genome-Wide Characterization of β-Glucosidase (TaBGLU) Genes in Bread Wheat and Their Expression Under Drought, Cold, and Combined Stress.

Glycoside hydrolase 1 (GH1) β-glucosidases were known to activate hormone conjugates and defense metabolites, yet their genomic organization and stress-response dynamics in wheat remained incompletely defined. We therefore performed an integrated characterization of TaBGLUs spanning phylogeny, gene structure and conserved motifs, subcellular localization, promoter cis-elements, Gene Ontology enrichment, protein-protein interaction networks, and targeted expression profiling. Wheat TaBGLUs partitioned into well-supported clades that shared canonical GH1 catalytic residues and a largely conserved motif scaffold. Subcellular localization predictions indicated predominant nuclear and chloroplast targeting, with a smaller cohort directed to secretory or endomembrane compartments. Promoters were enriched for light-responsive, hormone-related (ABA, JA/SA, auxin, GA) and stress-associated (MYB/WRKY, heat, low temperature) cis-elements, and functional annotations were consistent with roles in carbohydrate and cell-wall metabolism, hormone homeostasis, and defense. Network analysis revealed a densely connected TaBGLU submodule embedded within broader carbohydrate and defense interaction networks, suggesting coordinated or cooperative functions. Expression profiling under cold, drought, and combined drought and cold demonstrated broad stress inducibility, with early activation detected by 6 h, cold-responsive maxima typically at 12 h, drought-responsive peaks predominating at 24 h, and combined stress eliciting both earlier and more sustained expression maxima between 12-24 h. Representative strongly responsive genes included TaBGLU20, TaBGLU44, TaBGLU6, and TaBGLU23, which showed pronounced late induction under combined stress, TaBGLU30, which exhibited an earlier combined-stress peak, and TaBGLU12, which displayed a marked late drought-specific response. Taken together, this integrated genomic, regulatory, and expression atlas refined the wheat BGLU repertoire relative to previous gene model inventories, highlighted candidate TaBGLUs with central network positions and strong stress inducibility, and provided concrete entry points for functional validation and breeding for improved stress resilience.

Triticum↗

Annotated expressed sequence tags and cDNA microarrays for studies of brain and behavior in the honey bee.

To accelerate the molecular analysis of behavior in the honey bee (Apis mellifera), we created expressed sequence tag (EST) and cDNA microarray resources for the bee brain. Over 20,000 cDNA clones were partially sequenced from a normalized (and subsequently subtracted) library generated from adult A. mellifera brains. These sequences were processed to identify 15,311 high-quality ESTs representing 8912 putative transcripts. Putative transcripts were functionally annotated (using the Gene Ontology classification system) based on matching gene sequences in Drosophila melanogaster. The brain ESTs represent a broad range of molecular functions and biological processes, with neurobiological classifications particularly well represented. Roughly half of Drosophila genes currently implicated in synaptic transmission and/or behavior are represented in the Apis EST set. Of Apis sequences with open reading frames of at least 450 bp, 24% are highly diverged with no matches to known protein sequences. Additionally, over 100 Apis transcript sequences conserved with other organisms appear to have been lost from the Drosophila genome. DNA microarrays were fabricated with over 7000 EST cDNA clones putatively representing different transcripts. Using probe derived from single bee brain mRNA, microarrays detected gene expression for 90% of Apis cDNAs two standard deviations greater than exogenous control cDNAs. [The sequence data described in this paper have been submitted to Genbank data library under accession nos. BI502708-BI517278. The sequences are also available at http://titan.biotec.uiuc.edu/bee/honeybee_project.htm.]

Animals↗

The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.

While genome sequencing projects are advancing rapidly, EST sequencing and analysis remains a primary research tool for the identification and categorization of gene sequences in a wide variety of species and an important resource for annotation of genomic sequence. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi. shtml) are a collection of species-specific databases that use a highly refined protocol to analyze EST sequences in an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed by first clustering, then assembling EST and annotated gene sequences from GenBank for the targeted species. This process produces a set of unique, high-fidelity virtual transcripts, or Tentative Consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, to provide links between orthologous and paralogous genes and as a resource for comparative sequence analysis.

Animals↗

Annotated draft genomic sequence from a Streptococcus pneumoniae type 19F clinical isolate.

The public availability of numerous microbial genomes is enabling the analysis of bacterial biology in great detail and with an unprecedented, organism-wide and taxon-wide, broad scope. Streptococcus pneumoniae is one of the most important bacterial pathogens throughout the world. We present here sequences and functional annotations for 2.1-Mbp of pneumococcal DNA, covering more than 90% of the total estimated size of the genome. The sequenced strain is a clinical isolate resistant to macrolides and tetracycline. It carries a type 19F capsular locus, but multilocus sequence typing for several conserved genetic loci suggests that the strain sequenced belongs to a pneumococcal lineage that most often expresses a serotype 15 capsular polysaccharide. A total of 2,046 putative open reading frames (ORFs) longer than 100 amino acids were identified (average of 1,009 bp per ORF), including all described two-component systems and aminoacyl tRNA synthetases. Comparisons to other complete, or nearly complete, bacterial genomes were made and are presented in a graphical form for all the predicted proteins.

DNA, Bacterial↗

Can sequence determine function?

The functional annotation of proteins identified in genome sequencing projects is based on similarities to homologs in the databases. As a result of the possible strategies for divergent evolution, homologous enzymes frequently do not catalyze the same reaction, and we conclude that assignment of function from sequence information alone should be viewed with some skepticism.

Animals↗

Methods to map protein interactions in mammalian cells: different tools to address different questions.

In the post-genome era, functional annotation of the predicted gene-sets will be one of the most important upcoming challenges. So-called interactome analysis positions a protein in its subcellular environment by mapping its interaction partners. Such interaction maps are essential for an accurate insight into protein function since many cellular processes are organised to operate in protein complexes. These assemblies have dynamic structures and can interact with each other, two properties which are often controlled by regulated protein expression and modification. Various methods exist to unravel protein interaction circuitries, which can be roughly divided into biochemical and genetic strategies. In this review we focus on the different strategies to study protein-protein interactions in living mammalian cells. Recently developed analytical and screening methods are also addressed.

Animals↗

Annotated genome of the Atlantic dog whelk, Nucella lapillus.

Nucella lapillus is an important player in rocky shore food chains and has been a focal organism of ecological and evolutionary studies for decades. Despite poor dispersal, they have a broad geographic range, which makes them an ideal species to examine isolation by distance and selection across environmental gradients. Here we present the fully annotated genome of N. lapillus generated with Oxford Nanopore Techonology sequencing at ∼37× coverage. The genome assembly is 2.32 Gbp and consists of 2,525 contigs, with an N50 length of 2 Mbp. Repeat annotation identified 2,491 families that cover 67.56% of the genome, which is similar to other gastropods. Despite its large size and high proportion of repeats, the genome is of high quality. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis revealed a score of 96.8%. Functional annotation of the genome produced 45,848 protein-coding genes with a 96.6% BUSCO score. Genomic resources for mollusks lag behind that of other phyla, perhaps because many of their innate characteristics complicate DNA extraction, sequencing, and assembly. This new N. lapillus genome will increase our genomic understanding of the second largest phylum (and the most diverse class within said phylum) and serve as a key resource to advance studies on the organismal biology and population genetics of this iconic species as well as the connection between genomic variation and community-level processes.

Animals↗

Extracting knowledge from dynamics in gene expression.

Most investigations of coordinated gene expression have focused on identifying correlated expression patterns between genes by examining their normalized static expression levels. In this study, we focus on the dynamics of gene expression by seeking to identify correlated patterns of changes in genetic expression level. In doing so, we build upon methods developed in clinical informatics to detect temporal trends of laboratory and other clinical data. We construct relevance networks from Saccharomyces cerevisiae gene-expression dynamics data and find genes with related functional annotations grouped together. While some of these associations are also found using a standard expression level analysis, many are identified exclusively through the dynamic analysis. These results strongly suggest that the analysis of gene expression dynamics is a necessary and important tool for studying regulatory and other functional relationships among genes. The source code developed for this investigation is freely available to all non-commercial investigators by contacting the authors.

Cluster Analysis↗

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant↗

Chromosome-level genome assembly and annotation of Petunia hybrida.

Petunia hybrida is the world's most popular garden plant and is regarded as a supermodel for studying the biology associated with the Asterid clade, the largest of the two major groups of flowering plants. Unlike other Solanaceae, petunia has a base chromosome number of seven, not 12. This along with recombination suppression has previously hindered efforts to assemble its genome to chromosome level. Here we achieve a chromosome-level assembly for P. hybrida using a combination of short-read and long-read sequencing, optical mapping (Bionano) and Hi-C technologies. The resulting assembly spans 1253.6 Mb with a BUSCO score of 99.8%. A total of 35,089 genes were predicted and of those 29,655 were functionally annotated. Syntenic regions between petunia, tomato and pepper were identified, highlighting rearrangements that have occurred since their divergence indicating that the 12 chromosomes of Solanaceae did not originate from whole genome duplication of an ancestral species with seven chromosomes like petunia. This assembly will enhance trait mapping efficiency and serve as a valuable resource for functional genomic studies.

Petunia↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

Genome and protein evolution in eukaryotes.

The past year has seen the completion of the genome sequence of the flowering plant Arabidopsis thaliana and the initial sequence reports of the human genome. The availability of completely sequenced eukaryotic genomes from disparate phylogenetic lineages has opened the door to comparative analyses and a better understanding of the evolutionary processes shaping genomes. Complex many-to-many relationships between genes from different species appear to be the norm, suggesting that transfer of detailed functional annotation will not be straightforward. In addition to expansion and contraction of gene families, new genes evolve from recombination of pre-existing domains, although some domain families do appear to have evolved recently and to be specific to restricted phylogenetic lineages. The overall picture is of a huge diversity of gene content within eukaryotic genomes, reflecting different functional demands in different species.

Animals↗

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles↗