PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Functional screening of ZIP8 naturally occurring variants identifies pathogenic mutations and trafficking defects.

The rapid expansion of human genomic data has revealed a large number of naturally occurring variants, creating a major challenge for functional annotation. The human metal transporter SLC39A8 (ZIP8) is a clinically important divalent metal transporter, yet most of its documented variants remain uncharacterized. Here, we developed a workflow to functionally evaluate ZIP8 variants by integrating laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) with scaled-up cell-based transport assays. Using this method, we systematically analyzed 33 naturally occurring missense variants located in the extracellular domain (ECD) of ZIP8. The assay enables direct quantification of intracellular metal accumulation with substantially improved throughput (∼150 samples per hour). Functional screening identified 14 potential pathogenic variants with significantly reduced transport activity. Comparison with computational predictions revealed a moderate correlation between activity and AlphaMissense pathogenicity scores (R2 = 0.423), while an error rate of ∼20% for AlphaMissense underscores the need for experimental validation. Flow cytometry analysis showed that most loss-of-function variants exhibit impaired trafficking of the protein to the cell surface possibly due to mutation-caused protein misfolding or instability. Structural mapping of activity-compromised variants, together with functional assessment of the ZIP8-ECD, highlights the importance of this domain in ZIP8 expression and intracellular protein trafficking. Together, this work establishes a scalable approach for functional screening of metal transporter variants and provides new insights into the structure-function relationships of ZIP8.

Journal Article↗

Dental wastewater reveals a hidden reservoir of oral bacteriophage diversity.

Bacteriophages (phages) are being explored as alternatives or complements to antibiotics because of their ability to selectively kill bacterial pathogens. However, phages that infect many oral bacteria remain undiscovered. Here, we discovered that dental wastewater harbors previously underexplored phage diversity. Viral particles concentrated from dental wastewater displayed diverse morphologies, including abundant filamentous phage-like particles. Deep long-read metagenomic sequencing of concentrated viral particles generated 7.4 billion bases of sequence data and yielded 255 medium- to high-quality viral operational taxonomic units (vOTUs), including 46 predicted complete genomes. Comparison with large phage databases revealed that 63 of these 255 vOTUs had no detectable match, indicating that extensive sequencing of dental wastewater substantially expands the number of potential bacteriophages associated with the human oral microbiome. Host prediction linked many vOTUs to oral-associated bacterial taxa, including species with few or no previously reported phages, such as Porphyromonas gingivalis, Tannerella forsythia, and Candidatus Saccharibacteria. Functional annotation identified diverse genes associated with antiphage defense systems within a subset of vOTUs, suggesting that oral phages may contribute to the movement of genes encoding bacterial immune functions within the oral microbiome. Together, these findings expand the known oral phageome and show that dental wastewater contains a largely untapped diversity of phages.IMPORTANCEThe human oral cavity contains a diverse microbial community, but the bacteriophages (phages) that infect many oral bacteria remain poorly characterized. This gap limits our understanding of how phages shape oral microbial communities. Here, we show that dental wastewater is an underexplored source of oral phage diversity. Deep long-read metagenomic sequencing revealed 255 medium- to high-quality phage operational taxonomic units, many of which are not present in existing oral phage databases. These genomes include predicted phages of periodontal disease-associated bacteria and other oral taxa with few or no known phages. Dental wastewater therefore expands the known human oral phageome and reveals candidate phages linked to bacteria associated with oral health and disease.

Bacteriophages↗

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals↗

The Mycobacterium tuberculosis Transposon Sequencing Database (MtbTnDB): A Large-Scale Guide to Genetic Conditional Essentiality.

Characterizing genetic essentiality across various conditions is fundamental for understanding gene function. Transposon sequencing (TnSeq) is a powerful technique to generate genome-wide essentiality profiles in bacteria and has been extensively applied to Mycobacterium tuberculosis (Mtb). Dozens of TnSeq screens have yielded valuable insights into the biology of Mtb in vitro, inside macrophages, and in model host organisms. Despite their value, these Mtb TnSeq profiles have not been standardized or collated into a single, easily searchable database. This results in significant challenges when attempting to query and compare these resources, limiting our ability to obtain a comprehensive and consistent understanding of genetic conditional essentiality in Mtb. We address this problem by building a central repository of publicly available Mtb TnSeq screens, the Mtb transposon sequencing database (MtbTnDB). The MtbTnDB is a living resource that encompasses to date ≈150 standardized TnSeq screens, enabling open access to data, visualizations, and functional predictions through an interactive web app (www.mtbtndb.app). We conduct several statistical analyses on the complete database, such as demonstrating that (i) genes in the same genomic neighborhood have similar TnSeq profiles, and (ii) clusters of genes with similar TnSeq profiles are enriched for genes from similar functional categories. We further analyze the performance of machine learning models trained on TnSeq profiles to predict the functional annotation of orphan genes in Mtb. By facilitating the comparison of TnSeq screens across conditions, the MtbTnDB will accelerate the exploration of conditional genetic essentiality, provide insights into the functional organization of Mtb genes, and help predict gene function in this important human pathogen.

DNA Transposable Elements↗

Functional screening of ZIP8 naturally occurring variants identifies pathogenic mutations and trafficking defects.

The rapid expansion of human genomic data has revealed a large number of naturally occurring variants, creating a major challenge for functional annotation. The human metal transporter SLC39A8 (ZIP8) is a clinically important, promiscuous divalent metal transporter, yet most of its documented variants remain uncharacterized. Here, we developed a workflow to functionally evaluate ZIP8 variants by integrating laser ablation inductively coupled plasma time-of-flight mass spectrometry (LA-ICP-TOF-MS) with scaled-up cell-based transport assays. Using this method, we systematically analyzed 33 naturally occurring missense variants located in the extracellular domain (ECD) of ZIP8. The assay enables direct quantification of intracellular metal accumulation with substantially improved throughput (~150 samples per hour). Functional screening identified 14 potential pathogenic variants with significantly reduced transport activity. Comparison with computational predictions revealed a moderate correlation between activity and AlphaMissense pathogenicity scores (R2 = 0.423), while an error rate of ~20% underscores the need for experimental validation. Flow cytometry analysis showed that most loss-of-function variants exhibit impaired trafficking of the protein to the cell surface possibly due to mutation-caused protein misfolding or instability. Structural mapping of activity-compromised variants, together with functional assessment of the ZIP8-ECD, highlights the importance of this domain in ZIP8 expression and intracellular trafficking. Together, this work establishes a scalable approach for functional screening of metal transporter variants and provides new insights into the structure-function relationships of ZIP8.

Journal Article↗

Genes linked by fusion events are generally of the same functional category: a systematic analysis of 30 microbial genomes.

Recent work in computational genomics has shown that a functional association between two genes can be derived from the existence of a fusion of the two as one continuous sequence in another genome. For each of 30 completely sequenced microbial genomes, we established all such fusion links among its genes and determined the distribution of links within and among 15 broad functional categories. We found that 72% of all fusion links related genes of the same functional category. A comparison of the distribution of links to simulations on the basis of a random model further confirmed the significance of intracategory fusion links. Where a gene of annotated function is linked to an unclassified gene, the fusion link suggests that the two genes belong to the same functional category. The predictions based on fusion links are shown here for Methanobacterium thermoautotrophicum, and another 661 predictions are available at http://fusion.bu.edu.

Gene Expression Regulation, Bacterial↗

Genome-Wide Identification and Characterization of the TBL Gene Family and Temporal Expression Dynamics During Powdery Mildew Infection in Cucumber (Cucumis sativus).

Cell-wall polysaccharide O-acetylation contributes to cell-wall assembly, organ development, and plant-pathogen interactions, but the cucumber TBL gene family remains poorly characterized. Here, 37 CsTBL genes were identified genome-wide and analyzed using phylogenetic, syntenic, conserved-motif, gene-structure, promoter, protein-structure, Gene Ontology, and transcriptome approaches, followed by RT-qPCR analysis after powdery mildew inoculation. All CsTBL proteins contained the conserved GDS and DxxH motifs, whereas accessory motifs and predicted structural features varied among clades. Intraspecific analysis identified dispersed, WGD/segmental, and tandem duplication categories, and cross-species synteny was more extensive with melon than with Arabidopsis. Homology-derived annotations associated CsTBL genes with cell-wall polysaccharide metabolism, Golgi/endomembrane compartments, and O-acetyltransferase activity, including six genes assigned to xylan O-acetyltransferase-related annotations. Expression profiling revealed tissue- and developmental-stage-dependent patterns, whereas the publicly available powdery mildew RNA-seq dataset provided descriptive temporal expression profiles in Podosphaera xanthii-inoculated samples. Independent RT-qPCR analysis using time-matched mock controls revealed distinct post-inoculation responses among six selected genes. Relative to the corresponding mock controls, CsTBL2 was consistently repressed; CsTBL15 showed transient induction at 1 dpi followed by repression; CsTBL24 exhibited a biphasic response; CsTBL25 was induced at all sampled post-inoculation time points; CsTBL26 showed progressive induction; and CsTBL30 reached its highest observed expression level at 3 dpi. Integrated functional annotation and expression evidence highlighted CsTBL26 as a priority candidate for further functional characterization, while CsTBL24 and CsTBL25 represented fruit-associated candidates with distinct powdery mildew responses; CsTBL30 remained an additional strongly infection-responsive candidate. These findings provide an evolutionary and expression-based framework for the functional characterization of the cucumber TBL gene family.

O-acetylation↗

Genetic overlap between depression and C-reactive protein levels: Evidence from a cross-trait analysis.

Inflammation and depression have been consistently associated, with elevated C-reactive protein (CRP) levels observed in a significant subset of affected individuals. However, the genetic mechanisms underlying this association remain poorly understood. We integrated results from large-scale genome-wide association studies (GWAS) of depression and CRP levels in a cross-trait analysis specifically focusing on identifying horizontally pleiotropic loci. Identified variants were stratified as concordant versus discordant based on their direction of effects on the two traits and followed up using functional annotation, gene set enrichment, and colocalization analyses. We also explored causal relationships using Mendelian Randomization (MR) analysis with extensive sensitivity analyses, including adjustment for body mass index (BMI). We identified 9 novel loci. Functional analyses revealed that concordant loci were enriched in genes linked to immune and inflammatory processes, while discordant loci mostly mapped to metabolic pathways, including lipid regulation. MR provided strong evidence for body mass index driving a causal relationship between the genetic liability of depression on CRP levels. Our findings suggest that the association between depression and CRP levels is partly driven by shared genetic influences, pointing to different biological pathways depending on whether genetic effects are concordant or discordant. These results underscore the importance of considering effect direction when assessing the genetic overlap between depression and inflammatory processes. In addition, they highlight BMI as a key factor in the causal relationship between depression and systemic inflammation.

C-Reactive Protein↗

A protein-protein interaction map of the Caenorhabditis elegans 26S proteasome.

The ubiquitin-proteasome proteolytic pathway is pivotal in most biological processes. Despite a great level of information available for the eukaryotic 26S proteasome-the protease responsible for the degradation of ubiquitylated proteins-several structural and functional questions remain unanswered. To gain more insight into the assembly and function of the metazoan 26S proteasome, a two-hybrid-based protein interaction map was generated using 30 Caenorhabditis elegans proteasome subunits. The results recapitulate interactions reported for other organisms and reveal new potential interactions both within the 19S regulatory complex and between the 19S and 20S subcomplexes. Moreover, novel potential proteasome interactors were identified, including an E3 ubiquitin ligase, transcription factors, chaperone proteins and other proteins not yet functionally annotated. By providing a wealth of novel biological hypotheses, this interaction map constitutes a framework for further analysis of the ubiquitin-proteasome pathway in a multicellular organism amenable to both classical genetics and functional genomics.

Animals↗

RAREsim2: flexible simulation of rare variant genetic data using real haplotypes.

MOTIVATION: Realistic simulated data is critical for advancing methodological development and optimizing study design in genetics research. However, many genetic simulation tools are unable to replicate the distribution of rare variants or incorporate key genetic information, such as functional annotations and linkage disequilibrium. RAREsim, an accurate rare variant simulation algorithm that uses real genetic haplotypes, was developed to address these limitations. Here, we introduce RAREsim2, an update that provides both streamlined software and new functionalities for simulating individual-level differences (e.g., case-control status, technological or batch effects) and variant-level differences to represent a variety of causal models. RESULTS: We demonstrate RAREsim2's utility with three rare variant association methods (Burden, SKAT, and SKAT-O) across several simulation scenarios, including various genetic ancestries, gene sizes, strengths of association, and proportions of risk variants. Type I Error was maintained and the test with the highest power matched previously known patterns. Importantly, real genetic regions can be simulated to include known variant functions and disease associations. Ultimately, RAREsim2 offers additional flexibility and ease in simulating a multitude of realistic genetic scenarios. AVAILABILITY AND IMPLEMENTATION: The RAREsim2 Python package is publicly available on Github (https://github.com/Hendricks-Research-Team/RAREsim2), PyPI (https://pypi.org/project/raresim/), and Zenodo (https://doi.org/10.5281/zenodo.19442523). Code for the example demonstration can be found at https://github.com/JessMurphy/RAREsim2-demo.

Software↗

Genome-Wide Characterization of β-Glucosidase (TaBGLU) Genes in Bread Wheat and Their Expression Under Drought, Cold, and Combined Stress.

Glycoside hydrolase 1 (GH1) β-glucosidases were known to activate hormone conjugates and defense metabolites, yet their genomic organization and stress-response dynamics in wheat remained incompletely defined. We therefore performed an integrated characterization of TaBGLUs spanning phylogeny, gene structure and conserved motifs, subcellular localization, promoter cis-elements, Gene Ontology enrichment, protein-protein interaction networks, and targeted expression profiling. Wheat TaBGLUs partitioned into well-supported clades that shared canonical GH1 catalytic residues and a largely conserved motif scaffold. Subcellular localization predictions indicated predominant nuclear and chloroplast targeting, with a smaller cohort directed to secretory or endomembrane compartments. Promoters were enriched for light-responsive, hormone-related (ABA, JA/SA, auxin, GA) and stress-associated (MYB/WRKY, heat, low temperature) cis-elements, and functional annotations were consistent with roles in carbohydrate and cell-wall metabolism, hormone homeostasis, and defense. Network analysis revealed a densely connected TaBGLU submodule embedded within broader carbohydrate and defense interaction networks, suggesting coordinated or cooperative functions. Expression profiling under cold, drought, and combined drought and cold demonstrated broad stress inducibility, with early activation detected by 6 h, cold-responsive maxima typically at 12 h, drought-responsive peaks predominating at 24 h, and combined stress eliciting both earlier and more sustained expression maxima between 12-24 h. Representative strongly responsive genes included TaBGLU20, TaBGLU44, TaBGLU6, and TaBGLU23, which showed pronounced late induction under combined stress, TaBGLU30, which exhibited an earlier combined-stress peak, and TaBGLU12, which displayed a marked late drought-specific response. Taken together, this integrated genomic, regulatory, and expression atlas refined the wheat BGLU repertoire relative to previous gene model inventories, highlighted candidate TaBGLUs with central network positions and strong stress inducibility, and provided concrete entry points for functional validation and breeding for improved stress resilience.

Triticum↗

Unexpected catalytic site variation in phosphoprotein phosphatase homologues of cofactor-dependent phosphoglycerate mutase.

The cofactor-dependent phosphoglycerate mutase (dPGM) superfamily contains, besides mutases, a variety of phosphatases, both broadly and narrowly substrate-specific. Distant dPGM homologues, conspicuously abundant in microbial genomes, represent a challenge for functional annotation based on sequence comparison alone. Here we carry out sequence analysis and molecular modelling of two families of bacterial dPGM homologues, one the SixA phosphoprotein phosphatases, the other containing various proteins of no known molecular function. The models show how SixA proteins have adapted to phosphoprotein substrate and suggest that the second family may also encode phosphoprotein phosphatases. Unexpected variation in catalytic and substrate-binding residues is observed in the models.

Amino Acid Sequence↗

Annotated expressed sequence tags and cDNA microarrays for studies of brain and behavior in the honey bee.

To accelerate the molecular analysis of behavior in the honey bee (Apis mellifera), we created expressed sequence tag (EST) and cDNA microarray resources for the bee brain. Over 20,000 cDNA clones were partially sequenced from a normalized (and subsequently subtracted) library generated from adult A. mellifera brains. These sequences were processed to identify 15,311 high-quality ESTs representing 8912 putative transcripts. Putative transcripts were functionally annotated (using the Gene Ontology classification system) based on matching gene sequences in Drosophila melanogaster. The brain ESTs represent a broad range of molecular functions and biological processes, with neurobiological classifications particularly well represented. Roughly half of Drosophila genes currently implicated in synaptic transmission and/or behavior are represented in the Apis EST set. Of Apis sequences with open reading frames of at least 450 bp, 24% are highly diverged with no matches to known protein sequences. Additionally, over 100 Apis transcript sequences conserved with other organisms appear to have been lost from the Drosophila genome. DNA microarrays were fabricated with over 7000 EST cDNA clones putatively representing different transcripts. Using probe derived from single bee brain mRNA, microarrays detected gene expression for 90% of Apis cDNAs two standard deviations greater than exogenous control cDNAs. [The sequence data described in this paper have been submitted to Genbank data library under accession nos. BI502708-BI517278. The sequences are also available at http://titan.biotec.uiuc.edu/bee/honeybee_project.htm.]

Animals↗

The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.

While genome sequencing projects are advancing rapidly, EST sequencing and analysis remains a primary research tool for the identification and categorization of gene sequences in a wide variety of species and an important resource for annotation of genomic sequence. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi. shtml) are a collection of species-specific databases that use a highly refined protocol to analyze EST sequences in an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed by first clustering, then assembling EST and annotated gene sequences from GenBank for the targeted species. This process produces a set of unique, high-fidelity virtual transcripts, or Tentative Consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, to provide links between orthologous and paralogous genes and as a resource for comparative sequence analysis.

Animals↗

Gene expression profiling of prolonged cold ischemia and reperfusion in murine heart transplants.

BACKGROUND: Heart transplantation causes complex changes in the biological homeostasis of the graft. Current knowledge is restricted to a few genes and regulation of certain factors involved in ischemia-reperfusion (I/R) injury. Efficient strategies to prevent I/R injury, however, require a better understanding of its mechanisms. Using cDNA microarrays, we investigated gene expression profiles of murine cardiac isografts. METHODS: For microarray hybridization experiments, chips with 8,734 individual target sequences were used. Messenger RNA was extracted from hearts subjected to warm ischemia and different time periods of reperfusion or to prolonged cold ischemia or warm ischemia and transplantation. Native hearts served as controls. RESULTS: A set of 68 sequences was regulated in all hearts. In addition, grafts without cold ischemia showed differential expression of 65 sequences, which were not found in hearts transplanted after cold storage, and which in turn had 38 sequences regulated and not detected in grafts without cold ischemia. Overall, approximately 50% of regulated transcripts are expressed sequence tags (ESTs) with unknown function. Annotated genes encoded immune modulators (20% of sequences), receptor proteins, structural proteins, and proteins involved in metabolism. CONCLUSION: Our data demonstrate expression profiles of hearts subjected to prolonged cold ischemia or transplantation in an isogeneic setting. We have defined functional complexes and detected a substantial amount of ESTs encoding novel proteins. These studies may provide a molecular basis for further functional experiments and may help identify potential targets for modulation of postischemic inflammation.

Animals↗

Predicting phenotype from patterns of annotation.

MOTIVATION: Predicting the outcome of specific experiments (such as the growth of a particular mutant strain in a particular medium) has the potential to allow researchers to devote resources to experiments with higher expected numbers of 'hits'. RESULTS: We use decision trees to predict phenotypes associated with Saccharomyces cerevisiae genes on the basis of Gene Ontology (GO) functional annotations from the Saccharomyces Genome Database (SGD) and other phenotypic annotations from the Yeast Phenotype Catalog at the Munich Information Center for Protein Sequences (MIPS). We assess the methodology in three ways: (1) we use cross-validation on the phenotypic annotations listed in MIPS, and show ROC curves indicating the tradeoff between true-positive rate and false-positive rate; (2) we do a literature-search for 100 of the predicted gene-phenotype associations that are not listed in MIPS, and find evidence for 43 of them; (3) we use deletion strains to experimentally assess 61 predicted gene-phenotype associations not listed in MIPS; significantly more of these deletion strains show abnormal growth than would be expected by chance.

Algorithms↗

Annotated draft genomic sequence from a Streptococcus pneumoniae type 19F clinical isolate.

The public availability of numerous microbial genomes is enabling the analysis of bacterial biology in great detail and with an unprecedented, organism-wide and taxon-wide, broad scope. Streptococcus pneumoniae is one of the most important bacterial pathogens throughout the world. We present here sequences and functional annotations for 2.1-Mbp of pneumococcal DNA, covering more than 90% of the total estimated size of the genome. The sequenced strain is a clinical isolate resistant to macrolides and tetracycline. It carries a type 19F capsular locus, but multilocus sequence typing for several conserved genetic loci suggests that the strain sequenced belongs to a pneumococcal lineage that most often expresses a serotype 15 capsular polysaccharide. A total of 2,046 putative open reading frames (ORFs) longer than 100 amino acids were identified (average of 1,009 bp per ORF), including all described two-component systems and aminoacyl tRNA synthetases. Comparisons to other complete, or nearly complete, bacterial genomes were made and are presented in a graphical form for all the predicted proteins.

DNA, Bacterial↗

Can sequence determine function?

The functional annotation of proteins identified in genome sequencing projects is based on similarities to homologs in the databases. As a result of the possible strategies for divergent evolution, homologous enzymes frequently do not catalyze the same reaction, and we conclude that assignment of function from sequence information alone should be viewed with some skepticism.

Animals↗