PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Methods to map protein interactions in mammalian cells: different tools to address different questions.

In the post-genome era, functional annotation of the predicted gene-sets will be one of the most important upcoming challenges. So-called interactome analysis positions a protein in its subcellular environment by mapping its interaction partners. Such interaction maps are essential for an accurate insight into protein function since many cellular processes are organised to operate in protein complexes. These assemblies have dynamic structures and can interact with each other, two properties which are often controlled by regulated protein expression and modification. Various methods exist to unravel protein interaction circuitries, which can be roughly divided into biochemical and genetic strategies. In this review we focus on the different strategies to study protein-protein interactions in living mammalian cells. Recently developed analytical and screening methods are also addressed.

Animals↗

Annotated genome of the Atlantic dog whelk, Nucella lapillus.

Nucella lapillus is an important player in rocky shore food chains and has been a focal organism of ecological and evolutionary studies for decades. Despite poor dispersal, they have a broad geographic range, which makes them an ideal species to examine isolation by distance and selection across environmental gradients. Here we present the fully annotated genome of N. lapillus generated with Oxford Nanopore Techonology sequencing at ∼37× coverage. The genome assembly is 2.32 Gbp and consists of 2,525 contigs, with an N50 length of 2 Mbp. Repeat annotation identified 2,491 families that cover 67.56% of the genome, which is similar to other gastropods. Despite its large size and high proportion of repeats, the genome is of high quality. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis revealed a score of 96.8%. Functional annotation of the genome produced 45,848 protein-coding genes with a 96.6% BUSCO score. Genomic resources for mollusks lag behind that of other phyla, perhaps because many of their innate characteristics complicate DNA extraction, sequencing, and assembly. This new N. lapillus genome will increase our genomic understanding of the second largest phylum (and the most diverse class within said phylum) and serve as a key resource to advance studies on the organismal biology and population genetics of this iconic species as well as the connection between genomic variation and community-level processes.

Animals↗

Extracting knowledge from dynamics in gene expression.

Most investigations of coordinated gene expression have focused on identifying correlated expression patterns between genes by examining their normalized static expression levels. In this study, we focus on the dynamics of gene expression by seeking to identify correlated patterns of changes in genetic expression level. In doing so, we build upon methods developed in clinical informatics to detect temporal trends of laboratory and other clinical data. We construct relevance networks from Saccharomyces cerevisiae gene-expression dynamics data and find genes with related functional annotations grouped together. While some of these associations are also found using a standard expression level analysis, many are identified exclusively through the dynamic analysis. These results strongly suggest that the analysis of gene expression dynamics is a necessary and important tool for studying regulatory and other functional relationships among genes. The source code developed for this investigation is freely available to all non-commercial investigators by contacting the authors.

Cluster Analysis↗

Bioinformatics for venom and toxin sciences.

Venomous animals produce a myriad of important pharmacological components. The individual components, or venoms (toxins), are used in ion channel and receptor studies, drug discovery, and formulation of insecticides. The toxin data are scattered across public databases which provide sequence and structural descriptions, but very limited functional annotation. The exponential growth of newly identified toxin data has created a need for better data management. Venominformatics is a systematic bioinformatics approach in which classified, consolidated and cleaned venom data are stored into repositories and integrated with advanced bioinformatics tools for the analysis of structure and function of toxins. Venominformatics complements experimental studies and helps reduce the number of essential experiments.

Animals↗

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant↗

Chromosome-level genome assembly and annotation of Petunia hybrida.

Petunia hybrida is the world's most popular garden plant and is regarded as a supermodel for studying the biology associated with the Asterid clade, the largest of the two major groups of flowering plants. Unlike other Solanaceae, petunia has a base chromosome number of seven, not 12. This along with recombination suppression has previously hindered efforts to assemble its genome to chromosome level. Here we achieve a chromosome-level assembly for P. hybrida using a combination of short-read and long-read sequencing, optical mapping (Bionano) and Hi-C technologies. The resulting assembly spans 1253.6 Mb with a BUSCO score of 99.8%. A total of 35,089 genes were predicted and of those 29,655 were functionally annotated. Syntenic regions between petunia, tomato and pepper were identified, highlighting rearrangements that have occurred since their divergence indicating that the 12 chromosomes of Solanaceae did not originate from whole genome duplication of an ancestral species with seven chromosomes like petunia. This assembly will enhance trait mapping efficiency and serve as a valuable resource for functional genomic studies.

Petunia↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

Genome and protein evolution in eukaryotes.

The past year has seen the completion of the genome sequence of the flowering plant Arabidopsis thaliana and the initial sequence reports of the human genome. The availability of completely sequenced eukaryotic genomes from disparate phylogenetic lineages has opened the door to comparative analyses and a better understanding of the evolutionary processes shaping genomes. Complex many-to-many relationships between genes from different species appear to be the norm, suggesting that transfer of detailed functional annotation will not be straightforward. In addition to expansion and contraction of gene families, new genes evolve from recombination of pre-existing domains, although some domain families do appear to have evolved recently and to be specific to restricted phylogenetic lineages. The overall picture is of a huge diversity of gene content within eukaryotic genomes, reflecting different functional demands in different species.

Animals↗

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles↗

Complex Genetics and Regulatory Drivers of Hypermobile Ehlers-Danlos Syndrome: Insights from Genome-Wide Association Study Meta-analysis.

BACKGROUND: Hypermobile Ehlers-Danlos syndrome (hEDS) is the most common subtype of EDS, a group of heritable connective tissue disorders. Clinically, hEDS is defined by generalized joint hypermobility and chronic musculoskeletal pain, but its impact extends beyond the musculoskeletal system. Affected individuals frequently experience autonomic, gastrointestinal, immune, and neuropsychiatric involvement, highlighting both the multisystemic nature of the condition and challenges of diagnosis. In contrast to other EDS subtypes with defined genetic causes, the molecular basis of hEDS has remained elusive. METHODS: We conducted a genome-wide association study (GWAS) of hEDS across three case controls studies, including 1,815 cases and 5,008 ancestry-matched controls. Fixed-effects meta-analysis of 6.2 million variants was complemented with LDAK gene-based association testing, transcriptome-wide association studies, and integrative annotation across multiple tissues and cell types including eQTLs, enhancer marks and open chromatin accessibility profiles, supported by luciferase assays on one candidate variant. LD-score genetic correlations were assessed between hEDS and 19 frequently reported comorbid conditions. RESULTS: Two loci reached genome-wide significance, including a regulatory region near the atypical chemokine receptor 3 gene (ACKR3) on chromosome 2. Functional annotation supports ACKR3 risk alleles colocalize with eQTLs in tibial nerve, alter enhancer activity, and generate a de novo AHR transcription factor regulatory site, implicating neuroimmune and pain signaling pathways. Gene-based and transcriptome-wide analyses identified common variants in a locus containing multiple candidates, including SLC39A13, a zinc transporter critical for connective tissue development previously implicated in a rare form of EDS, and PSMC3, a gene involved in central nervous system development. LD-score regression revealed significant genetic correlations between hEDS and joint hypermobility, myalgic encephalomyelitis/chronic fatigue syndrome, fibromyalgia, depression, anxiety, autism spectrum disorder, migraine, and gastrointestinal diseases. CONCLUSIONS: These results establish the first evidence of common variant contributions to hEDS, supporting a complex, multisystem model involving neuroimmune-stromal dysregulation. Our findings add novel indications to hEDS pathogenesis and provide solid foundations for future molecular definition and therapeutic discovery.

Genome-wide association study↗

Metagenomic Analysis of Gut Microbiome of Persistent Pulmonary Hypertension of the Newborn.

Persistent pulmonary hypertension of the newborn (PPHN) is one of the most common diseases in the neonatal intensive care unit which severely affects neonatal survival. Gut microbes play an increasingly important role in human health, but there are rarely reported how gut microbiota contribute to PPHN. In our study, the metagenomic sequencing of feces from 12 PPHN's neonates and 8 controls were performed to expose the relation between neonatal gut microbes and PPHN disease. Firstly, we found that the abundance of Actinobacteria, Proteobacteria, Bacteroidetes were significantly increased in PPHN compared with controls, but the Firmicutes components was reduced. And some pathogenic strains (like Vibrio metschnikovii) were significantly enriched in the PPHN compared with controls. Secondly, functional annotation of genes found that PPHN up-regulated transmembrane transport, but down-regulated ribosome and ATP binding. Lastly, microbial metabolic pathway enrichment analysis indicated that some metabolic pathway in PPHN were conflicting and contradictory, showed that an abnormally increased metabolism, disturbed protein synthesis and genomic instability in the PPHN neonate. Our results contribute to understanding the changes in the species and function of gut microbiota in PPHN, thus providing a theoretical basis for the explanation and treatment of PPHN.

Gastrointestinal Microbiome↗

Mutation accumulation in a hybrid parthenogenetic vertebrate.

Asexual lineages are thought to experience elevated extinction rates compared with sexual species, yet direct evidence for the underlying genetic causes remains scarce. Muller's ratchet predicts that the absence of recombination in asexual organisms facilitates the accumulation of deleterious mutations, thereby reducing long-term fitness. Here, we test this hypothesis in the hybrid-origin, parthenogenetic whiptail lizard Aspidoscelis tesselatus by integrating short-read RNAseq and long-read IsoSeq data from both the asexual lineage and its parental sexual species. We reconstructed phased transcripts for A. tesselatus to quantify mutation accumulation relative to the parental sexual species. Comparative analyses revealed elevated ω ratios in both parental genomic complements (subgenomes) of the parthenogenetic lineage, consistent with accelerated accumulation of nonsynonymous mutations. Structural variant analyses identified multiple indels in expressed transcripts predicted to disrupt protein domains. Functional annotation indicated that genes affected by both single-nucleotide variants and indels were enriched for roles in chromatin organization, apoptosis regulation, and transcriptional control. While both parental subgenomes showed similar evolutionary patterns, the maternal complement exhibited more structural and missense mutations than the paternal complement. Together, these results provide evidence that mutations accumulate in asexual A. tesselatus in genes involved in core cellular functions, supporting theoretical predictions that Muller's ratchet contributes to mutation accumulation in asexual lineages.

Animals↗

Genome-resolved analysis reveals disruption of gut microbial vitamin B and K2 biosynthesis during Toxoplasma gondii infection in mice.

UNLABELLED: Toxoplasma gondii infection remodels the gut microbiome, yet its impact on microbial vitamin biosynthetic potential and host redox metabolism remains unclear. Here, we integrated mouse gut metagenomes with publicly available metagenome-assembled genomes (MAGs) to construct a genome-resolved atlas of B-vitamin and vitamin K2 biosynthesis. From 45,697 MAGs, we curated 4,771 representative genomes, of which 2,682 met high-quality criteria (completeness &#x2265;90%, contamination <5%). Functional annotation identified 229,717 vitamin-related genes corresponding to 177 Kyoto Encyclopedia of Genes and Genomes (KEGG) orthologs across de novo pathways for eight B vitamins, thiamine (B1), riboflavin (B2), niacin (B3), pantothenate (B5), pyridoxine (B6), biotin (B7), folate (B9), cobalamin (B12), and vitamin K2. Among the high-quality genomes, 1,665 encoded complete de novo pathways for at least one vitamin, highlighting functional specialization and community-level complementarity. Transcripts per million-normalized metagenomic read counts revealed significant differences in KEGG ortholog abundances across six of the nine vitamin pathways. Reanalysis of metagenomic data from infected mice (acute, chronic, and control; n = 10 per group) revealed a stage-dependent reduction in &#x3b1;-diversity of vitamin biosynthesis pathways during acute infection, and a clear &#x3b2;-diversity separation from chronic and control groups. Core niacin biosynthesis genes (nadB, nadA, nadC) displayed phylum-specific redistribution, indicating selective remodeling of microbial NAD+ precursor production under infection-induced metabolic stress. These results suggest that T. gondii infection disrupts cooperative vitamin biosynthetic networks while specifically modulating niacin pathways linked to host NAD+ metabolism. IMPORTANCE: Gut microbes can synthesize essential vitamins, but how infection alters this function is poorly understood. By integrating mouse gut metagenomes with genome-resolved microbial data, we show that Toxoplasma gondii infection reshapes the vitamin biosynthetic potential of the gut microbiome in a stage-dependent manner. Acute infection reduces the diversity of vitamin biosynthesis pathways and shifts the taxonomic distribution of key niacin biosynthesis genes involved in microbial NAD+ precursor production. These findings identify vitamin metabolism, especially niacin-related pathways, as a sensitive functional axis of microbiome remodeling during infection. Our work links microbial taxonomic changes to functional metabolic consequences and suggests that microbiome-mediated regulation of NAD+-related metabolism may contribute to host redox adaptation during T. gondii infection.

B vitamins↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

TAIR: a resource for integrated Arabidopsis data.

The Arabidopsis Information Resource (TAIR; http://arabidopsis.org) provides an integrated view of genomic data for Arabidopsis thaliana. The information is obtained from a battery of sources, including the Arabidopsis user community, the literature, and the major genome centers. Currently TAIR provides information about genes, markers, polymorphisms, maps, sequences, clones, DNA and seed stocks, gene families and proteins. In addition, users can find Arabidopsis publications and information about Arabidopsis researchers. Our emphasis is now on incorporating functional annotations of genes and gene products, genome-wide expression, and biochemical pathway data. Among the tools developed at TAIR, the most notable is the Sequence Viewer, which displays gene annotation, clones, transcripts, markers and polymorphisms on the Arabidopsis genome, and allows zooming in to the nucleotide level. A tool recently released is AraCyc, which is designed for visualization of biochemical pathways. We are also developing tools to extract information from the literature in a systematic way, and building controlled vocabularies to describe biological concepts in collaboration with other database groups. A significant new feature is the integration of the ABRC database functions and stock ordering system, which allows users to place orders for seed and DNA stocks directly from the TAIR site.

Arabidopsis↗

Beyond synexpression relationships: local clustering of time-shifted and inverted gene expression profiles identifies new, biologically relevant interactions.

The complexity of biological systems provides for a great diversity of relationships between genes. The current analysis of whole-genome expression data focuses on relationships based on global correlation over a whole time-course, identifying clusters of genes whose expression levels simultaneously rise and fall. There are, of course, other potential relationships between genes, which are missed by such global clustering. These include activation, where one expects a time-delay between related expression profiles, and inhibition, where one expects an inverted relationship. Here, we propose a new method, which we call local clustering, for identifying these time-delayed and inverted relationships. It is related to conventional gene-expression clustering in a fashion analogous to the way local sequence alignment (the Smith-Waterman algorithm) is derived from global alignment (Needleman-Wunsch). An integral part of our method is the use of random score distributions to assess the statistical significance of each cluster. We applied our method to the yeast cell-cycle expression dataset and were able to detect a considerable number of additional biological relationships between genes, beyond those resulting from conventional correlation. We related these new relationships between genes to their similarity in function (as determined from the MIPS scheme) or their having known protein-protein interactions (as determined from the large-scale two-hybrid experiment); we found that genes strongly related by local clustering were considerably more likely than random to have a known interaction or a similar cellular role. This suggests that local clustering may be useful in functional annotation of uncharacterized genes. We examined many of the new relationships in detail. Some of them were already well-documented examples of inhibition or activation, which provide corroboration for our results. For instance, we found an inverted expression profile relationship between genes YME1 and YNT20, where the latter has been experimentally documented as a bypass suppressor of the former. We also found new relationships involving uncharacterized yeast genes and were able to suggest functions for many of them. In particular, we found a time-delayed expression relationship between J0544 (which has not yet been functionally characterized) and four genes associated with the mitochondria. This suggests that J0544 may be involved in the control or activation of mitochondrial genes. We have also looked at other, less extensive datasets than the yeast cell-cycle and found further interesting relationships. Our clustering program and a detailed website of clustering results is available at http://www.bioinfo.mbb.yale.edu/expression/cluster (or http://www.genecensus.org/expression/cluster).

Algorithms↗

Ab initio protein structure prediction.

Steady progress has been made in the field of ab initio protein folding. A variety of methods now allow the prediction of low-resolution structures of small proteins or protein fragments up to approximately 100 amino acid residues in length. Such low-resolution structures may be sufficient for the functional annotation of protein sequences on a genome-wide scale. Although no consistently reliable algorithm is currently available, the essential challenges to developing a general theory or approach to protein structure prediction are better understood. The energy landscapes resulting from the structure prediction algorithms are only partially funneled to the native state of the protein. This review focuses on two areas of recent advances in ab initio structure prediction-improvements in the energy functions and strategies to search the caldera region of the energy landscapes.

Chemistry, Physical↗

Yeast genomic expression studies using DNA microarrays.

The exploration and characterization of yeast genomic expression programs is providing a wealth of information about yeast biology, as well as other organisms. The intriguing biology of yeast species invites characterization of genomic expression patterns to illuminate the details of cellular physiology. In addition to its value as an interesting organism, yeast maintains its role as an excellent model in which to characterize genomic expression programs. Microarray studies are quickly spreading to plant, animal, and microbial organisms that remain in the early stages of characterization. The extensive knowledge of yeast biology, as well as the relative ease with which yeast studies can be performed and controlled, facilitates interpretation of the genomic expression data. Importantly, existing information about yeast biology, including functional annotations for each gene, is captured and efficiently presented in databases such as the Saccharomyces Genome Database (SGD), the Munich Information Center Yeast Genome Database (MIPS), the Yeast and Pombe Protein Databases (YPD and PPD, respectively), and others. A number of databases also allow the exploration of published genomic expression studies, including the "Expression Connection" at SGD and the Microarray Global Viewer (yMGV) organized by Marc et al. Consulting these databases to retrieve known details about gene function and regulation vastly facilitates interpretation of the genomic expression data, allowing biological hypotheses to be formulated and tested. These hypotheses can be applied to other organisms that may execute genomic expression programs similar to those seen in yeast. Furthermore, as more genomic expression studies in multiple organisms emerge, large-scale data comparisons can be conducted, within and across organisms. Incorporating the results of yeast studies into such comparisons is certain to increase our understanding about the function, regulation, and evolution of genomic expression programs.

Carbocyanines↗