PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

16S rRNA and Metagenomic Datasets of Gastrointestinal Microbiota in Fetal and 7-Day-Old Goat Kids.

The perinatal period (from late gestation to the neonatal stage) in ruminants is a critical phase for fetal organ maturation, where ecological succession of gastrointestinal microbial communities significantly impacts livestock production efficiency. However, research remains insufficient regarding the distribution patterns and functional annotation of microbial communities across different gastrointestinal compartments during this period. This study characterized early microbiota dynamics in Hutianshi Goats using 16S rRNA sequencing (4 fetal goats at 90 ± 10 gestational days) and metagenomics (3 7-day-old goat kids). The fetal goat group generated 852,694 valid reads, yielding 688,277 high-quality reads after chimera removal for downstream analysis. The 7-day-old goat kids group produced 1,081,588,182 final valid reads, after data processing and assembly, 8,561,345 contigs were generated. Gene prediction identified 6,095,352 genes. Multi-database annotations (NR, KEGG, CAZy, etc.) revealed functional potential and antimicrobial resistance traits. The public release of this dataset facilitates academic understanding of microbial community dynamics and host-microbe interactions during this developmental stage, providing both theoretical foundations and data resources for ruminant developmental biology and precision breeding regulation.

Animals↗

Affinity purification-mass spectrometry. Powerful tools for the characterization of protein complexes.

Multi-protein complexes are emerging as important entities of biological activity inside cells that serve to create functional diversity by contextual combination of gene products and, at the same time, organize the large number of different proteins into functional units. Many a time, when studying protein complexes rather than individual proteins, the biological insight gained has been fundamental, particularly in cases in which proteins with no previous functional annotation could be placed into a functional context derived from their 'molecular environment'. In this minireview, we summarize the current state of the art for the retrieval of multiprotein complexes by affinity purification and their analysis by mass spectrometry. The advances in technology made over the past few years now enable the study of protein complexes on a proteomic scale and it can be anticipated that the knowledge gathered from such projects will fuel drug target discovery and validation pipelines and that the technology is also going to prove valuable in the emerging field of systems biology.

Animals↗

GUANinE v1.1 reveals complementarity of supervised and genomic language models.

There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.

Genomics↗

Incorporating gene functions as priors in model-based clustering of microarray gene expression data.

MOTIVATION: Cluster analysis of gene expression profiles has been widely applied to clustering genes for gene function discovery. Many approaches have been proposed. The rationale is that the genes with the same biological function or involved in the same biological process are more likely to co-express, hence they are more likely to form a cluster with similar gene expression patterns. However, most existing methods, including model-based clustering, ignore known gene functions in clustering. RESULTS: To take advantage of accumulating gene functional annotations, we propose incorporating known gene functions as prior probabilities in model-based clustering. In contrast to a global mixture model applicable to all the genes in the standard model-based clustering, we use a stratified mixture model: one stratum corresponds to the genes of unknown function while each of the other ones corresponding to the genes sharing the same biological function or pathway; the genes from the same stratum are assumed to have the same prior probability of coming from a cluster while those from different strata are allowed to have different prior probabilities of coming from the same cluster. We derive a simple EM algorithm that can be used to fit the stratified model. A simulation study and an application to gene function prediction demonstrate the advantage of our proposal over the standard method. CONTACT: weip@biostat.umn.edu

Algorithms↗

Incorporating biological knowledge into distance-based clustering analysis of microarray gene expression data.

MOTIVATION: Because co-expressed genes are likely to share the same biological function, cluster analysis of gene expression profiles has been applied for gene function discovery. Most existing clustering methods ignore known gene functions in the process of clustering. RESULTS: To take advantage of accumulating gene functional annotations, we propose incorporating known gene functions into a new distance metric, which shrinks a gene expression-based distance towards 0 if and only if the two genes share a common gene function. A two-step procedure is used. First, the shrinkage distance metric is used in any distance-based clustering method, e.g. K-medoids or hierarchical clustering, to cluster the genes with known functions. Second, while keeping the clustering results from the first step for the genes with known functions, the expression-based distance metric is used to cluster the remaining genes of unknown function, assigning each of them to either one of the clusters obtained in the first step or some new clusters. A simulation study and an application to gene function prediction for the yeast demonstrate the advantage of our proposal over the standard method.

Algorithms↗

Comprehensive analysis of orthologous protein domains using the HOPS database.

One of the most reliable methods for protein function annotation is to transfer experimentally known functions from orthologous proteins in other organisms. Most methods for identifying orthologs operate on a subset of organisms with a completely sequenced genome, and treat proteins as single-domain units. However, it is well known that proteins are often made up of several independent domains, and there is a wealth of protein sequences from genomes that are not completely sequenced. A comprehensive set of protein domain families is found in the Pfam database. We wanted to apply orthology detection to Pfam families, but first some issues needed to be addressed. First, orthology detection becomes impractical and unreliable when too many species are included. Second, shorter domains contain less information. It is therefore important to assess the quality of the orthology assignment and avoid very short domains altogether. We present a database of orthologous protein domains in Pfam called HOPS: Hierarchical grouping of Orthologous and Paralogous Sequences. Orthology is inferred in a hierarchic system of phylogenetic subgroups using ortholog bootstrapping. To avoid the frequent errors stemming from horizontally transferred genes in bacteria, the analysis is presently limited to eukaryotic genes. The results are accessible in the graphical browser NIFAS, a Java tool originally developed for analyzing phylogenetic relations within Pfam families. The method was tested on a set of curated orthologs with experimentally verified function. In comparison to tree reconciliation with a complete species tree, our approach finds significantly more orthologs in the test set. Examples for investigating gene fusions and domain recombination using HOPS are given.

Animals↗

Unravelling the genomic potential of sponge-associated Streptomyces sp. BLC 17-3 from Indonesia for mannooligosaccharide production.

This research aims to show the promising capacity of Streptomyces sp. BLC 17-3 to produce high β-mannanase enzymes and generate mannooligosaccharide (MOS) such as mannobiose, mannotriose, mannotetraose and mannopentaose when exposed to mannan polymers. Streptomyces sp. BLC 17-3 was isolated from the sponge (Rhabdastrella globostellata) Put4 obtained from the marine waters of Putus Island in Bitung, North Sulawesi, Indonesia. The characterization results showed that the peak enzyme activity was achieved at 50 mM sodium acetate, 6.0 pH, and 60 °C temperature on the seventh day of production with a value of 155.77 ± 3.21 U/mL. The SDS-PAGE and zymograms also showed that the size of the enzyme molecule was approximately ±34.8-49.1 kDa. Moreover, whole-genome sequencing was conducted to identify the genetic basis of MOS-synthesizing capabilities in the selected strain, followed by functional annotation of genes encoding mannan degradation and associated functions. The results showed an 8,248,862 Mb complete draft genome of the strain which comprised 111 predicted gene models. Gene annotation also provided important information about the location and function of protein-encoding genes. A total of 6 mannan degradation-related genes encoding mannanase-related metabolism were identified and the three-dimensional structures were predicted using AlphaFold 3. This characterization and modeling further enhanced the bioprospecting and development of this strain which exhibited efficient mannose metabolism. The results showed Streptomyces sp. BLC 17-3 as a promising microorganism for the future bioproduction of MOS which were discovered to have the capability of serving as a potential prebiotic substance to enhance digestion and promote health.

Bioprospecting↗

Development and evaluation of an automated annotation pipeline and cDNA annotation system.

Manual curation has long been held to be the "gold standard" for functional annotation of DNA sequence. Our experience with the annotation of more than 20,000 full-length cDNA sequences revealed problems with this approach, including inaccurate and inconsistent assignment of gene names, as well as many good assignments that were difficult to reproduce using only computational methods. For the FANTOM2 annotation of more than 60,000 cDNA clones, we developed a number of methods and tools to circumvent some of these problems, including an automated annotation pipeline that provides high-quality preliminary annotation for each sequence by introducing an "uninformative filter" that eliminates uninformative annotations, controlled vocabularies to accurately reflect both the functional assignments and the evidence supporting them, and a highly refined, Web-based manual annotation tool that allows users to view a wide array of sequence analyses and to assign gene names and putative functions using a consistent nomenclature. The ultimate utility of our approach is reflected in the low rate of reassignment of automated assignments by manual curation. Based on these results, we propose a new standard for large-scale annotation, in which the initial automated annotations are manually investigated and then computational methods are iteratively modified and improved based on the results of manual curation.

Animals↗

Whole blood transcriptome profile identifies motor neurone disease RNA biomarker signatures.

Blood-based biomarkers for motor neuron disease are needed for better diagnosis, progression prediction, and clinical trial monitoring. We used whole blood-derived total RNA and performed whole transcriptome analysis to compare the gene expression profiles in (motor neurone disease) MND patients to the control subjects. We compared 42 MND patients to 42 aged and sex-matched healthy controls and described the whole transcriptome profile characteristic for MND. In addition to the formal differential analysis, we performed functional annotation of the genomics data and identified the molecular pathways that are differentially regulated in MND patients. We identified 12,972 genes differentially expressed in the blood of MND patients compared to age and sex-matched controls. Functional genomic annotation identified activation of the pathways related to neurodegeneration, RNA transcription, RNA splicing and extracellular matrix reorganisation. Blood-based whole transcriptomic analysis can reliably differentiate MND patients from controls and can provide useful information for the clinical management of the disease and clinical trials.

Humans↗

Optimization of protoplast based DNA isolation and genome analysis in a gamma-irradiated Aspergillus niger mutant strain.

Aspergillus niger is an important industrial fungus widely used for citric acid production and a range of biotechnological applications. In this study, a protoplast-based DNA isolation protocol was optimized for a gamma-irradiated A. niger AN-L103_M1 mutant strain, followed by whole-genome sequencing and functional genome analysis. Protoplast yield was strongly influenced by enzyme concentration and the molarity of the osmotic stabilizer. The highest yield was achieved at an enzyme concentration of 50&#xa0;mg/mL (2.487&#x2009;&#xb1;&#x2009;0.04&#x2009;&#xd7;&#x2009;10&#x2078; cells/mL) and 0.8&#xa0;M KCl (2.550&#x2009;&#xb1;&#x2009;0.06&#x2009;&#xd7;&#x2009;10&#x2078; cells/mL), with both factors showing significant effects (p&#x2009;<&#x2009;0.0001) in GraphPad Prism 11.0.0. Whole-genome sequencing performed using an Illumina NovaSeq 6000 platform yielded a 37.06&#xa0;Mb draft genome assembled into 537 contigs, with an N50 of 363,084&#xa0;bp and a GC content of 48.2%. BUSCO 14 analysis showed high completeness (97.95% complete BUSCOs). Functional annotation and KEGG pathway mapping identified genes involved in glycolysis, the tricarboxylic acid cycle, and citrate biosynthesis, while biosynthetic gene cluster analysis revealed diverse potential for secondary metabolite production. These findings provide an optimized workflow for protoplast-based DNA isolation and genome-scale functional analysis in A. niger, proposing a basis for future comparative genomics, transformation studies, and experimentally validated metabolic engineering.

Aspergillus niger↗

Pinpointing genomic regions conferring herbicide tolerance in cassava via genome-wide association mapping.

Cassava (Manihot esculenta Crantz) is a tropical crop of major socioeconomic importance, whose productivity can be limited by sensitivity to herbicides used for weed management. This study aimed to perform a genome-wide association study (GWAS) in 194 cassava genotypes to identify genomic regions associated with tolerance to the herbicides mesotrione, S-metolachlor, and chloransulam-methyl. The evaluations performed at 3, 6, 9, 15, and 30 days after application (DAA) were used to characterize the temporal progression of phytotoxicity. Based on this analysis, the phenotype obtained at 9 days after application (PhytoX9DAA) was selected for genome-wide association analyses because it represented the period of greatest symptom expression and the highest discrimination among genotypes. GWAS analyses were performed using de-regressed BLUPs and the MLM, MLMM, and BLINK models, incorporating kinship (K) and population structure (Q) matrices. Significant markers were detected across multiple chromosomes, and the corresponding genomic windows contained candidate genes with functional annotations related to herbicide response. The predominant functional categories included membrane transport, channel activity, signal peptide processing, protein phosphorylation, cellular signaling, and metabolic regulation. Key candidate genes included Manes.02G151900 and Manes.02G152700 (chromosome 2), associated with transmembrane transport and signal peptide processing; Manes.09G060900 (chromosome 9), associated with protein kinase activity, ATP binding, and protein phosphorylation; and Manes.15G083800 and Manes.15G084000 (chromosome 15), associated with S-adenosylmethionine-dependent methyltransferase activity, membrane-related functions, and protein phosphorylation. These genes participate in biochemical pathways involved in cellular signaling, membrane transport, and metabolic regulation that may contribute to herbicide tolerance. Overall, the results demonstrate that herbicide tolerance in cassava is a quantitative and polygenic trait governed by numerous small-effect loci. The integration of cellular signaling, metabolic regulation, and membrane transport supports the physiological resilience of the species under chemical exposure, providing valuable insights for breeding strategies and marker-assisted selection.

Genome-Wide Association Study↗

A genetic screen implicates miRNA-372 and miRNA-373 as oncogenes in testicular germ cell tumors.

Endogenous small RNAs (miRNAs) regulate gene expression by mechanisms conserved across metazoans. While the number of verified human miRNAs is still expanding, only few have been functionally annotated. To perform genetic screens for novel functions of miRNAs, we developed a library of vectors expressing the majority of cloned human miRNAs and created corresponding DNA barcode arrays. In a screen for miRNAs that cooperate with oncogenes in cellular transformation, we identified miR-372 and miR-373, each permitting proliferation and tumorigenesis of primary human cells that harbor both oncogenic RAS and active wild-type p53. These miRNAs neutralize p53-mediated CDK inhibition, possibly through direct inhibition of the expression of the tumor-suppressor LATS2. We provide evidence that these miRNAs are potential novel oncogenes participating in the development of human testicular germ cell tumors by numbing the p53 pathway, thus allowing tumorigenic growth in the presence of wild-type p53.

Cells, Cultured↗

A wiring of the human nucleolus.

Recent proteomic efforts have created an extensive inventory of the human nucleolar proteome. However, approximately 30% of the identified proteins lack functional annotation. We present an approach of assigning function to uncharacterized nucleolar proteins by data integration coupled to a machine-learning method. By assembling protein complexes, we present a first draft of the human ribosome biogenesis pathway encompassing 74 proteins and hereby assign function to 49 previously uncharacterized proteins. Moreover, the functional diversity of the nucleolus is underlined by the identification of a number of protein complexes with functions beyond ribosome biogenesis. Finally, we were able to obtain experimental evidence of nucleolar localization of 11 proteins, which were predicted by our platform to be associates of nucleolar complexes. We believe other biological organelles or systems could be "wired" in a similar fashion, integrating different types of data with high-throughput proteomics, followed by a detailed biological analysis and experimental validation.

Artificial Intelligence↗

Expression changes in tolerant murine cardiac allografts after gene therapy with a lentiviral vector expressing alpha1,3 galactosyltransferase.

Comparison of intragraft gene expression changes in tolerant cardiac allograft models may provide the basis for identifying pathways involved in graft survival. Our laboratory has previously demonstrated that tolerance to the gal alpha1,3 gal epitope, the major target of rejection of wild-type pig hearts in human cardiac transplantation, can be achieved after transplantation with bone marrow transduced with a lentiviral vector expressing alpha1,3 galactosyltransferase. We now present intracardiac gene expression changes associated with long-term tolerance in this model. Biotin-labeled cRNA was hybridized to Affymetrix GeneChip 430 2.0 Mouse Genome Arrays. Data were subjected to functional annotation analysis to identify genes of known function in which expression was increased or decreased by at least 2-fold (t-test, P < .05) in tolerant gal+/+ wild-type hearts as compared to transplanted syngeneic controls. Tolerant hearts demonstrated increased expression of genes associated with the stress response, modulation of immune function and cell survival (HSPa9a, CD56, and Akt1s1), and decreased expression of several immunoregulatory genes (CD209, CD26, and PDE4b). These data suggest that tolerance may be associated with activation of immunomodulatory and survival pathways.

Animals↗

Prediction of protein function from protein sequence and structure.

The sequence of a genome contains the plans of the possible life of an organism, but implementation of genetic information depends on the functions of the proteins and nucleic acids that it encodes. Many individual proteins of known sequence and structure present challenges to the understanding of their function. In particular, a number of genes responsible for diseases have been identified but their specific functions are unknown. Whole-genome sequencing projects are a major source of proteins of unknown function. Annotation of a genome involves assignment of functions to gene products, in most cases on the basis of amino-acid sequence alone. 3D structure can aid the assignment of function, motivating the challenge of structural genomics projects to make structural information available for novel uncharacterized proteins. Structure-based identification of homologues often succeeds where sequence-alone-based methods fail, because in many cases evolution retains the folding pattern long after sequence similarity becomes undetectable. Nevertheless, prediction of protein function from sequence and structure is a difficult problem, because homologous proteins often have different functions. Many methods of function prediction rely on identifying similarity in sequence and/or structure between a protein of unknown function and one or more well-understood proteins. Alternative methods include inferring conservation patterns in members of a functionally uncharacterized family for which many sequences and structures are known. However, these inferences are tenuous. Such methods provide reasonable guesses at function, but are far from foolproof. It is therefore fortunate that the development of whole-organism approaches and comparative genomics permits other approaches to function prediction when the data are available. These include the use of protein-protein interaction patterns, and correlations between occurrences of related proteins in different organisms, as indicators of functional properties. Even if it is possible to ascribe a particular function to a gene product, the protein may have multiple functions. A fundamental problem is that function is in many cases an ill-defined concept. In this article we review the state of the art in function prediction and describe some of the underlying difficulties and successes.

Amino Acid Sequence↗

Getting to synaptic complexes through systems biology.

Large numbers of synaptic components have been identified, but the effect so far on our understanding of synaptic function is limited. Now, network maps and annotated functions of individual components have been used in a systems biology approach to analyzing the function of NMDA receptor complexes at synapses, identifying biologically relevant modular networks within the complex.

Animals↗

Association algorithm to mine the rules that govern enzyme definition and to classify protein sequences.

BACKGROUND: The number of sequences compiled in many genome projects is growing exponentially, but most of them have not been characterized experimentally. An automatic annotation scheme must be in an urgent need to reduce the gap between the amount of new sequences produced and reliable functional annotation. This work proposes rules for automatically classifying the fungus genes. The approach involves elucidating the enzyme classifying rule that is hidden in UniProt protein knowledgebase and then applying it for classification. The association algorithm, Apriori, is utilized to mine the relationship between the enzyme class and significant InterPro entries. The candidate rules are evaluated for their classificatory capacity. RESULTS: There were five datasets collected from the Swiss-Prot for establishing the annotation rules. These were treated as the training sets. The TrEMBL entries were treated as the testing set. A correct enzyme classification rate of 70% was obtained for the prokaryote datasets and a similar rate of about 80% was obtained for the eukaryote datasets. The fungus training dataset which lacks an enzyme class description was also used to evaluate the fungus candidate rules. A total of 88 out of 5085 test entries were matched with the fungus rule set. These were otherwise poorly annotated using their functional descriptions. CONCLUSION: The feasibility of using the method presented here to classify enzyme classes based on the enzyme domain rules is evident. The rules may be also employed by the protein annotators in manual annotation or implemented in an automatic annotation flowchart.

Algorithms↗

Genome-scale gene function prediction using multiple sources of high-throughput data in yeast Saccharomyces cerevisiae.

Characterizing gene function is one of the major challenging tasks in the post-genomic era. To address this challenge, we have developed GeneFAS (Gene Function Annotation System), a new integrated probabilistic method for cellular function prediction by combining information from protein-protein interactions, protein complexes, microarray gene expression profiles, and annotations of known proteins through an integrative statistical model. Our approach is based on a novel assessment for the relationship between (1) the interaction/correlation of two proteins' high-throughput data and (2) their functional relationship in terms of their Gene Ontology (GO) hierarchy. We have developed a Web server for the predictions. We have applied our method to yeast Saccharomyces cerevisiae and predicted functions for 1548 out of 2472 unannotated proteins.

Automation↗