PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

EyeSite: a semi-automated database of protein families in the eye.

The EyeSite is a web-based database of protein families for proteins that function in the eye and their homologous sequences. The resource clusters proteins at different levels of homology in order to facilitate functional annotation of sequences and modelling of proteins from structural homologues. Eye proteins are organized into the tissue types in which they function and are clustered into homologous families using a novel protocol employing the TribeMCL algorithm. Homologous families are further subdivided into sequence clusters for which multiple sequence alignments are generated. Structural annotations from the CATH domain database are provided for nearly 90% of the sequences, and protein family annotations from the Pfam database for approximately 86%. Homology models have also been generated where appropriate. The EyeSite is stored in a relational database and is extensively linked to other online bioinformatics resources to help relate allelic variants, annotations and clinical details to the derived data in the database. The EyeSite is available for online search, sequence information and model retrieval at http://eyesite.cryst.bbk.ac.uk/.

Amino Acid Sequence↗

Data-driven analysis approach for biomarker discovery using molecular-profiling technologies.

High-throughput molecular-profiling technologies provide rapid, efficient and systematic approaches to search for biomarkers. Supervised learning algorithms are naturally suited to analyse a large amount of data generated using these technologies in biomarker discovery efforts. The study demonstrates with two examples a data-driven analysis approach to analysis of large complicated datasets collected in high-throughput technologies in the context of biomarker discovery. The approach consists of two analytic steps: an initial unsupervised analysis to obtain accurate knowledge about sample clustering, followed by a second supervised analysis to identify a small set of putative biomarkers for further experimental characterization. By comparing the most widely applied clustering algorithms using a leukaemia DNA microarray dataset, it was established that principal component analysis-assisted projections of samples from a high-dimensional molecular feature space into a few low dimensional subspaces provides a more effective and accurate way to explore visually and identify data structures that confirm intended experimental effects based on expected group membership. A supervised analysis method, shrunken centroid algorithm, was chosen to take knowledge of sample clustering gained or confirmed by the first step of the analysis to identify a small set of molecules as candidate biomarkers for further experimentation. The approach was applied to two molecular-profiling studies. In the first study, PCA-assisted analysis of DNA microarray data revealed that discrete data structures exist in rat liver gene expression and correlated with blood clinical chemistry and liver pathological damage in response to a chemical toxicant diethylhexylphthalate, a peroxisome-proliferator-activator receptor agonist. Sixteen genes were then identified by shrunken centroid algorithm as the best candidate biomarkers for liver damage. Functional annotations of these genes revealed roles in acute phase response, lipid and fatty acid metabolism and they are functionally relevant to the observed toxicities. In the second study, 26 urine ions identified from a GC/MS spectrum, two of which were glucose fragment ions included as positive controls, showed robust changes with the development of diabetes in Zucker diabetic fatty rats. Further experiments are needed to define their chemical identities and establish functional relevancy to disease development.

Algorithms↗

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans↗

Annotating the human proteome.

The completion of the human genome has shifted the attention from deciphering the sequence to the identification and characterization of the encoded components. The identification and functional annotation of the proteome is here of special interest and starts with the identification of genes and transcripts as a prerequisite of proteome annotation. Gene predictions are very powerful in predicting most of the exons in a genome, but reliable gene structure predictions of both known and novel genes are dependent on existing transcript and protein information. An enormous amount of data already exists on the function of many human proteins, but this is scattered over many resources. Public domain databases are required to manage and collate this information and present it to the user community in both a human and machine readable manner.

Databases, Factual↗

AgBase: a unified resource for functional analysis in agriculture.

Analysis of functional genomics (transcriptomics and proteomics) datasets is hindered in agricultural species because agricultural genome sequences have relatively poor structural and functional annotation. To facilitate systems biology in these species we have established the curated, web-accessible, public resource 'AgBase' (www.agbase.msstate.edu). We have improved the structural annotation of agriculturally important genomes by experimentally confirming the in vivo expression of electronically predicted proteins and by proteogenomic mapping. Proteogenomic data are available from the AgBase proteogenomics link. We contribute Gene Ontology (GO) annotations and we provide a two tier system of GO annotations for users. The 'GO Consortium' gene association file contains the most rigorous GO annotations based solely on experimental data. The 'Community' gene association file contains GO annotations based on expert community knowledge (annotations based directly from author statements and submitted annotations from the community) and annotations for predicted proteins. We have developed two tools for proteomics analysis and these are freely available on request. A suite of tools for analyzing functional genomics datasets using the GO is available online at the AgBase site. We encourage and publicly acknowledge GO annotations from researchers and provide an online mechanism for agricultural researchers to submit requests for GO annotations.

Agriculture↗

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44 Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92 Mb and 29.61 Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals↗

Protein complexes and functional modules in molecular networks.

Proteins, nucleic acids, and small molecules form a dense network of molecular interactions in a cell. Molecules are nodes of this network, and the interactions between them are edges. The architecture of molecular networks can reveal important principles of cellular organization and function, similarly to the way that protein structure tells us about the function and organization of a protein. Computational analysis of molecular networks has been primarily concerned with node degree [Wagner, A. & Fell, D. A. (2001) Proc. R. Soc. London Ser. B 268, 1803-1810; Jeong, H., Tombor, B., Albert, R., Oltvai, Z. N. & Barabasi, A. L. (2000) Nature 407, 651-654] or degree correlation [Maslov, S. & Sneppen, K. (2002) Science 296, 910-913], and hence focused on single/two-body properties of these networks. Here, by analyzing the multibody structure of the network of protein-protein interactions, we discovered molecular modules that are densely connected within themselves but sparsely connected with the rest of the network. Comparison with experimental data and functional annotation of genes showed two types of modules: (i) protein complexes (splicing machinery, transcription factors, etc.) and (ii) dynamic functional units (signaling cascades, cell-cycle regulation, etc.). Discovered modules are highly statistically significant, as is evident from comparison with random graphs, and are robust to noise in the data. Our results provide strong support for the network modularity principle introduced by Hartwell et al. [Hartwell, L. H., Hopfield, J. J., Leibler, S. & Murray, A. W. (1999) Nature 402, C47-C52], suggesting that found modules constitute the "building blocks" of molecular networks.

Biophysical Phenomena↗

Applications of InterPro in protein annotation and genome analysis.

The applications of InterPro span a range of biologically important areas that includes automatic annotation of protein sequences and genome analysis. In automatic annotation of protein sequences InterPro has been utilised to provide reliable characterisation of sequences, identifying them as candidates for functional annotation. Rules based on the InterPro characterisation are stored and operated through a database called RuleBase. RuleBase is used as the main tool in the sequence database group at the EBI to apply automatic annotation to unknown sequences. The annotated sequences are stored and distributed in the TrEMBL protein sequence database. InterPro also provides a means to carry out statistical and comparative analyses of whole genomes. In the Proteome Analysis Database, InterPro analyses have been combined with other analyses based on CluSTr, the Gene Ontology (GO) and structural information on the proteins.

Amino Acid Sequence↗

Graph-based analysis and visualization of experimental results with ONDEX.

MOTIVATION: Assembling the relevant information needed to interpret the output from high-throughput, genome scale, experiments such as gene expression microarrays is challenging. Analysis reveals genes that show statistically significant changes in expression levels, but more information is needed to determine their biological relevance. The challenge is to bring these genes together with biological information distributed across hundreds of databases or buried in the scientific literature (millions of articles). Software tools are needed to automate this task which at present is labor-intensive and requires considerable informatics and biological expertise. RESULTS: This article describes ONDEX and how it can be applied to the task of interpreting gene expression results. ONDEX is a database system that combines the features of semantic database integration and text mining with methods for graph-based analysis. An overview of the ONDEX system is presented, concentrating on recently developed features for graph-based analysis and visualization. A case study is used to show how ONDEX can help to identify causal relationships between stress response genes and metabolic pathways from gene expression data. ONDEX also discovered functional annotations for most of the genes that emerged as significant in the microarray experiment, but were previously of unknown function.

Algorithms↗

Analysis of multiple genomic sequence alignments: a web resource, online tools, and lessons learned from analysis of mammalian SCL loci.

Comparative analysis of genomic sequences is becoming a standard technique for studying gene regulation. However, only a limited number of tools are currently available for the analysis of multiple genomic sequences. An extensive data set for the testing and training of such tools is provided by the SCL gene locus. Here we have expanded the data set to eight vertebrate species by sequencing the dog SCL locus and by annotating the dog and rat SCL loci. To provide a resource for the bioinformatics community, all SCL sequences and functional annotations, comprising a collation of the extensive experimental evidence pertaining to SCL regulation, have been made available via a Web server. A Web interface to new tools specifically designed for the display and analysis of multiple sequence alignments was also implemented. The unique SCL data set and new sequence comparison tools allowed us to perform a rigorous examination of the true benefits of multiple sequence comparisons. We demonstrate that multiple sequence alignments are, overall, superior to pairwise alignments for identification of mammalian regulatory regions. In the search for individual transcription factor binding sites, multiple alignments markedly increase the signal-to-noise ratio compared to pairwise alignments.

Animals↗

Finding functional features in Saccharomyces genomes by phylogenetic footprinting.

The sifting and winnowing of DNA sequence that occur during evolution cause nonfunctional sequences to diverge, leaving phylogenetic footprints of functional sequence elements in comparisons of genome sequences. We searched for such footprints among the genome sequences of six Saccharomyces species and identified potentially functional sequences. Comparison of these sequences allowed us to revise the catalog of yeast genes and identify sequence motifs that may be targets of transcriptional regulatory proteins. Some of these conserved sequence motifs reside upstream of genes with similar functional annotations or similar expression patterns or those bound by the same transcription factor and are thus good candidates for functional regulatory sequences.

Algorithms↗

Systematic discovery of retina-enriched Rik genes identifies 1190005I06Rik as a novel modulator of visual signalling.

BACKGROUND: High‑throughput transcriptome projects have revealed thousands of mammalian genes with little or no functional annotation. Among these are hundreds of loci assigned provisional “Rik” identifiers following discovery in the RIKEN cDNA annotation effort. Although often dismissed as genomic dark matter, such genes may encode tissue‑restricted proteins that modulate physiologic functions and influence disease. The retina is a highly specialised neural tissue and a common site of inherited disorders; understanding its molecular repertoire could illuminate novel therapeutic avenues. METHODS: We integrated bulk RNA‑seq from ten adult mouse tissues, evolutionary and domain analysis, single‑cell RNA‑seq, and CRISPR/Cas9 gene disruption to systematically catalogue protein‑coding Rik genes enriched in the retina and test the function of a representative gene. RESULTS: A rigorous differential expression analysis identified 44 Rik genes with robust retina‑specific expression compared with nine non‑retinal tissues. Many of these genes lack orthologues beyond rodents, while others show broad conservation, illustrating a continuum from lineage‑restricted to conserved retinopathy candidates. Single‑cell transcriptomics revealed that these genes are expressed across retinal cell types, with the highest aggregate expression in cone photoreceptors and inner interneurons. To evaluate physiological significance, we generated a 1190005I06Rik knockout mouse. Although retinal architecture appeared normal, loss of 1190005I06Rik enhanced electroretinogram b‑wave amplitudes and altered light‑avoidance behaviour, indicating that this previously uncharacterised gene acts as a negative modulator of visual signalling. CONCLUSIONS: We present a curated atlas of retina‑enriched Rik genes and demonstrate that 1190005I06RIK modulates retinal circuit function. This resource expands the molecular landscape of the retina and provides new candidates for the genetic basis of inherited retinal disease. Our findings underscore that unannotated genes may exert measurable effects on sensory processing and warrant systematic exploration in the context of human ocular disorders.

Animals↗

Predicting pathway perturbations in Down syndrome.

Comparative annotation of human chromosome 21 genomic sequence with homologous regions of mouse chromosomes 16, 17 and 10 has identified 170 orthologous gene pairs. Functional annotation of these genes, based on literature reports and computationally-derived predictions, shows that a broad range of cellular processes are represented. A goal of Down syndrome research is to determine which of these processes are perturbed by overexpression of chromosome 21 genes, and which may, therefore, contribute to the cognitive deficits that characterize Down syndrome. Eleven chromosome 21 genes are annotated to interact with or be affected by components of the MAP Kinase pathway and eight are involved in Ca2+/calcineurin signaling. Both pathways are critical for normal neurological function, and consequently their perturbations are proposed as candidates for phenotypic relevance. We present evidence suggesting that the MAP Kinase pathway is perturbed in the Ts65Dn mouse model of Down syndrome at 4-6 months of age. Analysis is complicated by the observation that overexpression of chromosome 21 genes in trisomy may be affected by method of detection, organism, tissue or brain region, and/or developmental age.

Animals↗

VMD: a community annotation database for oomycetes and microbial genomes.

The VBI Microbial Database (VMD) is a database system designed to host a range of microbial genome sequences. At present, the database contains genome sequence and annotation data of two plant pathogens Phytophthora sojae and Phytophthora ramorum. With the completion of the draft genome sequences of these pathogens in collaboration with the DOE Joint Genome Institute (JGI), we have created this resource to make the sequences publicly available. The genome sequences (95 MB for P.sojae and 65 MB for P.ramorum) were annotated with approximately 19,000 and approximately 16,000 gene models, respectively. We used two different statistical methods to validate these gene models, Fickett's and a log-likelihood method. Functional annotation of the gene models is based on results from BlastX and InterProScan screens. From the InterProScan results, we could assign putative functions to 17,694 genes in P.sojae and 14,700 genes in P.ramorum. We created an easy-to-use genome browser to view the genome sequence data, which opens to detailed annotation pages for each gene model. A community annotation interface is available for registered community members to add or edit annotations. There are approximately 1600 gene models for P.sojae and approximately 700 models for P.ramorum that have already been manually curated. A toolkit is provided as an additional resource for users to perform a variety of sequence analysis jobs. The database is publicly available at http://phytophthora.vbi.vt.edu/.

Databases, Nucleic Acid↗

Confirmation of the expression of a large set of conserved hypothetical proteins in Shewanella oneidensis MR-1.

High-throughput "omic" technologies have allowed for a relatively rapid, yet comprehensive analysis of the global expression patterns within an organism in response to perturbations. In the current study, 9503 different tryptic peptides were identified with high confidence from capillary liquid chromatography-mass spectrometry analysis of 26 chemostat cultures of Shewanella oneidensis MR-1 under various conditions. Using at least one distinctive and a total of two total peptide identifications per protein, we detected the expression of 758 conserved hypothetical proteins. This included 359 such proteins previously described [Kolker, E., Picone, A.F., Galperin, M.Y., Romine, M.F., Higdon, R., Makarova, K.S., Kolker, N., Anderson, G.A., Qiu, X., Auberry, K.J., Babnigg, G., Beliaev, A.S., Edlefsen, P., Elias, D.A., Gorby, Y.A., Holzman, T., Klappenbach, J.A., Konstantinidis, K.T., Land, M.L., Lipton, M.S., McCue, L.A., Monroe, M., Pasa-Tolic, L., Pinchuk, G., Purvine, S., Serres, M.H., Tsapin, S., Zakrajsek, B.A., Zhu, W., Zhou, J., Larimer, F.W., Lawrence, C.E., Riley, M., Collart, F.R., Yates, J.R., III, Smith, R.D., Giometti, C.S., Nealson, K.H., Fredrickson, J.K., Tiedje, J.M., 2005. Global profiling of Shewanella oneidensis MR-1: expression of hypothetical genes and improved functional annotations. Proc Natl Acad Sci U S A 102, 2099-2104] with an additional 399 reported herein for the first time. The latter 399 proteins ranged from 5.3 to 208.3 kDa, with 44 being of 100 amino acid residues or less. Using a combination of information including peptide detection in cells grown under specific culture conditions and predictive algorithms such as PSORT and PSORT-B, possible/plausible functions are proposed for some conserved hypothetical proteins. Such proteins were found not only to be expressed, but 19 were only expressed under certain culturing conditions, thereby providing insight into potential functions. These findings also impact the genomic annotation for S. oneidensis MR-1 by confirming that these genes code for expressed proteins. Our results indicate that 399 proteins can now be upgraded from "conserved hypothetical protein" to "expressed protein in Shewanella," 19 of which appeared to be expressed under specific culture conditions.

Bacterial Proteins↗

DeepES: deep learning-based enzyme screening to identify orphan enzyme genes.

MOTIVATION: Progress in sequencing technology has led to determination of large numbers of protein sequences, and large enzyme databases are now available. Although many computational tools for enzyme annotation were developed, sequence information is unavailable for many enzymes, known as orphan enzymes. These orphan enzymes hinder sequence similarity-based functional annotation, leading gaps in understanding the association between sequences and enzymatic reactions. RESULTS: Therefore, we developed DeepES, a deep learning-based tool for enzyme screening to identify orphan enzyme genes, focusing on biosynthetic gene clusters and reaction class. DeepES uses protein sequences as inputs and evaluates whether the input genes contain biosynthetic gene clusters of interest by integrating the outputs of the binary classifier for each reaction class. The validation results suggested that DeepES can capture functional similarity between protein sequences, and it can be implemented to explore orphan enzyme genes. By applying DeepES to 4744 metagenome-assembled genomes, we identified candidate genes for 236 orphan enzymes, including those involved in short-chain fatty acid production as a characteristic pathway in human gut bacteria. AVAILABILITY AND IMPLEMENTATION: DeepES is available at https://github.com/yamada-lab/DeepES. Model weights and the candidate genes are available at Zenodo (https://doi.org/10.5281/zenodo.11123900).

Deep Learning↗

Proteome-wide functional classification and identification of prokaryotic transmembrane proteins by transmembrane topology similarity comparison.

We propose a new method for classifying and identifying transmembrane (TM) protein functions in proteome-scale by applying a single-linkage clustering method based on TM topology similarity, which is calculated simply from comparing the lengths of loop regions. In this study, we focused on 87 prokaryotic TM proteomes consisting of 31 proteobacteria, 22 gram-positive bacteria, 19 other bacteria, and 15 archaea. Prior to performing the clustering, we first categorized individual TM protein sequences as "known," "putative" (similar to "known" sequences), or "unknown" by using the homology search and the sequence similarity comparison against SWISS-PROT to assess the current status of the functional annotation of the TM proteomes based on sequence similarity only. More than three-quarters, that is, 75.7% of the TM protein sequences are functionally "unknown," with only 3.8% and 20.5% of them being classified as "known" and "putative," respectively. Using our clustering approach based on TM topology similarity, we succeeded in increasing the rate of TM protein sequences functionally classified and identified from 24.3% to 60.9%. Obtained clusters correspond well to functional superfamilies or families, and the functional classification and identification are successfully achieved by this approach. For example, in an obtained cluster of TM proteins with six TM segments, 109 sequences out of 119 sequences annotated as "ATP-binding cassette transporter" are properly included and 122 "unknown" sequences are also contained.

Algorithms↗

Systems for the detection and analysis of protein-protein interactions.

The analysis of protein-protein interactions is important for developing a better understanding of the functional annotations of proteins that are involved in various biochemical reactions in vivo. The discovery that a protein with an unknown function binds to a protein with a known function could provide a significant clue to the cellular pathway concerning the unknown protein. Therefore, information on protein-protein interactions obtained by the comprehensive analysis of all gene products is available for the construction of interactive networks consisting of individual protein-protein interactions, which, in turn, permit elaborate biological phenomena to be understood. Systems for detecting protein-protein interactions in vitro and in vivo have been developed, and have been modified to compensate for limitations. Using these novel approaches, comprehensive and reliable information on protein-protein interactions can be determined. Systems that permit this to be achieved are described in this review.

Chromatography, Affinity↗