PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Functional fingerprints of folds: evidence for correlated structure-function evolution.

Using structural similarity clustering of protein domains: protein domain universe graph (PDUG), and a hierarchical functional annotation: gene ontology (GO) as two evolutionary lenses, we find that each structural cluster (domain fold) exhibits a distribution of functions that is unique to it. These functional distributions are functional fingerprints that are specific to characteristic structural clusters and vary from cluster to cluster. Furthermore, as structural similarity threshold for domain clustering in the PDUG is relaxed we observe an influx of earlier-diverged domains into clusters. These domains join clusters without destroying the functional fingerprint. These results can be understood in light of a divergent evolution scenario that posits correlated divergence of structural and functional traits in protein domains from one or few progenitors.

Adenosine Triphosphate↗

VIRGO: computational prediction of gene functions.

Dramatic advances in sequencing technology and sophisticated experimental assays that interrogate the cell, combined with the public availability of the resulting data, herald the era of systems biology. However, the biological functions of more than 40% of the genes in sequenced genomes are unknown, posing a fundamental barrier to progress in systems biology. The large scale and diversity of available data requires the development of techniques that can automatically utilize these datasets to make quantified and robust predictions of gene function that can be experimentally verified. We present a service called the VIRtual Gene Ontology (VIRGO) that (i) constructs a functional linkage network (FLN) from gene expression and molecular interaction data, (ii) labels genes in the FLN with their functional annotations in the Gene Ontology and (iii) systematically propagates these labels across the FLN in order to precisely predict the functions of unlabelled genes. VIRGO assigns confidence estimates to predicted functions so that a biologist can prioritize predictions for further experimental study. For each prediction, VIRGO also provides an informative 'propagation diagram' that traces the flow of information in the FLN that led to the prediction. VIRGO is available at http://whipple.cs.vt.edu:8080/virgo.

Algorithms↗

Open-source toolkit for simple XML annotation.

Use of Extensible Markup Language (XML) is increasingly prevalent among medical informatics projects. Many of these projects involve, at some point, the interaction between a researcher and specialized XML documents for the purpose of annotating the XML data. We offer a simple toolkit to assist these researchers. Our solution is a simple, yet fully functional, annotation system that can easily be adapted to the needs of the researcher. All of the materials for this toolkit are freely available.

Algorithms↗

Functional grouping based on signatures in protein termini.

The two ends of each protein are known as the amino (N-) and carboxyl (C-) termini. Short signatures in a protein's termini often carry vital cellular function. No systematic research has been conducted to address the importance of short signatures (3 to 10 amino acids) in protein termini at the proteomic level. Specifically, it is unknown whether such signatures are evolutionarily conserved, and if so, whether this conservation confers shared biological functions. Current signature detection methods fail to detect such short signatures due to inadequate statistical scores. The findings presented in this study strongly support the notion that functional significance of protein sets may be captured by short signatures at their termini. A positional search method was applied to over one million proteins from the UniProt database. The result is a collection of about a thousand significant signature groups (SIGs) that include previously identified as well as many novel signatures in protein termini. These SIGs represent protein sets with minimal or no overall sequence similarity excepting the similarity at their termini. The most significant SIGs are assigned by their strong correspondence to functional annotations derived from external databases such as Gene Ontology. Each of the SIGs is associated with the statistical significance of its functional association. These SIGs provide a valuable source for testing previously overlooked signatures in protein termini and allow for the investigation of the role played by such signatures throughout evolution. The SIGs archive and advanced search options are available at http://www.proteus.cs.huji.ac.il.

Amino Acid Sequence↗

Inferring functional relationships of proteins from local sequence and spatial surface patterns.

We describe a novel approach for inferring functional relationship of proteins by detecting sequence and spatial patterns of protein surfaces. Well-formed concave surface regions in the form of pockets and voids are examined to identify similarity relationship that might be directly related to protein function. We first exhaustively identify and measure analytically all 910,379 surface pockets and interior voids on 12,177 protein structures from the Protein Data Bank. The similarity of patterns of residues forming pockets and voids are then assessed in sequence, in spatial arrangement, and in orientational arrangement. Statistical significance in the form of E and p-values is then estimated for each of the three types of similarity measurements. Our method is fully automated without human intervention and can be used without input of query patterns. It does not assume any prior knowledge of functional residues of a protein, and can detect similarity based on surface patterns small and large. It also tolerates, to some extent, conformational flexibility of functional sites. We show with examples that this method can detect functional relationship with specificity for members of the same protein family and superfamily, as well as remotely related functional surfaces from proteins of different fold structures. We envision that this method can be used for discovering novel functional relationship of protein surfaces, for functional annotation of protein structures with unknown biological roles, and for further inquiries on evolutionary origins of structural elements important for protein function.

Amino Acid Sequence↗

A network-based analysis of allergen-challenged CD4+ T cells from patients with allergic rhinitis.

We performed a network-based analysis of DNA microarray data from allergen-challenged CD4(+) T cells from patients with seasonal allergic rhinitis. Differentially expressed genes were organized into a functionally annotated network using the Ingenuity Knowledge Database, which is based on manual review of more than 200,000 publications. The main function of this network is the regulation of lymphocyte apoptosis, a role associated with several genes of the tuber necrosis factor superfamily. The expression of TNFRSF4, one of the genes in this family, was found to be 48 times higher in allergen-challenged cells than in diluent-challenged cells. TNFRSF4 is known to inhibit apoptosis and to enhance Th2 proliferation. Examination of a different material of allergen-stimulated peripheral blood mononuclear cells showed a higher number of interleukin-4(+) type 2 CD4(+) T (Th2) cells in patients than in controls (P<0.01), as well as a higher number of non-apoptotic Th2 cells in patients (P<0.01). The number of Th2 cells expressing TNFRSF4, TNFSF7 and TNFRSF1B was also significantly higher in patients. Treatment with anti-TNFSF4 resulted in a significantly decreased number of Th2 cells (P<0.05). A logical inference from all this is that the proliferation of allergen-challenged Th2 cells is associated with a decreased apoptosis of Th2 cells and an increase in TNFRSF4 signalling.

Adolescent↗

Global molecular and morphological effects of 24-hour chromium(VI) exposure on Shewanella oneidensis MR-1.

The biological impact of 24-h ("chronic") chromium(VI) [Cr(VI) or chromate] exposure on Shewanella oneidensis MR-1 was assessed by analyzing cellular morphology as well as genome-wide differential gene and protein expression profiles. Cells challenged aerobically with an initial chromate concentration of 0.3 mM in complex growth medium were compared to untreated control cells grown in the absence of chromate. At the 24-h time point at which cells were harvested for transcriptome and proteome analyses, no residual Cr(VI) was detected in the culture supernatant, thus suggesting the complete uptake and/or reduction of this metal by cells. In contrast to the untreated control cells, Cr(VI)-exposed cells formed apparently aseptate, nonmotile filaments that tended to aggregate. Transcriptome profiling and mass spectrometry-based proteomic characterization revealed that the principal molecular response to 24-h Cr(VI) exposure was the induction of prophage-related genes and their encoded products as well as a number of functionally undefined hypothetical genes that were located within the integrated phage regions of the MR-1 genome. In addition, genes with annotated functions in DNA metabolism, cell division, biosynthesis and degradation of the murein (peptidoglycan) sacculus, membrane response, and general environmental stress protection were upregulated, while genes encoding chemotaxis, motility, and transport/binding proteins were largely repressed under conditions of 24-h chromate treatment.

Bacterial Proteins↗

Expressed sequence tags (ESTs) and simple sequence repeat (SSR) markers from octoploid strawberry (Fragaria x ananassa).

BACKGROUND: Cultivated strawberry (Fragaria x ananassa) represents one of the most valued fruit crops in the United States. Despite its economic importance, the octoploid genome presents a formidable barrier to efficient study of genome structure and molecular mechanisms that underlie agriculturally-relevant traits. Many potentially fruitful research avenues, especially large-scale gene expression surveys and development of molecular genetic markers have been limited by a lack of sequence information in public databases. As a first step to remedy this discrepancy a cDNA library has been developed from salicylate-treated, whole-plant tissues and over 1800 expressed sequence tags (EST's) have been sequenced and analyzed. RESULTS: A putative unigene set of 1304 sequences--133 contigs and 1171 singlets--has been developed, and the transcripts have been functionally annotated. Homology searches indicate that 89.5% of sequences share significant similarity to known/putative proteins or Rosaceae ESTs. The ESTs have been functionally characterized and genes relevant to specific physiological processes of economic importance have been identified. A set of tools useful for SSR development and mapping is presented. CONCLUSION: Sequences derived from this effort may be used to speed gene discovery efforts in Fragaria and the Rosaceae in general and also open avenues of comparative mapping. This report represents a first step in expanding molecular-genetic analyses in strawberry and demonstrates how computational tools can be used to optimally mine a large body of useful information from a relatively small data set.

Chromosome Mapping↗

Pscroph, a parasitic plant EST database enriched for parasite associated transcripts.

BACKGROUND: Parasitic plants in the Orobanchaceae develop invasive root haustoria upon contact with host roots or root factors. The development of haustoria can be visually monitored and is rapid, highly synchronous, and strongly dependent on host factor exposure; therefore it provides a tractable system for studying chemical communications between roots of different plants. DESCRIPTION: Triphysaria is a facultative parasitic plant that initiates haustorium development within minutes after contact with host plant roots, root exudates, or purified haustorium-inducing phenolics. In order to identify genes associated with host root identification and early haustorium development, we sequenced suppression subtractive libraries (SSH) enriched for transcripts regulated in Triphysaria roots within five hours of exposure to Arabidopsis roots or the purified haustorium-inducing factor 2,6 dimethoxybenzoquinone. The sequences of over nine thousand ESTs from three SSH libraries and their subsequent assemblies are available at the Pscroph database http://pscroph.ucdavis.edu. The web site also provides BLAST functions and allows keyword searches of functional annotations. CONCLUSION: Libraries prepared from Triphysaria roots treated with host roots or haustorium inducing factors were enriched for transcripts predicted to function in stress responses, electron transport or protein metabolism. In addition to parasitic plant investigations, the Pscroph database provides a useful resource for investigations in rhizosphere interactions, chemical signaling between organisms, and plant development and evolution.

Arabidopsis↗

Effect of ionizing irradiation on human esophageal cancer cell lines by cDNA microarray gene expression analysis.

To provide new insights into the molecular mechanisms underlying the effect of irradiation on esophageal squamous cell carcinomas (ESCCs), we used a cDNA microarray screening of more than 4,000 genes with known functions to identify genes involved in the early response to ionizing irradiation. Two human ESCC cell lines, one each of well (TE-1) and poorly (TE-2) differentiated phenotypes were screened. Subconfluent cells of each phenotype were treated with single doses of 2.0 Gy or 8.0 Gy irradiations. After a 15 min incubation time-point, the cells were collected and analyzed. Compared with non-irradiated cells, many genes revealed at least 2-fold upregulation or downregulation at both doses in well or poorly differentiated ESCC cells. The common upregulated genes in well and poorly differentiated cell types at both irradiation doses included SCYA5, CYP51, SMARCD2, COX6C, MAPK8, FOS, UBE2M, RPL6, PDGFRL, TRAF2, TNFAIP6, ITGB4, GSTM3, and SP3 and common downregulated genes involved NFIL3, SMARCA2, CAPZA1, MetAP2, CITED2, DAP3, MGAT2, ATRX, CIAO1, and STAT6. Several of these genes were novel and not previously known to be associated with irradiation. Functional annotations of the modulated genes suggested that at the molecular level, irradiation appears to induce a regularizing balance in ESCC cell function. The genes modulated in the early response to irradiation may be useful in our understanding of the molecular basis of radiotherapy and in developing strategies to augment its effect or establish novel less hazardous alternative adjuvant therapies.

Carcinoma, Squamous Cell↗

RNAs everywhere: genome-wide annotation of structured RNAs.

Starting with the discovery of microRNAs and the advent of genome-wide transcriptomics, non-protein-coding transcripts have moved from a fringe topic to a central field research in molecular biology. In this contribution we review the state of the art of "computational RNomics", i.e., the bioinformatics approaches to genome-wide RNA annotation. Instead of rehashing results from recently published surveys in detail, we focus here on the open problem in the field, namely (functional) annotation of the plethora of putative RNAs. A series of exploratory studies are used to provide non-trivial examples for the discussion of some of the difficulties.

Biological Evolution↗

The current and future perspective of ChickenGTEx project and its applications in precision breeding.

The Chicken Genotype-Tissue Expression (ChickenGTEx) project was established to systematically characterize the regulatory landscape of the chicken genome and to accelerate the translation of functional genomics into precision breeding. By integrating whole-genome sequencing with multi-tissue transcriptomic profiling, ChickenGTEx provides a comprehensive atlas of gene expression regulation across diverse tissues and physiological systems. Current findings demonstrate that complex production traits are governed by coordinated regulatory networks rather than isolated loci, with substantial contributions from tissue-specific gene expression, structural variation, and genotype-by-sex interactions. Sex-dependent regulatory effects further refine the genetic architecture of metabolic, immune, and reproductive traits, highlighting the importance of incorporating sex as a biological variable in genomic analyses. Application of integrative omics frameworks within elite layer populations has revealed multilayer regulatory mechanisms underlying extended laying performance, feed efficiency, metabolic health, and eggshell quality. By partitioning phenotypic variance into genetic, regulatory, and host-microbiome components, these approaches move beyond association-based mapping toward causal inference and biological interpretation. Importantly, validated regulatory loci identified through ChickenGTEx and related analyses provide actionable markers for genomic selection and rational targets for precision genome modification. Looking forward, continued expansion of regulatory atlases, incorporation of single-cell and longitudinal data in diverse environmental conditions, and integration of functional annotation into breeding pipelines will further enhance prediction accuracy and sustainable genetic improvement. The ChickenGTEx project thus represents a foundational platform bridging functional genomics and practical poultry breeding.

Animals↗

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32&#x2009;Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59&#x2009;Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant↗

Accounting for recombination rate variation improves inference of barrier loci and reveals the role of both natural and sexual selection in an incipient bird radiation.

Examining genomic patterns of differentiation across lineage pairs at different stages of the speciation continuum, in combination with recombination maps, can help disentangle the effects of linked and divergent selection and identify lineage-specific targets of selection that may act as barrier loci during speciation. Here, we apply this framework to genomic data from African and Indian Ocean bird species of the genus Zosterops (Zosteropidae) to identify candidate barrier loci between ecologically, phenotypically, and genetically distinct Reunion gray white-eye (Zosterops borbonicus) parapatric geographic forms. Using analyses that account for recombination rate variation, we show that putative targets of divergent selection are primarily located on the Z chromosome, except in comparisons between geographic forms that differ in their ecologies. Functional annotation revealed that candidate barrier loci between forms with similar environmental niches are associated with genes involved in song formation and immune function, whereas those between forms with different environmental niches are associated with adaptation to altitude, morphology, and song behavior. Our results highlight the combined roles of natural and sexual selection in the evolution of reproductive barriers in this incipient species radiation.

Animals↗

Identification of novel transcriptional networks in response to treatment with the anticarcinogen 3H-1,2-dithiole-3-thione.

3H-1,2-dithiole-3-thione (D3T), an inducer of antioxidant and phase 2 genes, is known to enhance the detoxification of environmental carcinogens, prevent neoplasia, and elicit other protective effects. However, a comprehensive view of the regulatory pathways induced by this compound has not yet been elaborated. Fischer F344 rats were gavaged daily for 5 days with vehicle or D3T (0.3 mmol/kg). The global changes of gene expression in liver were measured with Affymetrix RG-U34A chips. With the use of functional class scoring, a semi-supervised method exploring both the expression pattern and the functional annotation of the genes, the Gene Ontology classes were ranked according to the significance of the impact of D3T treatment. Two unexpected functional classes were identified for the D3T treatment, cytosolic ribosome constituents with 90% of those genes increased, and cholesterol biosynthesis with 91% of the genes repressed. In another novel approach, the differentially expressed genes were evaluated by the Ingenuity computational pathway analysis tool to identify specific regulatory networks and canonical pathways responsive to D3T treatment. In addition to the known glutathione metabolism pathway (P = 0.0011), several other significant pathways were also revealed, including antigen presentation (P = 0.000476), androgen/estrogen biosynthesis (P = 0.000551), fatty acid (P = 0.000216), and tryptophan metabolism (P = 0.000331) pathways. These findings showed a profound impact of D3T on lipid metabolism and anti-inflammatory/immune-suppressive response, indicating a broader cytoprotective effect of this compound than previously expected.

Animals↗

Sequence- and structure-based protein function prediction from genomic information.

Existing functional annotation transfer is fraught with inaccuracies that may hinder forward interpretation and mining of genomic data. Hand-curation of the annotation placed into databases is not practical. In lieu of experimental evidence, computational biological approaches offer high-throughput tools to predict function accurately; however, these methods are still notably deficient in defining and describing the complexity of protein function. Enriching genomic sequences obtained from sequencing efforts and expression array methods with protein function information and classification will be an efficient first step for incorporating genomic data into drug discovery programs.

Computational Biology↗

Update on genome completion and annotations: Protein Information Resource.

The Protein Information Resource (PIR) recently joined the European Bioinformatics Institute (EBI) and Swiss Institute of Bioinformatics (SIB) to establish UniProt--the Universal Protein Resource--which now unifies the PIR, Swiss-Prot and TrEMBL databases. The PIRSF (SuperFamily) classification system is central to the PIR/UniProt functional annotation of proteins, providing classifications of whole proteins into a network structure to reflect their evolutionary relationships. Data integration and associative studies of protein family, function and structure are supported by the iProClass database, which offers value-added descriptions of all UniProt proteins with highly informative links to more than 50 other databases. The PIR system allows consistent, rich and accurate protein annotation for all investigators.

Animals↗

POCUS: mining genomic sequence annotation to predict disease genes.

Here we present POCUS (prioritization of candidate genes using statistics), a novel computational approach to prioritize candidate disease genes that is based on over-representation of functional annotation between loci for the same disease. We show that POCUS can provide high (up to 81-fold) enrichment of real disease genes in the candidate-gene shortlists it produces compared with the original large sets of positional candidates. In contrast to existing methods, POCUS can also suggest counterintuitive candidates.

Autistic Disorder↗