PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Using the CATH domain database to assign structures and functions to the genome sequences.

The CATH database of protein structures contains approximately 18000 domains organized according to their (C)lass, (A)rchitecture, (T)opology and (H)omologous superfamily. Relationships between evolutionary related structures (homologues) within the database have been used to test the sensitivity of various sequence search methods in order to identify relatives in Genbank and other sequence databases. Subsequent application of the most sensitive and efficient algorithms, gapped blast and the profile based method, Position Specific Iterated Basic Local Alignment Tool (PSI-BLAST), could be used to assign structural data to between 22 and 36 % of microbial genomes in order to improve functional annotation and enhance understanding of biological mechanism. However, on a cautionary note, an analysis of functional conservation within fold groups and homologous superfamilies in the CATH database, revealed that whilst function was conserved in nearly 55% of enzyme families, function had diverged considerably, in some highly populated families. In these families, functional properties should be inherited far more cautiously and the probable effects of substitutions in key functional residues carefully assessed.

Algorithms↗

Preferential duplication in the sparse part of yeast protein interaction network.

Gene duplication is an important mechanism driving the evolution of biomolecular network. Thus, it is expected that there should be a strong relationship between a gene's duplicability and the interactions of its protein product with other proteins in the network. We studied this question in the context of the protein interaction network (PIN) of Saccharomyces cerevisiae. We found that duplicates have, on average, significantly lower clustering coefficient (CC) than singletons, and the proportion of duplicates (PD) decreases steadily with CC. Furthermore, using functional annotation data, we observed a strong negative correlation between PD and the mean CC for functional categories. By partitioning the network into modules and assigning each protein a modularity measure Q(n), we found that CC of a protein is a reflection of its modularity. Moreover, the core components of complexes identified in a recent high-throughput experiment, characterized by high CC, have lower PD than that of the attachments. Subsequently, 2 types of hub were identified by their degree, CC and Q(n). Although PD of intramodular hubs is much less than the network average, PD of intermodular hubs is comparable to, or even higher than, the network average. Our results suggest that high CC, and thus high modularity, pose strong evolutionary constraints on gene duplicability, and gene duplication prefers to happen in the sparse part of PINs.

Cluster Analysis↗

Design and characterization of a functional library for NMR screening against novel protein targets.

In the past few years, NMR has been extensively utilized as a screening tool for drug discovery using various types of compound libraries. The designs of NMR specific chemical libraries that utilize a fragment-based approach based on drug-like characteristics have been previously reported. In this article, a new type of compound library will be described that focuses on aiding in the functional annotation of novel proteins that have been identified from various ongoing genomics efforts. The NMR functional chemical library is comprised of small molecules with known biological activity such as: co-factors, inhibitors, metabolites and substrates. This functional library was developed through an extensive manual effort of mining several databases based on known ligand interactions with protein systems. In order to increase the efficiency of screening the NMR functional library, the compounds are screened as mixtures of 3-4 compounds that avoids the need to deconvolute positive hits by maintaining a unique NMR resonance and function for each compound in the mixture. The functional library has been used in the identification of general biological function of hypothetical proteins identified from the Protein Structure Initiative.

Binding Sites↗

[DNA arrays: technological aspects and applications].

The Human Genome Project has allowed considerable progress in the construction of physical and genetic maps and the identification of genes involved in human sicknesses. The accelerated accumulation of biological information and knowledge is due in large part to the sequencing projects of other organisms, which in fact paved the way for the Human Genome Project. In parallel, recently developed techniques which take advantage of genomic sequences allow large scale molecular analyses resulting in the functional annotation of many of the proteins represented by these genes. This is the goal of functional genomics. These progresses are at the origin of the present revolution in biomedical research. DNA microarrays are playing a dominant role compared to the other developing technologies since they are relatively easy to make and use and are applicable to numerous scientific inquiries. They allow the simultaneous analysis of several thousands of genes in biological samples from sick or healthy tissues, at the genome or transcriptome level. The data obtained is expected to result in major advances in the health sciences. In addition to an improved understanding of the complex molecular interaction networks of healthy cells and tissues, a more precise genetic characterization of the molecular mechanisms involved in pathology should result in the identification of new therapeutic targets and the development of new medicines. The genetic profiles thus obtained should also permit the definition of new pathologic subclasses not recognizable by traditional clinical factors, as well as new markers for susceptibility to certain illnesses, and new prognostic markers or methods of predicting responses to treatment. In this article, we present the different approaches and potential applications of DNA microarray technology, in particular as applied to cancer research.

Chromosome Mapping↗

Development of FuGO: an ontology for functional genomics investigations.

The development of the Functional Genomics Investigation Ontology (FuGO) is a collaborative, international effort that will provide a resource for annotating functional genomics investigations, including the study design, protocols and instrumentation used, the data generated and the types of analysis performed on the data. FuGO will contain both terms that are universal to all functional genomics investigations and those that are domain specific. In this way, the ontology will serve as the "semantic glue" to provide a common understanding of data from across these disparate data sources. In addition, FuGO will reference out to existing mature ontologies to avoid the need to duplicate these resources, and will do so in such a way as to enable their ease of use in annotation. This project is in the early stages of development; the paper will describe efforts to initiate the project, the scope and organization of the project, the work accomplished to date, and the challenges encountered, as well as future plans.

Biomedical Research↗

Integrated analysis of the genome and the transcriptome by FANTOM.

The key to reliable annotation of a mammalian genome is broad characterisation of the transcriptional output, the transcriptome. FANTOM, the functional annotation of mouse cDNA, is a large-scale analysis of both the genome and the transcriptome of the mouse. In the early days of this work, the transcripts were characterised using our sophisticated methods. After the timely release of the first draft of mouse genome sequences, interesting information was obtained by its integration with these one-by-one annotations. Moreover, each transcript included its expression profile. Here, the two integrated annotation methods used by FANTOM are reviewed: one-by-one and categorised. One-by-one annotation refers to naming carried out based on well-known transcripts or its fragments using the top-down-style pipeline developed mostly by the FANTOM project. Categorised annotation, which refers to transcript grouping, not only helps naming of unknown transcripts, but will be the most utilised method for integration of the genome and the transcriptome from now on.

Abstracting and Indexing↗

Structural characterization of the human proteome.

This paper reports an analysis of the encoded proteins (the proteome) of the genomes of human, fly, worm, yeast, and representatives of bacteria and archaea in terms of the three-dimensional structures of their globular domains together with a general sequence-based study. We show that 39% of the human proteome can be assigned to known structures. We estimate that for 77% of the proteome, there is some functional annotation, but only 26% of the proteome can be assigned to standard sequence motifs that characterize function. Of the human protein sequences, 13% are transmembrane proteins, but only 3% of the residues in the proteome form membrane-spanning regions. There are substantial differences in the composition of globular domains of transmembrane proteins between the proteomes we have analyzed. Commonly occurring structural superfamilies are identified within the proteome. The frequencies of these superfamilies enable us to estimate that 98% of the human proteome evolved by domain duplication, with four of the 10 most duplicated superfamilies specific for multicellular organisms. The zinc-finger superfamily is massively duplicated in human compared to fly and worm, and occurrence of domains in repeats is more common in metazoa than in single cellular organisms. Structural superfamilies over- and underrepresented in human disease genes have been identified. Data and results can be downloaded and analyzed via web-based applications at http://www.sbg.bio.ic.ac.uk.

Algorithms↗

Deciphering the genetic background of an industrial 2-ketogluconic acid-producing strain Pseudomonas plecoglossicida JUIM01 using whole-genome sequencing.

2-Ketogluconic acid (2KGA) is an important precursor for the food antioxidant erythorbic acid, currently produced via microbial fermentation using Pseudomonas species. To facilitate the genetic improvement of production strains, the complete genome of an industrial 2KGA producer P. plecoglossicida JUIM01 was sequenced and analyzed. The genome consists of a 5.13-Mb circular chromosome with a GC content of 63.58%, encoding 4,517 predicted proteins. Comprehensive functional annotation identified a putative global regulatory network comprising 75 core regulators, which were classified into six functionally cooperative modules, potentially governing the strain's metabolism and environmental adaptability. We further delineated the genetic determinants hypothetically linked to efficient 2KGA synthesis, including glucose metabolism, fatty acid metabolism, and the oxidative phosphorylation system. These outputs could provide the genomic resource for elucidating high productivity and robustness, and rationally engineering the high-performance chassis cells toward robust 2KGA production.

P. plecoglossicida↗

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi↗

Mouse functional genomics requires standardization of mouse handling and housing conditions.

The study of mouse models is crucial for the functional annotation of the human genome. The recent improvements in mouse genetics now moved the bottleneck in mouse functional genomics from the generation of mutant mice lines to the phenotypic analysis of these mice lines. Simple, validated, and reproducible phenotyping tests are a prerequisite to improving this phenotyping bottleneck. We analyzed here the impact of simple variations in animal handling and housing procedures, such as cage density, diet, gender, length of fasting, as well as site (retro-orbital vs. tail), timing, and anesthesia used during venipuncture, on biochemical, hematological, and metabolic/endocrine parameters in adult C57BL/6J mice. Our results, which show that minor changes in procedures can profoundly affect biological variables, underscore the importance of establishing uniform and validated animal procedures to improve reproducibility of mouse phenotypic data.

Animals↗

Development of a citrus genome-wide EST collection and cDNA microarray as resources for genomic studies.

A functional genomics project has been initiated to approach the molecular characterization of the main biological and agronomical traits of citrus. As a key part of this project, a citrus EST collection has been generated from 25 cDNA libraries covering different tissues, developmental stages and stress conditions. The collection includes a total of 22,635 high-quality ESTs, grouped in 11,836 putative unigenes, which represent at least one third of the estimated number of genes in the citrus genome. Functional annotation of unigenes which have Arabidopsis orthologues (68% of all unigenes) revealed gene representation in every major functional category, suggesting that a genome-wide EST collection was obtained. A Citrus clementina Hort. ex Tan. cv. Clemenules genomic library, that will contribute to further characterization of relevant genes, has also been constructed. To initiate the analysis of citrus transcriptome, we have developed a cDNA microarray containing 12,672 probes corresponding to 6875 putative unigenes of the collection. Technical characterization of the microarray showed high intra- and inter-array reproducibility, as well as a good range of sensitivity. We have also validated gene expression data achieved with this microarray through an independent technique such as RNA gel blot analysis.

Citrus↗

Comparative transcriptome analysis provides insights into dorso-ventral color pattern formation of Holothuria edulis.

Animal body color patterns are highly diverse and play critical roles in camouflage, intraspecific communication, and environmental adaptation. Holothuria edulis, an important echinoderm inhabiting tropical waters, exhibits a typical dorsoventral dichromatism. This unique body color difference represents a key phenotypic trait for its habitat adaptation; however, the core differential genes regulating this trait remain to be elucidated. In this study, comparative transcriptome sequencing was performed on the dorsal and ventral body wall tissues of H. edulis, leading to the identification of a number of differentially expressed genes (DEGs), followed by GO functional annotation and KEGG pathway enrichment analysis. GO enrichment analysis indicated that the DEGs were significantly enriched in functional categories such as extracellular region, peptidase inhibitor activity, and tetrapyrrole binding. KEGG pathway analysis further revealed significant enrichment of protein digestion and absorption, the TNF signaling pathway, and cholesterol metabolism. Notably, the pigmentation-related gene FMO2 was highly expressed in the dorsal body wall tissue, whereas cyp1a1, ZIC1, Slc7a11, WNT-1, and ADAMTS20 were highly expressed in the ventral body wall tissue. This study identified DEGs and enriched pathways associated with dorsoventral body color differences in H. edulis, providing new insights into the molecular regulatory mechanisms underlying body color pattern formation. From the perspective of aquaculture applications, body color is one of the important traits affecting the quality and market value of sea cucumber products. Elucidating the molecular mechanisms of body color variation can provide a scientific basis for molecular marker-assisted breeding of superior sea cucumber variety.

Animals↗

Unravelling sex differences in the genetic architecture of anxiety.

BACKGROUND: Anxiety disorders show striking sex differences in prevalence, symptoms, and clinical characteristics, shaping how they manifest and are experienced. METHODS: Here, we report the first sex-specific meta-analysis of genome-wide association studies (GWAS) of anxiety, leveraging two of the largest biobank datasets, UK Biobank and All of Us, comprising 85,042 female cases with 196,789 controls and 36,732 male cases with 136,924 controls. Functional annotation, sex-specific polygenic scores (PGS), and genetic correlations were performed to assess genetic differences and functional implications. RESULTS: In females, 21 lead SNPs were significantly associated with anxiety, compared to five in males. Although the genetic correlation between sexes was high, it was significantly different from one, indicating partially distinct genetic architectures. In addition, both the SNP-based observed and liability-scale heritabilities (assuming a 2:1 female-to-male prevalence ratio) were significantly higher in females. Gene-based tests and functional prioritization identified different genes associated with anxiety in females and males. Moreover, genetic correlation analyses revealed stronger associations of female anxiety with attention-deficit/hyperactivity disorder (ADHD) and body mass index (BMI), whereas male anxiety showed stronger correlations with waist-hip-ratio-adjusted BMI. CONCLUSIONS: While the overall genetic architecture of anxiety is largely shared, our findings reveal distinct sex-specific genetic associations and correlations, highlighting the value of analyzing the sexes separately to uncover genetic signals that may be masked in sex-combined samples.

Female↗

Diffusion kernel-based logistic regression models for protein function prediction.

Assigning functions to unknown proteins is one of the most important problems in proteomics. Several approaches have used protein-protein interaction data to predict protein functions. We previously developed a Markov random field (MRF) based method to infer a protein's functions using protein-protein interaction data and the functional annotations of its protein interaction partners. In the original model, only direct interactions were considered and each function was considered separately. In this study, we develop a new model which extends direct interactions to all neighboring proteins, and one function to multiple functions. The goal is to understand a protein's function based on information on all the neighboring proteins in the interaction network. We first developed a novel kernel logistic regression (KLR) method based on diffusion kernels for protein interaction networks. The diffusion kernels provide means to incorporate all neighbors of proteins in the network. Second, we identified a set of functions that are highly correlated with the function of interest, referred to as the correlated functions, using the chi-square test. Third, the correlated functions were incorporated into our new KLR model. Fourth, we extended our model by incorporating multiple biological data sources such as protein domains, protein complexes, and gene expressions by converting them into networks. We showed that the KLR approach of incorporating all protein neighbors significantly improved the accuracy of protein function predictions over the MRF model. The incorporation of multiple data sets also improved prediction accuracy. The prediction accuracy is comparable to another protein function classifier based on the support vector machine (SVM), using a diffusion kernel. The advantages of the KLR model include its simplicity as well as its ability to explore the contribution of neighbors to the functions of proteins of interest.

Databases, Protein↗

Identification of transmembrane protein functions by binary topology patterns.

We propose a novel method for identifying and classifying the functions of transmembrane (TM) proteins based on their TM topology [the number of TM segments (tms), the loop length and the N-terminus location]. In this method, the TM topology is expressed as a string of '0' and '1', and this is designated the binary topology pattern (BTP). We focused on TM proteins with up to 12 tms, with the exception of 1 and 9 tms, and classified them into 37 functional groups by the number of tms and the functional annotation. These grouped TM protein sequences were used to determine BTPs which are specific to the individual functional groups. Since the evaluated accuracies (sensitivity, specificity and self-consistency) of these patterns in functional identification were quite high overall, i.e. 0.940, 0.934 and 0.935, respectively, as averaged over the 37 functional groups, we confirmed that TM protein function can be identified by the number of tms and the characteristics of loop lengths, i.e. BTPs.

Data Interpretation, Statistical↗

LocustDB: a relational database for the transcriptome and biology of the migratory locust (Locusta migratoria).

BACKGROUND: The migratory locust (Locusta migratoria) is an orthopteran pest and a representative member of hemimetabolous insects for biological studies. Its transcriptomic data provide invaluable information for molecular entomology and pave a way for the comparative research of other medically, agronomically, and ecologically relevant insects. We developed the first transcriptomic database of the locust (LocustDB), building necessary infrastructures to integrate, organize, and retrieve data that are either currently available or to be acquired in the future. DESCRIPTION: LocustDB currently hosts 45,474 high-quality EST sequences from the locust, which were assembled into 12,161 unigenes. It, through user-friendly web interfaces, allows investigators to freely access sequence data, including homologous/orthologous sequences, functional annotations, and pathway analysis, based on conserved orthologous groups (COG), gene ontology (GO), protein domain (InterPro), and functional pathways (KEGG). It also provides information from comparative analysis based on data from the migratory locust and five other invertebrate species, including the silkworm, the honeybee, the fruitfly, the mosquito and the nematode. The website address of LocustDB is http://locustdb.genomics.org.cn/. CONCLUSION: LocustDB starts with the first transcriptome information for an orthopteran and hemimetabolous insect and will be extended to provide a framework for incorporating in-coming genomic data of relevant insect groups and a workbench for cross-species comparative studies.

Animals↗

Automatic rule generation for protein annotation with the C4.5 data mining algorithm applied on SWISS-PROT.

MOTIVATION: The gap between the amount of newly submitted protein data and reliable functional annotation in public databases is growing. Traditional manual annotation by literature curation and sequence analysis tools without the use of automated annotation systems is not able to keep up with the ever increasing quantity of data that is submitted. Automated supplements to manually curated databases such as TrEMBL or GenPept cover raw data but provide only limited annotation. To improve this situation automatic tools are needed that support manual annotation, automatically increase the amount of reliable information and help to detect inconsistencies in manually generated annotations. RESULTS: A standard data mining algorithm was successfully applied to gain knowledge about the Keyword annotation in SWISS-PROT. 11 306 rules were generated, which are provided in a database and can be applied to yet unannotated protein sequences and viewed using a web browser. They rely on the taxonomy of the organism, in which the protein was found and on signature matches of its sequence. The statistical evaluation of the generated rules by cross-validation suggests that by applying them on arbitrary proteins 33% of their keyword annotation can be generated with an error rate of 1.5%. The coverage rate of the keyword annotation can be increased to 60% by tolerating a higher error rate of 5%. AVAILABILITY: The results of the automatic data mining process can be browsed on http://golgi.ebi.ac.uk:8080/Spearmint/ Source code is available upon request. CONTACT: kretsch@ebi.ac.uk.

Algorithms↗

HOPPSIGEN: a database of human and mouse processed pseudogenes.

Processed pseudogenes result from reverse transcribed mRNAs. In general, because processed pseudogenes lack promoters, they are no longer functional from the moment they are inserted into the genome. Subsequently, they freely accumulate substitutions, insertions and deletions. Moreover, the ancestral structure of processed pseudogenes could be easily inferred using the sequence of their functional homologous genes. Owing to these characteristics, processed pseudogenes represent good neutral markers for studying genome evolution. Recently, there is an increasing interest for these markers, particularly to help gene prediction in the field of genome annotation, functional genomics and genome evolution analysis (patterns of substitution). For these reasons, we have developed a method to annotate processed pseudogenes in complete genomes. To make them useful to different fields of research, we stored them in a nucleic acid database after having annotated them. In this work, we screened both mouse and human complete genomes from ENSEMBL to find processed pseudogenes generated from functional genes with introns. We used a conservative method to detect processed pseudogenes in order to minimize the rate of false positive sequences. Within processed pseudogenes, some are still having a conserved open reading frame and some have overlapping gene locations. We designated as retroelements all reverse transcribed sequences and more strictly, we designated as processed pseudogenes, all retroelements not falling in the two former categories (having a conserved open reading or overlapping gene locations). We annotated 5823 retroelements (5206 processed pseudogenes) in the human genome and 3934 (3428 processed pseudogenes) in the mouse genome. Compared to previous estimations, the total number of processed pseudogenes was underestimated but the aim of this procedure was to generate a high-quality dataset. To facilitate the use of processed pseudogenes in studying genome structure and evolution, DNA sequences from processed pseudogenes, and their functional reverse transcribed homologs, are now stored in a nucleic acid database, HOPPSIGEN. HOPPSIGEN can be browsed on the PBIL (Pole Bioinformatique Lyonnais) World Wide Web server (http://pbil.univ-lyon1.fr/) or fully downloaded for local installation.

Animals↗