PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

GeConT: gene context analysis.

SUMMARY: The fact that adjacent genes in bacteria are often functionally related is widely known. GeConT (Gene Context Tool) is a web interface designed to visualize genome context of a gene or a group of genes and their orthologs in all the completely sequenced genomes. The graphical information of GeConT can be used to analyze genome annotation, functional ortholog identification or to verify the genomic context congruence of any set of genes that share a common property. AVAILABILITY: http://www.ibt.unam.mx/biocomputo/gecont.html

Chromosome Mapping↗

Differential brain transcriptome of beta4 nAChR subunit-deficient mice: is it the effect of the null mutation or the background strain?

Studies using mice with beta4 nicotinic acetylcholine receptor (nAChR) subunit deficiency (beta4-/- mice) helped reveal the roles of this subunit in bradycardiac response to vagal stimulation, nicotine-induced seizure activity and anxiety. To identify genes that might be related to beta4-containing nAChRs activity, we compared the mRNA expression profiles of brains from beta4-/- and wild-type mice using Affymetrix U74Av2 microarray. Seventy-seven genes significantly differentiated between these two experimental groups. Of them, the two most downregulated were spastic paraplegia 21 (human) homolog (Spg21) and 6-pyruvoyl-tetrahydropterin synthase (Pts) genes. Since the targeted mutagenesis of the beta4 nAChR subunit was done by using two mouse strains, 129SvEv and C57BL/6J, it is possible that the genes closely linked to the mutated beta4 gene represent the 129SvEv allele and not the control C57BL/6J-driven allele. We examined this possibility by using public database and quantitative RT-PCR. The expression levels of Spg21 and Pts genes that, like the beta4 gene, are localized on mouse chromosome 9, as well as the expression levels of other genes located on this chromosome, were dependent on the mouse background strain. The 67 differentially expressed genes that are not located on chromosome 9 were further analyzed for overrepresented functional annotations and transcription regulatory elements compared with the entire microarray. Genes encoding for proteins involved in tyrosine phosphatase activity, calcium ion binding, cell growth and/or maintenance, and chromosome organization were overrepresented. Our data enhance the understanding of the molecular interactions involved in the beta4 nAChR subunit function. They also emphasize the need for careful interpretation of expression microarray studies done on genetically manipulated animals.

Animals↗

MMDB: annotating protein sequences with Entrez's 3D-structure database.

Three-dimensional (3D) structure is now known for a large fraction of all protein families. Thus, it has become rather likely that one will find a homolog with known 3D structure when searching a sequence database with an arbitrary query sequence. Depending on the extent of similarity, such neighbor relationships may allow one to infer biological function and to identify functional sites such as binding motifs or catalytic centers. Entrez's 3D-structure database, the Molecular Modeling Database (MMDB), provides easy access to the richness of 3D structure data and its large potential for functional annotation. Entrez's search engine offers several tools to assist biologist users: (i) links between databases, such as between protein sequences and structures, (ii) pre-computed sequence and structure neighbors, (iii) visualization of structure and sequence/structure alignment. Here, we describe an annotation service that combines some of these tools automatically, Entrez's 'Related Structure' links. For all proteins in Entrez, similar sequences with known 3D structure are detected by BLAST and alignments are recorded. The 'Related Structure' service summarizes this information and presents 3D views mapping sequence residues onto all 3D structures available in MMDB (http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=structure).

Databases, Protein↗

Extracting knowledge from dynamics in gene expression.

Most investigations of coordinated gene expression have focused on identifying correlated expression patterns between genes by examining their normalized static expression levels. In this study, we focus on the dynamics of gene expression by seeking to identify correlated patterns of changes in genetic expression level. In doing so, we build upon methods developed in clinical informatics to detect temporal trends of laboratory and other clinical data. We construct relevance networks from Saccharomyces cerevisiae gene-expression dynamics data and find genes with related functional annotations grouped together. While some of these associations are also found using a standard expression level analysis, many are identified exclusively through the dynamic analysis. These results strongly suggest that the analysis of gene expression dynamics is a necessary and important tool for studying regulatory and other functional relationships among genes. The source code developed for this investigation is freely available to all non-commercial investigators by contacting the authors.

Cluster Analysis↗

Bioinformatics for venom and toxin sciences.

Venomous animals produce a myriad of important pharmacological components. The individual components, or venoms (toxins), are used in ion channel and receptor studies, drug discovery, and formulation of insecticides. The toxin data are scattered across public databases which provide sequence and structural descriptions, but very limited functional annotation. The exponential growth of newly identified toxin data has created a need for better data management. Venominformatics is a systematic bioinformatics approach in which classified, consolidated and cleaned venom data are stored into repositories and integrated with advanced bioinformatics tools for the analysis of structure and function of toxins. Venominformatics complements experimental studies and helps reduce the number of essential experiments.

Animals↗

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant↗

Functional network analysis reveals extended gliomagenesis pathway maps and three novel MYC-interacting genes in human gliomas.

Gene expression profiling has proven useful in subclassification and outcome prognostication for human glial brain tumors. The analysis of biological significance of the hundreds or thousands of alterations in gene expression found in genomic profiling remains a major challenge. Moreover, it is increasingly evident that genes do not act as individual units but collaborate in overlapping networks, the deregulation of which is a hallmark of cancer. Thus, we have here applied refined network knowledge to the analysis of key functions and pathways associated with gliomagenesis in a set of 50 human gliomas of various histogenesis, using cDNA microarrays, inferential and descriptive statistics, and dynamic mapping of gene expression data into a functional annotation database. Highest-significance networks were assembled around the myc oncogene in gliomagenesis and around the integrin signaling pathway in the glioblastoma subtype, which is paradigmatic for its strong migratory and invasive behavior. Three novel MYC-interacting genes (UBE2C, EMP1, and FBXW7) with cancer-related functions were identified as network constituents differentially expressed in gliomas, as was CD151 as a new component of a network that mediates glioblastoma cell invasion. Complementary, unsupervised relevance network analysis showed a conserved self-organization of modules of interconnected genes with functions in cell cycle regulation in human gliomas. This approach has extended existing knowledge about the organizational pattern of gene expression in human gliomas and identified potential novel targets for future therapeutic development.

Adult↗

Comparative assessment of performance and genome dependence among phylogenetic profiling methods.

BACKGROUND: The rapidly increasing speed with which genome sequence data can be generated will be accompanied by an exponential increase in the number of sequenced eukaryotes. With the increasing number of sequenced eukaryotic genomes comes a need for bioinformatic techniques to aid in functional annotation. Ideally, genome context based techniques such as proximity, fusion, and phylogenetic profiling, which have been so successful in prokaryotes, could be utilized in eukaryotes. Here we explore the application of phylogenetic profiling, a method that exploits the evolutionary co-occurrence of genes in the assignment of functional linkages, to eukaryotic genomes. RESULTS: In order to evaluate the performance of phylogenetic profiling in eukaryotes, we assessed the relative performance of commonly used profile construction techniques and genome compositions in predicting functional linkages in both prokaryotic and eukaryotic organisms. When predicting linkages in E. coli with a prokaryotic profile, the use of continuous values constructed from transformed BLAST bit-scores performed better than profiles composed of discretized E-values; the use of discretized E-values resulted in more accurate linkages when using S. cerevisiae as the query organism. Extending this analysis by incorporating several eukaryotic genomes in profiles containing a majority of prokaryotes resulted in similar overall accuracy, but with a surprising reduction in pathway diversity among the most significant linkages. Furthermore, the application of phylogenetic profiling using profiles composed of only eukaryotes resulted in the loss of the strong correlation between common KEGG pathway membership and profile similarity score. Profile construction methods, orthology definitions, ontology and domain complexity were explored as possible sources of the poor performance of eukaryotic profiles, but with no improvement in results. CONCLUSION: Given the current set of completely sequenced eukaryotic organisms, phylogenetic profiling using profiles generated from any of the commonly used techniques was found to yield extremely poor results. These findings imply genome-specific requirements for constructing functionally relevant phylogenetic profiles, and suggest that differences in the evolutionary history between different kingdoms might generally limit the usefulness of phylogenetic profiling in eukaryotes.

Bacterial Proteins↗

Genome-wide reverse genetics framework to identify novel functions of the vertebrate secretome.

BACKGROUND: Understanding the functional role(s) of the more than 20,000 proteins of the vertebrate genome is a major next step in the post-genome era. The approximately 4,000 co-translationally translocated (CTT) proteins - representing the vertebrate secretome - are important for such vertebrate-critical processes as organogenesis. However, the role(s) for most of these genes is currently unknown. RESULTS: We identified 585 putative full-length zebrafish CTT proteins using cross-species genomic and EST-based comparative sequence analyses. We further investigated 150 of these genes (Figure 1) for unique function using morpholino-based analysis in zebrafish embryos. 12% of the CTT protein-deficient embryos resulted in specific developmental defects, a notably higher rate of gene function annotation than the 2%-3% estimate from random gene mutagenesis studies. CONCLUSION: This initial collection includes novel genes required for the development of vascular, hematopoietic, pigmentation, and craniofacial tissues, as well as lipid metabolism, and organogenesis. This study provides a framework utilizing zebrafish for the systematic assignment of biological function in a vertebrate genome.

Amino Acid Sequence↗

Chromosome-level genome assembly and annotation of Petunia hybrida.

Petunia hybrida is the world's most popular garden plant and is regarded as a supermodel for studying the biology associated with the Asterid clade, the largest of the two major groups of flowering plants. Unlike other Solanaceae, petunia has a base chromosome number of seven, not 12. This along with recombination suppression has previously hindered efforts to assemble its genome to chromosome level. Here we achieve a chromosome-level assembly for P. hybrida using a combination of short-read and long-read sequencing, optical mapping (Bionano) and Hi-C technologies. The resulting assembly spans 1253.6 Mb with a BUSCO score of 99.8%. A total of 35,089 genes were predicted and of those 29,655 were functionally annotated. Syntenic regions between petunia, tomato and pepper were identified, highlighting rearrangements that have occurred since their divergence indicating that the 12 chromosomes of Solanaceae did not originate from whole genome duplication of an ancestral species with seven chromosomes like petunia. This assembly will enhance trait mapping efficiency and serve as a valuable resource for functional genomic studies.

Petunia↗

Does everything now make (anti)sense?

The data generated by the FANTOM (Functional Annotation of Mouse) consortium, Compugen and Affymetrix have collectively provided evidence that most of the mammalian genomes are actively transcribed. The emergence of an antisense RNA world brings new practical complexities to the study and detection of gene expression. However, we also need to address the fundamental questions regarding the functional importance of these molecules. In this brief paper, we focus on non-coding natural antisense transcription, as it appears to be a potentially powerful mechanism for extending the complexity of the protein coding genome, which is currently unable to explain inter-species diversification.

Animals↗

Putting microarrays in a context: integrated analysis of diverse biological data.

In recent years, multiple types of high-throughput functional genomic data that facilitate rapid functional annotation of sequenced genomes have become available. Gene expression microarrays are the most commonly available source of such data. However, genomic data often sacrifice specificity for scale, yielding very large quantities of relatively lower-quality data than traditional experimental methods. Thus sophisticated analysis methods are necessary to make accurate functional interpretation of these large-scale data sets. This review presents an overview of recently developed methods that integrate the analysis of microarray data with sequence, interaction, localisation and literature data, and further outlines current challenges in the field. The focus of this review is on the use of such methods for gene function prediction, understanding of protein regulation and modelling of biological networks.

Algorithms↗

A two-dimensional electrophoresis proteomic reference map and systematic identification of 1367 proteins from a cell suspension culture of the model legume Medicago truncatula.

The proteome of a Medicago truncatula cell suspension culture was analyzed using two-dimensional electrophoresis and nanoscale HPLC coupled to a tandem Q-TOF mass spectrometer (QSTAR Pulsar i) to yield an extensive protein reference map. Coomassie Brilliant Blue R-250 was used to visualize more than 1661 proteins, which were excised, subjected to in-gel trypsin digestion, and analyzed using nanoscale HPLC/MS/MS. The resulting spectral data were queried against a custom legume protein database using the MASCOT search engine. A total of 1367 of the 1661 proteins were identified with high rigor, yielding an identification success rate of 83% and 907 unique protein accession numbers. Functional annotation of the M. truncatula suspension cell proteins revealed a complete tricarboxylic acid cycle, a nearly complete glycolytic pathway, a significant portion of the ubiquitin pathway with the associated proteolytic and regulatory complexes, and many enzymes involved in secondary metabolism such as flavonoid/isoflavonoid, chalcone, and lignin biosynthesis. Proteins were also identified from most other functional classes including primary metabolism, energy production, disease/defense, protein destination/storage, protein synthesis, transcription, cell growth/division, and signal transduction. This work represents the most extensive proteomic description of M. truncatula suspension cells to date and provides a reference map for future comparative proteomic and functional genomic studies of the response of these cells to biotic and abiotic stress.

Amino Acid Sequence↗

A new family of NAD(P)H-dependent oxidoreductases distinct from conventional Rossmann-fold proteins.

A new family of NAD(P)H-dependent oxidoreductases is now recognized as a protein family distinct from conventional Rossmann-fold proteins. Numerous putative proteins belonging to the family have been annotated as malate dehydrogenase (MDH) or lactate dehydrogenase (LDH) according to the previous classification as type-2 malate/L-lactate dehydrogenases. However, recent biochemical and genetic studies have revealed that the protein family consists of a wide variety of enzymes with unique catalytic activities other than MDH or LDH activity. Based on their sequence homologies and plausible functions, the family proteins can be grouped into eight clades. This classification would be useful for reliable functional annotation of the new family of NAD(P)H-dependent oxidoreductases.

Amino Acid Sequence↗

Genome analysis of the glycosphingolipid-producing green alga tetraselmis sp. NKG400013.

Microalgae are gaining attention as sustainable resources for the production of valuable compounds, including biofuels, pigments, and bioactive metabolites. To support metabolic engineering and genome editing approaches aimed at enhancing these traits, high-quality genome assemblies are essential; however, genomic information remains limited for many microalgal lineages. Tetraselmis sp. NKG400013 is a green alga known for high glycosphingolipid accumulation with distinctive structural features. Here, we report a draft genome assembly of this strain generated using PacBio HiFi sequencing and transcriptome-supported annotation. The assembled genome spans 423.7 Mbp, with 74.5% repetitive sequences and 15,322 predicted protein-coding genes. Comparative analyses across 11 green algal species revealed a positive correlation between genome sizes and repeat contents, indicating that transposable element expansion, particularly long terminal repeat retrotransposons, has substantially contributed to genome enlargement in Tetraselmis. Genome-wide functional annotation and ortholog inference identified core enzymes required for glycosylceramide biosynthesis. Both sphingolipid Δ4 and Δ8 desaturases were identified in Tetraselmis and their coexistence suggests an expanded capacity for long-chain base modification that may underlie its distinctive glycosphingolipid profile. These results establish a genomic framework for understanding the high glycosphingolipid-producing capacity of NKG400013 and provide insights into the evolutionary diversification of sphingolipid metabolism in green algae.

Chlorophyta↗

C. elegans: an invaluable model organism for the proteomics studies of the cholesterol-mediated signaling pathway.

With the availability of its complete genome sequence and unique biological features relevant to human disease, Caenorhabditis elegans has become an invaluable model organism for the studies of proteomics, leading to the elucidation of nematode gene function. A journey from the genome to proteome of C. elegans may begin with preparation of expressed proteins, which enables a large-scale analysis of all possible proteins expressed under specific physiological conditions. Although various techniques have been used for proteomic analysis of C. elegans, systematic high-throughput analysis is still to come in order to accommodate studies of post-translational modification and quantitative analysis. Given that no integrated C. elegans protein expression database is available, it is about time that a global C. elegans proteome project is launched through which datasets of transcriptomes, protein-protein interaction and functional annotation can be integrated. As an initial target of a pilot project of the C. elegans proteome project, the cholesterol-mediated signaling pathway will be an excellent example since, like in other organisms, it is one of the key controlling pathways in cell growth and development in C. elegans. As this field tends to broaden to functional proteomics, there is a high demand to develop the versatile proteome informatics tools that can mange many different data in an integrative manner.

Animals↗

Genome and protein evolution in eukaryotes.

The past year has seen the completion of the genome sequence of the flowering plant Arabidopsis thaliana and the initial sequence reports of the human genome. The availability of completely sequenced eukaryotic genomes from disparate phylogenetic lineages has opened the door to comparative analyses and a better understanding of the evolutionary processes shaping genomes. Complex many-to-many relationships between genes from different species appear to be the norm, suggesting that transfer of detailed functional annotation will not be straightforward. In addition to expansion and contraction of gene families, new genes evolve from recombination of pre-existing domains, although some domain families do appear to have evolved recently and to be specific to restricted phylogenetic lineages. The overall picture is of a huge diversity of gene content within eukaryotic genomes, reflecting different functional demands in different species.

Animals↗

NEMBASE: a resource for parasitic nematode ESTs.

NEMBASE (available at http://www.nematodes.org) is a publicly available online database providing access to the sequence and associated meta-data currently being generated as part of the Edinburgh-Wellcome Trust Sanger Institute parasitic nematode EST project. NEMBASE currently holds approximately 100 000 sequences from 10 different species of nematode. To facilitate ease of use, sequences have been processed to generate a non-redundant set of gene objects ('partial genome') for each species. Users may query the database on the basis of BLAST annotation, sequence similarity or expression profiles. NEMBASE also features an interactive Java-based tool (SimiTri) which allows the simultaneous display and analysis of the relative similarity relationships of groups of sequences to three different databases. NEMBASE is currently being expanded to include sequence data from other nematode species. Other developments include access to accurate peptide predictions, improved functional annotation and incorporation of automated processes allowing rapid analysis of nematode-specific gene families.

Animals↗