PubMed Health⌕ Search

Biomedical subjects

Seung Y Rhee

Publications and source records attributed to Seung Y Rhee.

At least 19 recordsLinked to original sources

The plant structure ontology, a unified vocabulary of anatomy and morphology of a flowering plant.

Formal description of plant phenotypes and standardized annotation of gene expression and protein localization data require uniform terminology that accurately describes plant anatomy and morphology. This facilitates cross species comparative studies and quantitative comparison of phenotypes and expression patterns. A major drawback is variable terminology that is used to describe plant anatomy and morphology in publications and genomic databases for different species. The same terms are sometimes applied to different plant structures in different taxonomic groups. Conversely, similar structures are named by their species-specific terms. To address this problem, we created the Plant Structure Ontology (PSO), the first generic ontological representation of anatomy and morphology of a flowering plant. The PSO is intended for a broad plant research community, including bench scientists, curators in genomic databases, and bioinformaticians. The initial releases of the PSO integrated existing ontologies for Arabidopsis (Arabidopsis thaliana), maize (Zea mays), and rice (Oryza sativa); more recent versions of the ontology encompass terms relevant to Fabaceae, Solanaceae, additional cereal crops, and poplar (Populus spp.). Databases such as The Arabidopsis Information Resource, Nottingham Arabidopsis Stock Centre, Gramene, MaizeGDB, and SOL Genomics Network are using the PSO to describe expression patterns of genes and phenotypes of mutants and natural variants and are regularly contributing new annotations to the Plant Ontology database. The PSO is also used in specialized public databases, such as BRENDA, GENEVESTIGATOR, NASCArrays, and others. Over 10,000 gene annotations and phenotype descriptions from participating databases can be queried and retrieved using the Plant Ontology browser. The PSO, as well as contributed gene associations, can be obtained at www.plantontology.org.

Gene Expression Regulation, Plant↗

Whole-plant growth stage ontology for angiosperms and its application in plant biology.

Plant growth stages are identified as distinct morphological landmarks in a continuous developmental process. The terms describing these developmental stages record the morphological appearance of the plant at a specific point in its life cycle. The widely differing morphology of plant species consequently gave rise to heterogeneous vocabularies describing growth and development. Each species or family specific community developed distinct terminologies for describing whole-plant growth stages. This semantic heterogeneity made it impossible to use growth stage description contained within plant biology databases to make meaningful computational comparisons. The Plant Ontology Consortium (http://www.plantontology.org) was founded to develop standard ontologies describing plant anatomical as well as growth and developmental stages that can be used for annotation of gene expression patterns and phenotypes of all flowering plants. In this article, we describe the development of a generic whole-plant growth stage ontology that describes the spatiotemporal stages of plant growth as a set of landmark events that progress from germination to senescence. This ontology represents a synthesis and integration of terms and concepts from a variety of species-specific vocabularies previously used for describing phenotypes and genomic information. It provides a common platform for annotating gene function and gene expression in relation to the developmental trajectory of a plant described at the organismal level. As proof of concept the Plant Ontology Consortium used the plant ontology growth stage ontology to annotate genes and phenotypes in plants with initial emphasis on those represented in The Arabidopsis Information Resource, Gramene database, and MaizeGDB.

Arabidopsis↗

Systematic analysis of Arabidopsis organelles and a protein localization database for facilitating fluorescent tagging of full-length Arabidopsis proteins.

Cells are organized into a complex network of subcellular compartments that are specialized for various biological functions. Subcellular location is an important attribute of protein function. To facilitate systematic elucidation of protein subcellular location, we analyzed experimentally verified protein localization data of 1,300 Arabidopsis (Arabidopsis thaliana) proteins. The 1,300 experimentally verified proteins are distributed among 40 different compartments, with most of the proteins localized to four compartments: mitochondria (36%), nucleus (28%), plastid (17%), and cytosol (13.3%). About 19% of the proteins are found in multiple compartments, in which a high proportion (36.4%) is localized to both cytosol and nucleus. Characterization of the overrepresented Gene Ontology molecular functions and biological processes suggests that the Golgi apparatus and peroxisome may play more diverse functions but are involved in more specialized processes than other compartments. To support systematic empirical determination of protein subcellular localization using a technology called fluorescent tagging of full-length proteins, we developed a database and Web application to provide preselected green fluorescent protein insertion position and primer sequences for all Arabidopsis proteins to study their subcellular localization and to store experimentally verified protein localization images, videos, and their annotations of proteins generated using the fluorescent tagging of full-length proteins technology. The database can be searched, browsed, and downloaded using a Web browser at http://aztec.stanford.edu/gfp/. The software can also be downloaded from the same Web site for local installation.

Arabidopsis↗

MetaCyc: a multiorganism database of metabolic pathways and enzymes.

MetaCyc is a database of metabolic pathways and enzymes located at http://MetaCyc.org/. Its goal is to serve as a metabolic encyclopedia, containing a collection of non-redundant pathways central to small molecule metabolism, which have been reported in the experimental literature. Most of the pathways in MetaCyc occur in microorganisms and plants, although animal pathways are also represented. MetaCyc contains metabolic pathways, enzymatic reactions, enzymes, chemical compounds, genes and review-level comments. Enzyme information includes substrate specificity, kinetic properties, activators, inhibitors, cofactor requirements and links to sequence and structure databases. Data are curated from the primary literature by curators with expertise in biochemistry and molecular biology. MetaCyc serves as a readily accessible comprehensive resource on microbial and plant pathways for genome analysis, basic research, education, metabolic engineering and systems biology. Querying, visualization and curation of the database is supported by SRI's Pathway Tools software. The PathoLogic component of Pathway Tools is used in conjunction with MetaCyc to predict the metabolic network of an organism from its annotated genome. SRI and the European Bioinformatics Institute employed this tool to create pathway/genome databases (PGDBs) for 165 organisms, available at the BioCyc.org website. These PGDBs also include predicted operons and pathway hole fillers.

Animals↗

Taking the first steps towards a standard for reporting on phylogenies: Minimum Information About a Phylogenetic Analysis (MIAPA).

In the eight years since phylogenomics was introduced as the intersection of genomics and phylogenetics, the field has provided fundamental insights into gene function, genome history and organismal relationships. The utility of phylogenomics is growing with the increase in the number and diversity of taxa for which whole genome and large transcriptome sequence sets are being generated. We assert that the synergy between genomic and phylogenetic perspectives in comparative biology would be enhanced by the development and refinement of minimal reporting standards for phylogenetic analyses. Encouraged by the development of the Minimum Information About a Microarray Experiment (MIAME) standard, we propose a similar roadmap for the development of a Minimal Information About a Phylogenetic Analysis (MIAPA) standard. Key in the successful development and implementation of such a standard will be broad participation by developers of phylogenetic analysis software, phylogenetic database developers, practitioners of phylogenomics, and journal editors.

Genomics↗

PatMatch: a program for finding patterns in peptide and nucleotide sequences.

Here, we present PatMatch, an efficient, web-based pattern-matching program that enables searches for short nucleotide or peptide sequences such as cis-elements in nucleotide sequences or small domains and motifs in protein sequences. The program can be used to find matches to a user-specified sequence pattern that can be described using ambiguous sequence codes and a powerful and flexible pattern syntax based on regular expressions. A recent upgrade has improved performance and now supports both mismatches and wildcards in a single pattern. This enhancement has been achieved by replacing the previous searching algorithm, scan_for_matches [D'Souza et al. (1997), Trends in Genetics, 13, 497-498], with nondeterministic-reverse grep (NR-grep), a general pattern matching tool that allows for approximate string matching [Navarro (2001), Software Practice and Experience, 31, 1265-1312]. We have tailored NR-grep to be used for DNA and protein searches with PatMatch. The stand-alone version of the software can be adapted for use with any sequence dataset and is available for download at The Arabidopsis Information Resource (TAIR) at ftp://ftp.arabidopsis.org/home/tair/Software/Patmatch/. The PatMatch server is available on the web at http://www.arabidopsis.org/cgi-bin/patmatch/nph-patmatch.pl for searching Arabidopsis thaliana sequences.

Arabidopsis↗

An ontology for cell types.

We describe an ontology for cell types that covers the prokaryotic, fungal, animal and plant worlds. It includes over 680 cell types. These cell types are classified under several generic categories and are organized as a directed acyclic graph. The ontology is available in the formats adopted by the Open Biological Ontologies umbrella and is designed to be used in the context of model organism genome and other biological databases. The ontology is freely available at http://obo.sourceforge.net/ and can be viewed using standard ontology visualization tools such as OBO-Edit and COBrA.

Animals↗

Community-based gene structure annotation.

Uncertainty and inconsistency of gene structure annotation remain limitations on research in the genome era, frustrating both biologists and bioinformaticians, who have to sort out annotation errors for their genes of interest or to generate trustworthy datasets for algorithmic development. It is unrealistic to hope for better software solutions in the near future that would solve all the problems. The issue is all the more urgent with more species being sequenced and analyzed by comparative genomics - erroneous annotations could easily propagate, whereas correct annotations in one species will greatly facilitate annotation of novel genomes. We propose a dynamic, economically feasible solution to the annotation predicament: broad-based, web-technology-enabled community annotation, a prototype of which is now in use for Arabidopsis.

Arabidopsis↗

MetaCyc and AraCyc. Metabolic pathway databases for plant research.

MetaCyc (http://metacyc.org) contains experimentally determined biochemical pathways to be used as a reference database for metabolism. In conjunction with the Pathway Tools software, MetaCyc can be used to computationally predict the metabolic pathway complement of an annotated genome. To increase the breadth of pathways and enzymes, more than 60 plant-specific pathways have been added or updated in MetaCyc recently. In contrast to MetaCyc, which contains metabolic data for a wide range of organisms, AraCyc is a species-specific database containing only enzymes and pathways found in the model plant Arabidopsis (Arabidopsis thaliana). AraCyc (http://arabidopsis.org/tools/aracyc/) was the first computationally predicted plant metabolism database derived from MetaCyc. Since its initial computational build, AraCyc has been under continued curation to enhance data quality and to increase breadth of pathway coverage. Twenty-eight pathways have been manually curated from the literature recently. Pathway predictions in AraCyc have also been recently updated with the latest functional annotations of Arabidopsis genes that use controlled vocabulary and literature evidence. AraCyc currently features 1,418 unique genes mapped onto 204 pathways with 1,156 literature citations. The Omics Viewer, a user data visualization and analysis tool, allows a list of genes, enzymes, or metabolites with experimental values to be painted on a diagram of the full pathway map of AraCyc. Other recent enhancements to both MetaCyc and AraCyc include implementation of an evidence ontology, which has been used to provide information on data quality, expansion of the secondary metabolism node of the pathway ontology to accommodate curation of secondary metabolic pathways, and enhancement of the cellular component ontology for storing and displaying enzyme and pathway locations within subcellular compartments.

4-Hydroxyphenylpyruvate Dioxygenase↗

Functional annotation of the Arabidopsis genome using controlled vocabularies.

Controlled vocabularies are increasingly used by databases to describe genes and gene products because they facilitate identification of similar genes within an organism or among different organisms. One of The Arabidopsis Information Resource's goals is to associate all Arabidopsis genes with terms developed by the Gene Ontology Consortium that describe the molecular function, biological process, and subcellular location of a gene product. We have also developed terms describing Arabidopsis anatomy and developmental stages and use these to annotate published gene expression data. As of March 2004, we used computational and manual annotation methods to make 85,666 annotations representing 26,624 unique loci. We focus on associating genes to controlled vocabulary terms based on experimental data from the literature and use The Arabidopsis Information Resource-developed PubSearch software to facilitate this process. Each annotation is tagged with a combination of evidence codes, evidence descriptions, and references that provide a robust means to assess data quality. Annotation of all Arabidopsis genes will allow quantitative comparisons between sets of genes derived from sources such as microarray experiments. The Arabidopsis annotation data will also facilitate annotation of newly sequenced plant genomes by using sequence similarity to transfer annotations to homologous genes. In addition, complete and up-to-date annotations will make unknown genes easy to identify and target for experimentation. Here, we describe the process of Arabidopsis functional annotation using a variety of data sources and illustrate several ways in which this information can be accessed and used to infer knowledge about Arabidopsis and other plant species.

Arabidopsis↗

MetaCyc: a multiorganism database of metabolic pathways and enzymes.

The MetaCyc database (see URL http://MetaCyc.org) is a collection of metabolic pathways and enzymes from a wide variety of organisms, primarily microorganisms and plants. The goal of MetaCyc is to contain a representative sample of each experimentally elucidated pathway, and thereby to catalog the universe of metabolism. MetaCyc also describes reactions, chemical compounds and genes. Many of the pathways and enzymes in MetaCyc contain extensive information, including comments and literature citations. SRI's Pathway Tools software supports querying, visualization and curation of MetaCyc. With its wide breadth and depth of metabolic information, MetaCyc is a valuable resource for a variety of applications. MetaCyc is the reference database of pathways and enzymes that is used in conjunction with SRI's metabolic pathway prediction program to create Pathway/Genome Databases that can be augmented with curation from the scientific literature and published on the world wide web. MetaCyc also serves as a readily accessible comprehensive resource on microbial and plant pathways for genome analysis, basic research, education, metabolic engineering and systems biology. In the past 2 years the data content and the Pathway Tools software used to query, visualize and edit MetaCyc have been expanded significantly. These enhancements are described in this paper.

Biochemical Phenomena↗

High-throughput fluorescent tagging of full-length Arabidopsis gene products in planta.

We developed a high-throughput methodology, termed fluorescent tagging of full-length proteins (FTFLP), to analyze expression patterns and subcellular localization of Arabidopsis gene products in planta. Determination of these parameters is a logical first step in functional characterization of the approximately one-third of all known Arabidopsis genes that encode novel proteins of unknown function. Our FTFLP-based approach offers two significant advantages: first, it produces internally-tagged full-length proteins that are likely to exhibit native intracellular localization, and second, it yields information about the tissue specificity of gene expression by the use of native promoters. To demonstrate how FTFLP may be used for characterization of the Arabidopsis proteome, we tagged a series of known proteins with diverse subcellular targeting patterns as well as several proteins with unknown function and unassigned subcellular localization.

Arabidopsis↗

MAPMAN: a user-driven tool to display genomics data sets onto diagrams of metabolic pathways and other biological processes.

MAPMAN is a user-driven tool that displays large data sets onto diagrams of metabolic pathways or other processes. SCAVENGER modules assign the measured parameters to hierarchical categories (formed 'BINs', 'subBINs'). A first build of TRANSCRIPTSCAVENGER groups genes on the Arabidopsis Affymetrix 22K array into >200 hierarchical categories, providing a breakdown of central metabolism (for several pathways, down to the single enzyme level), and an overview of secondary metabolism and cellular processes. METABOLITESCAVENGER groups hundreds of metabolites into pathways or groups of structurally related compounds. An IMAGEANNOTATOR module uses these groupings to organise and display experimental data sets onto diagrams of the users' choice. A modular structure allows users to edit existing categories, add new categories and develop SCAVENGER modules for other sorts of data. MAPMAN is used to analyse two sets of 22K Affymetrix arrays that investigate the response of Arabidopsis rosettes to low sugar: one investigates the response to a 6-h extension of the night, and the other compares wild-type Columbia-0 (Col-0) and the starchless pgm mutant (plastid phosphoglucomutase) at the end of the night. There were qualitatively similar responses in both treatments. Many genes involved in photosynthesis, nutrient acquisition, amino acid, nucleotide, lipid and cell wall synthesis, cell wall modification, and RNA and protein synthesis were repressed. Many genes assigned to amino acid, nucleotide, lipid and cell wall breakdown were induced. Changed expression of genes for trehalose metabolism point to a role for trehalose-6-phosphate (Tre6P) as a starvation signal. Widespread changes in the expression of genes encoding receptor kinases, transcription factors, components of signalling pathways, proteins involved in post-translational modification and turnover, and proteins involved in the synthesis and sensing of cytokinins, abscisic acid (ABA) and ethylene revealing large-scale rewiring of the regulatory network is an early response to sugar depletion.

Abscisic Acid↗

Freezing-sensitive tomato has a functional CBF cold response pathway, but a CBF regulon that differs from that of freezing-tolerant Arabidopsis.

Many plants increase in freezing tolerance in response to low temperature, a process known as cold acclimation. In Arabidopsis, cold acclimation involves action of the CBF cold response pathway. Key components of the pathway include rapid cold-induced expression of three homologous genes encoding transcriptional activators, CBF1, 2 and 3 (also known as DREB1b, c and a, respectively), followed by expression of CBF-targeted genes, the CBF regulon, that increase freezing tolerance. Unlike Arabidopsis, tomato cannot cold acclimate raising the question of whether it has a functional CBF cold response pathway. Here we show that tomato, like Arabidopsis, encodes three CBF homologs, LeCBF1-3 (Lycopersicon esculentum CBF1-3), that are present in tandem array in the genome. Only the tomato LeCBF1 gene, however, was found to be cold-inducible. As is the case for Arabidopsis CBF1-3, transcripts for LeCBF1-3 did accumulate in response to mechanical agitation, but not in response to drought, ABA or high salinity. Constitutive overexpression of LeCBF1 in transgenic Arabidopsis plants induced expression of CBF-targeted genes and increased freezing tolerance indicating that LeCBF1 encodes a functional homolog of the Arabidopsis CBF1-3 proteins. However, constitutive overexpression of either LeCBF1 or AtCBF3 in transgenic tomato plants did not increase freezing tolerance. Gene expression studies, including the use of a cDNA microarray representing approximately 8000 tomato genes, identified only four genes that were induced 2.5-fold or more in the LeCBF1 or AtCBF3 overexpressing plants, three of which were putative members of the tomato CBF regulon as they were also upregulated in response to low temperature. Additional experiments indicated that of eight tomato genes that were likely orthologs of Arabidopsis CBF regulon genes, none were responsive to CBF overexpression in tomato. From these results, we conclude that tomato has a complete CBF cold response pathway, but that the tomato CBF regulon differs from that of Arabidopsis and appears to be considerably smaller and less diverse in function.

Acclimatization↗

Strategies for avoiding reinventing the precollege education and outreach wheel.

The National Science Foundation's recent mandate that all Principal Investigators address the broader impacts of their research has prompted an unprecedented number of scientists to seek opportunities to participate in precollege education and outreach. To help interested geneticists avoid duplicating efforts and make use of existing resources, we examined several precollege genetics, genomics, and biotechnology education efforts and noted the elements that contributed to their success, indicated by program expansion, participant satisfaction, or participant learning. Identifying a specific audience and their needs and resources, involving K-12 teachers in program development, and evaluating program efforts are integral to program success. We highlighted a few innovative programs to illustrate these findings. Challenges that may compromise further development and dissemination of these programs include absence of reward systems for participation in outreach as well as lack of training for scientists doing outreach. Several programs and institutions are tackling these issues in ways that will help sustain outreach efforts while allowing them to be modified to meet the changing needs of their participants, including scientists, teachers, and students. Most importantly, resources and personnel are available to facilitate greater and deeper involvement of scientists in precollege and public education.

Biotechnology↗

Microspore separation in the quartet 3 mutants of Arabidopsis is impaired by a defect in a developmentally regulated polygalacturonase required for pollen mother cell wall degradation.

Mutations in the QUARTET loci in Arabidopsis result in failure of microspore separation during pollen development due to a defect in degradation of the pollen mother cell wall during late stages of pollen development. Mutations in a new locus required for microspore separation, QRT3, were isolated, and the corresponding gene was cloned by T-DNA tagging. QRT3 encodes a protein that is approximately 30% similar to an endopolygalacturonase from peach (Prunus persica). The QRT3 protein was expressed in yeast (Saccharomyces cerevisiae) and found to exhibit polygalacturonase activity. In situ hybridization experiments showed that QRT3 is specifically and transiently expressed in the tapetum during the phase when microspores separate from their meiotic siblings. Immunohistochemical localization of QRT3 indicated that the protein is secreted from tapetal cells during the early microspore stage. Thus, QRT3 plays a direct role in degrading the pollen mother cell wall during microspore development.

Amino Acid Sequence↗

AraCyc: a biochemical pathway database for Arabidopsis.

AraCyc is a database containing biochemical pathways of Arabidopsis, developed at The Arabidopsis Information Resource (http://www.arabidopsis.org). The aim of AraCyc is to represent Arabidopsis metabolism as completely as possible with a user-friendly Web-based interface. It presently features more than 170 pathways that include information on compounds, intermediates, cofactors, reactions, genes, proteins, and protein subcellular locations. The database uses Pathway Tools software, which allows the users to visualize a bird's eye view of all pathways in the database down to the individual chemical structures of the compounds. The database was built using Pathway Tools' Pathologic module with MetaCyc, a collection of pathways from more than 150 species, as a reference database. This initial build was manually refined and annotated. More than 20 plant-specific pathways, including carotenoid, brassinosteroid, and gibberellin biosyntheses have been added from the literature. A list of more than 40 plant pathways will be added in the coming months. The quality of the initial, automatic build of the database was compared with the manually improved version, and with EcoCyc, an Escherichia coli database using the same software system that has been manually annotated for many years. In addition, a Perl interface, PerlCyc, was developed that allows programmers to access Pathway Tools databases from the popular Perl language. AraCyc is available at the tools section of The Arabidopsis Information Resource Web site (http://www.arabidopsis.org/tools/aracyc).

Arabidopsis↗