PubMed Health⌕ Search

Biomedical subjects

Norman Morrison

Publications and source records attributed to Norman Morrison.

11 recordsLinked to original sources

The MGED Ontology: a resource for semantics-based description of microarray experiments.

MOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: Stoeckrt@pcbi.upenn.edu SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Computational Biology↗

Meeting report: eGenomics: Cataloguing our Complete Genome Collection II.

This article summarizes the proceedings of the "eGenomics: Cataloguing our Complete Genome Collection II" workshop held November 10-11, 2005, at the European Bioinformatics Institute. This exploratory workshop, organized by members of the Genomic Standards Consortium (GSC), brought together researchers from the genomic, functional OMICS, and computational biology communities to discuss standardization activities across a range of projects. The workshop proceedings and outcomes are set to help guide the development of the GSC's Minimal Information about a Genome Sequence (MIGS) specification.

Animals↗

Concept of sample in OMICS technology.

Fundamental biological processes can now be studied by applying the full range of OMICS technologies (genomics, transcriptomics, proteomics, metabolomics, and beyond) to the same biological sample. Clearly, it would be desirable if the concept of sample were shared among these technologies, especially as up until the time a biological sample is prepared for use in a specific OMICS assay, its description is inherently technology independent. Sharing a common informatic representation would encourage data sharing (rather than data replication), thereby reducing redundant data capture and the potential for error. This would result in a significant degree of harmonization across different OMICS data standardization activities, a task that is critical if we are to integrate data from these different data sources. Here, we review the current concept of sample in OMICS technologies as it is being dealt with by different OMICS standardization initiatives and discuss the special role that the newly formed Genomic Standards Consortium (GSC) might have to play in this domain.

Animals↗

A strategy capitalizing on synergies: the Reporting Structure for Biological Investigation (RSBI) working group.

In this article we present the Reporting Structure for Biological Investigation (RSBI), a working group under the Microarray Gene Expression Data (MGED) Society umbrella. RSBI brings together several communities to tackle the challenges associated with integrating data and representing complex biological investigations, employing multiple OMICS technologies. Currently, RSBI includes environmental genomics, nutrigenomics and toxicogenomics communities, where independent activities are underway to develop databases and establish data communication standards within their respective domains. The RSBI working group has been conceived as a "single point of focus" for these communities, conforming to general accepted view that duplication and incompatibility should be avoided where possible. This endeavour has aimed to synergize insular solutions into one common terminology between biologically driven standardisation efforts and has also resulted in strong collaborations and shared understanding between those in the technological domain. Through extensive liaisons with many standards efforts, several threads have been woven with the hope that ultimately technology-centered standards and their specific extensions into biological domains of interest will not only stand alone, but will also be able to function together, as interchangeable modules.

Databases, Genetic↗

Annotation of environmental OMICS data: application to the transcriptomics domain.

Researchers working on environmentally relevant organisms, populations, and communities are increasingly turning to the application of OMICS technologies to answer fundamental questions about the natural world, how it changes over time, and how it is influenced by anthropogenic factors. In doing so, the need to capture meta-data that accurately describes the biological "source" material used in such experiments is growing in importance. Here, we provide an overview of the formation of the "Env" community of environmental OMICS researchers and its efforts at considering the meta-data capture needs of those working in environmental OMICS. Specifically, we discuss the development to date of the Env specification, an informal specification including descriptors related to geographic location, environment, organism relationship, and phenotype. We then describe its application to the description of environmental transcriptomic experiments and how we have used it to extend the Minimum Information About a Microarray Experiment (MIAME) data standard to create a domain-specific extension that we have termed MIAME/Env. Finally, we make an open call to the community for participation in the Env Community and its future activities.

Ecology↗

Development of FuGO: an ontology for functional genomics investigations.

The development of the Functional Genomics Investigation Ontology (FuGO) is a collaborative, international effort that will provide a resource for annotating functional genomics investigations, including the study design, protocols and instrumentation used, the data generated and the types of analysis performed on the data. FuGO will contain both terms that are universal to all functional genomics investigations and those that are domain specific. In this way, the ontology will serve as the "semantic glue" to provide a common understanding of data from across these disparate data sources. In addition, FuGO will reference out to existing mature ontologies to avoid the need to duplicate these resources, and will do so in such a way as to enable their ease of use in annotation. This project is in the early stages of development; the paper will describe efforts to initiate the project, the scope and organization of the project, the work accomplished to date, and the challenges encountered, as well as future plans.

Biomedical Research↗

maxdLoad2 and maxdBrowse: standards-compliant tools for microarray experimental annotation, data management and dissemination.

BACKGROUND: maxdLoad2 is a relational database schema and Java application for microarray experimental annotation and storage. It is compliant with all standards for microarray meta-data capture; including the specification of what data should be recorded, extensive use of standard ontologies and support for data exchange formats. The output from maxdLoad2 is of a form acceptable for submission to the ArrayExpress microarray repository at the European Bioinformatics Institute. maxdBrowse is a PHP web-application that makes contents of maxdLoad2 databases accessible via web-browser, the command-line and web-service environments. It thus acts as both a dissemination and data-mining tool. RESULTS: maxdLoad2 presents an easy-to-use interface to an underlying relational database and provides a full complement of facilities for browsing, searching and editing. There is a tree-based visualization of data connectivity and the ability to explore the links between any pair of data elements, irrespective of how many intermediate links lie between them. Its principle novel features are: the flexibility of the meta-data that can be captured, the tools provided for importing data from spreadsheets and other tabular representations, the tools provided for the automatic creation of structured documents, the ability to browse and access the data via web and web-services interfaces. Within maxdLoad2 it is very straightforward to customise the meta-data that is being captured or change the definitions of the meta-data. These meta-data definitions are stored within the database itself allowing client software to connect properly to a modified database without having to be specially configured. The meta-data definitions (configuration file) can also be centralized allowing changes made in response to revisions of standards or terminologies to be propagated to clients without user intervention.maxdBrowse is hosted on a web-server and presents multiple interfaces to the contents of maxd databases. maxdBrowse emulates many of the browse and search features available in the maxdLoad2 application via a web-browser. This allows users who are not familiar with maxdLoad2 to browse and export microarray data from the database for their own analysis. The same browse and search features are also available via command-line and SOAP server interfaces. This both enables scripting of data export for use embedded in data repositories and analysis environments, and allows access to the maxd databases via web-service architectures. CONCLUSION: maxdLoad2 http://www.bioinf.man.ac.uk/microarray/maxd/ and maxdBrowse http://dbk.ch.umist.ac.uk/maxdBrowse are portable and compatible with all common operating systems and major database servers. They provide a powerful, flexible package for annotation of microarray experiments and a convenient dissemination environment. They are available for download and open sourced under the Artistic License.

Data Interpretation, Statistical↗

Chemical effects in biological systems--data dictionary (CEBS-DD): a compendium of terms for the capture and integration of biological study design description, conventional phenotypes, and 'omics data.

A critical component in the design of the Chemical Effects in Biological Systems (CEBS) Knowledgebase is a strategy to capture toxicogenomics study protocols and the toxicity endpoint data (clinical pathology and histopathology). A Study is generally an experiment carried out during a period of time for the purpose of obtaining data, and the Study Design Description captures the methods, timing, and organization of the Study. The CEBS Data Dictionary (CEBS-DD) has been designed to define and organize terms in an attempt to standardize nomenclature needed to describe a toxicogenomics Study in a structured yet intuitive format and provide a flexible means to describe a Study as conceptualized by the investigator. The CEBS-DD will organize and annotate information from a variety of sources, thereby facilitating the capture and display of toxicogenomics data in biological context in CEBS, i.e., associating molecular events detected in highly-parallel data with the toxicology/pathology phenotype as observed in the individual Study Subjects and linked to the experimental treatments. The CEBS-DD has been developed with a focus on acute toxicity studies, but with a design that will permit it to be extended to other areas of toxicology and biology with the addition of domain-specific terms. To illustrate the utility of the CEBS-DD, we present an example of integrating data from two proteomics and transcriptomics studies of the response to acute acetaminophen toxicity (A. N. Heinloth et al., 2004, Toxicol. Sci. 80, 193-202).

Acetaminophen↗

PEDRo: a database for storing, searching and disseminating experimental proteomics data.

BACKGROUND: Proteomics is rapidly evolving into a high-throughput technology, in which substantial and systematic studies are conducted on samples from a wide range of physiological, developmental, or pathological conditions. Reference maps from 2D gels are widely circulated. However, there is, as yet, no formally accepted standard representation to support the sharing of proteomics data, and little systematic dissemination of comprehensive proteomic data sets. RESULTS: This paper describes the design, implementation and use of a Proteome Experimental Data Repository (PEDRo), which makes comprehensive proteomics data sets available for browsing, searching and downloading. It is also serves to extend the debate on the level of detail at which proteomics data should be captured, the sorts of facilities that should be provided by proteome data management systems, and the techniques by which such facilities can be made available. CONCLUSIONS: The PEDRo database provides access to a collection of comprehensive descriptions of experimental data sets in proteomics. Not only are these data sets interesting in and of themselves, they also provide a useful early validation of the PEDRo data model, which has served as a starting point for the ongoing standardisation activity through the Proteome Standards Initiative of the Human Proteome Organisation.

Animals↗

Biological implications of Mycobacterium leprae gene expression during infection.

The genome of Mycobacterium leprae, the etiologic agent of leprosy, has been sequenced and annotated revealing a genome in apparent disarray and in stark contrast to the genome of the related human pathogen, M. tuberculosis. With less than 50% coding capacity of a 3.3-Mb genome and 1,116 pseudogenes, the remaining genes help define the minimal gene set necessary for in vivo survival of this mycobacterial pathogen as well as genes potentially required for infection and pathogenesis seen in leprosy. To identify genes transcribed during infection, we surveyed gene transcripts from M. leprae growing in athymic nude mice using reverse transcriptase-polymerase chain reaction (RT-PCR) and cross-species DNA microarray technologies. Transcripts were detected for 221 open reading frames, which included genes involved in DNA replication, cell division, SecA-dependent protein secretion, energy production, intermediary metabolism, iron transport and storage and genes associated with virulence. These results suggest that M. leprae actively catabolizes fatty acids for energy, produces a large number of secretory proteins, utilizes the full array of sigma factors available, produces several proteins involved in iron transport, storage and regulation in the absence of recognizable genes encoding iron scavengers and transcribes several genes associated with virulence in M. tuberculosis. When transcript levels of 9 of these genes were compared from M. leprae derived from lesions of multibacillary leprosy patients and infected nude mouse foot pad tissue using quantitative real-time RT-PCR, gene transcript levels were comparable for all but one of these genes, supporting the continued use of the foot pad infection model for M. leprae gene expression profiling. Identifying genes associated with growth and survival during infection should lead to a more comprehensive understanding of the ability of M. leprae to cause disease.

Animals↗