PubMed Health⌕ Search

PubMed · 12651726

Gene expression data preprocessing.

Abstract

We present an interactive web tool for preprocessing microarray gene expression data. It analyses the data, suggests the most appropriate transformations and proceeds with them after user agreement. The normal preprocessing steps include scale transformations, management of missing values, replicate handling, flat pattern filtering and pattern standardization and they are required before performing any pattern analysis. The processed data set can be sent to other pattern analysis tools.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J Herrero, R Díaz-Uriarte, J Dopazo. 2003-03-22. Gene expression data preprocessing.. https://doi.org/10.1093/bioinformatics%2Fbtg040

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

The Longhorn Array Database (LAD): an open-source, MIAME compliant implementation of the Stanford Microarray Database (SMD).

BACKGROUND: The power of microarray analysis can be realized only if data is systematically archived and linked to biological annotations as well as analysis algorithms. DESCRIPTION: The Longhorn Array Database (LAD) is a MIAME compliant microarray database that operates on PostgreSQL and Linux. It is a fully open source version of the Stanford Microarray Database (SMD), one of the largest microarray databases. LAD is available at http://www.longhornarraydatabase.org CONCLUSIONS: Our development of LAD provides a simple, free, open, reliable and proven solution for storage and analysis of two-color microarray data.

Database Management Systems↗

Digital extractor: analysis of digital differential display output.

Digital Extractor is a program for the high-throughput processing of data sets derived from digital differential display-based comparisons of EST libraries. These comparisons can be utilized to identify discrete subsets of genes whose expression is restricted to distinct tissue types. The program facilitates these investigations by permitting parallel annotation of genes identified as being differentially expressed.

Database Management Systems↗

Achieving evolvable Web-database bioscience applications using the EAV/CR framework: recent advances.

The EAV/CR framework, designed for database support of rapidly evolving scientific domains, utilizes metadata to facilitate schema maintenance and automatic generation of Web-enabled browsing interfaces to the data. EAV/CR is used in SenseLab, a neuroscience database that is part of the national Human Brain Project. This report describes various enhancements to the framework. These include (1) the ability to create "portals" that present different subsets of the schema to users with a particular research focus, (2) a generic XML-based protocol to assist data extraction and population of the database by external agents, (3) a limited form of ad hoc data query, and (4) semantic descriptors for interclass relationships and links to controlled vocabularies such as the UMLS.

Database Management Systems↗