PubMed Health⌕ Search

PubMed · 11808684

DNA array analysis in a Microsoft Windows environment.

Abstract

Microsoft Windows-based computers have evolved to the point that they provide sufficient computational and visualization power for robust analysis of DNA array data. In fact, smaller laboratories might prefer to carry out some or all of their analyses and visualization in a Windows environment, rather than alternative platforms such as UNIX. We have developed a series of manually executed macros written in Visual Basic for Microsoft Excel spreadsheets, that allows for rapid and comprehensive gene expression data analysis. The first macro assigns gene names to spots on the DNA array and normalizes individual hybridizations by expressing the signal intensity for each gene as a percentage of the sum of all gene intensities. The second macro streamlines statistical consideration of the confidence in individual gene measurements for sets of experimental replicates by calculating probability values with the Student's t test. The third macro introduces a threshold value, calculates expression ratios between experimental conditions, and calculates the standard deviation of the mean of the log ratio values. Selected columns of data are copied by a fourth macro to create a processed data set suitable for entry into a Microsoft Access database. An Access database structure is described that allows simple queries across multiple experiments and export of data into third-party data visualization software packages. These analysis tools can be used in their present form by others working with commercial E. coli membrane arrays, or they may be adapted for use with other systems. The Excel spreadsheets with embedded Visual Basic macros and detailed instructions for their use are available at http://www.ou.edu/microarray.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

T Conway, B Kraus, D L Tucker, D J Smalley, A F Dorman, L McKibben. 2002. DNA array analysis in a Microsoft Windows environment.. https://doi.org/10.2144/02321bc02

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

The INSDC specifications-foundations for a FAIR and global INSDC.

Members of the International Nucleotide Sequence Database Collaboration (INSDC; https://www.insdc.org/) collect, exchange, and preserve comprehensive open nucleotide sequence information and provide tools for its access. The INSDC has stated its commitment to welcoming new members into the collaboration to be more representative of the global community of data and users. To reach this goal, a comprehensive definition of the INSDC data model and minimum requirements for data acceptance have been established. Here we describe the processes used to arrive upon these INSDC Specifications and lay out strategies for their continued upkeep to remain current and relevant. Database URL:  https://www.insdc.org/.

Databases, Nucleic Acid↗

Using expressed sequence tag databases to identify ovarian genes of interest.

GenBank contains 4879 expressed sequence tags (EST) derived from four non-normalized human ovarian cDNA libraries. Of these EST, 2646 are contributors to UniGene clusters and have UniGene numbers. The EST map to 1206 distinct UniGenes. A gene expression profile was established for the human ovary by identifying the abundance of each UniGene cluster and its corresponding annotation. The most highly expressed transcripts were for proteins associated with protein synthesis (ribosomal proteins, elongation factors, thymosins, etc.). However, there are also transcripts for genes of unknown function that are ovary-specific. This ovarian gene expression profile provides useful data for the design of DNA microarrays targeted at ovarian function and highlights novel sequences that warrant further investigation.

Databases, Nucleic Acid↗

Identification of Ugandan HIV type 1 variants with unique patterns of recombination in pol involving subtypes A and D.

Most HIV-1 infections in Uganda are caused by subtypes A and D. The prevalence of recombination and the sites of specific breakpoints between these subtypes have not been reported. HIV-1 pol sequences encoding protease (amino acids 1-99) and reverse transcriptase (amino acids 1-324) from 102 pregnant Ugandan women were analyzed by the Recombinant Identification Program, SimPlot, and examination of phylogenetically informative sites to identify sites of recombination between sequence segments belonging to different subtypes. Thirteen percent (13 of 102) of the pol sequences contained strong evidence of recombination between subtypes A and D. At least nine different patterns of recombination were observed. Five women infected with a recombinant virus transmitted the recombinant virus perinatally. In this population-based study, intersubtype recombinants were common. The large number of different types of pol recombinants identified suggests that recombination occurs readily in the pol region. Perinatal transmission of the recombinant viruses demonstrates their evolutionary stability.

Databases, Nucleic Acid↗