PubMed Health⌕ Search

Biomedical subjects

Barbara J Wold

Publications and source records attributed to Barbara J Wold.

14 recordsLinked to original sources

Connectivity in the yeast cell cycle transcription network: inferences from neural networks.

A current challenge is to develop computational approaches to infer gene network regulatory relationships based on multiple types of large-scale functional genomic data. We find that single-layer feed-forward artificial neural network (ANN) models can effectively discover gene network structure by integrating global in vivo protein:DNA interaction data (ChIP/Array) with genome-wide microarray RNA data. We test this on the yeast cell cycle transcription network, which is composed of several hundred genes with phase-specific RNA outputs. These ANNs were robust to noise in data and to a variety of perturbations. They reliably identified and ranked 10 of 12 known major cell cycle factors at the top of a set of 204, based on a sum-of-squared weights metric. Comparative analysis of motif occurrences among multiple yeast species independently confirmed relationships inferred from ANN weights analysis. ANN models can capitalize on properties of biological gene networks that other kinds of models do not. ANNs naturally take advantage of patterns of absence, as well as presence, of factor binding associated with specific expression output; they are easily subjected to in silico "mutation" to uncover biological redundancies; and they can use the full range of factor binding values. A prominent feature of cell cycle ANNs suggested an analogous property might exist in the biological network. This postulated that "network-local discrimination" occurs when regulatory connections (here between MBF and target genes) are explicitly disfavored in one network module (G2), relative to others and to the class of genes outside the mitotic network. If correct, this predicts that MBF motifs will be significantly depleted from the discriminated class and that the discrimination will persist through evolution. Analysis of distantly related Schizosaccharomyces pombe confirmed this, suggesting that network-local discrimination is real and complements well-known enrichment of MBF sites in G1 class genes.

Artificial Intelligence↗

Spatiometabolic stratification of Shewanella oneidensis biofilms.

Biofilms, or surface-attached microbial communities, are both ubiquitous and resilient in the environment. Although much is known about how biofilms form, develop, and detach, very little is understood about how these events are related to metabolism and its dynamics. It is commonly thought that large subpopulations of cells within biofilms are not actively producing proteins or generating energy and are therefore dead. An alternative hypothesis is that within the growth-inactive domains of biofilms, significant populations of living cells persist and retain the capacity to dynamically regulate their metabolism. To test this, we employed unstable fluorescent reporters to measure growth activity and protein synthesis in vivo over the course of biofilm development and created a quantitative routine to compare domains of activity in independently grown biofilms. Here we report that Shewanella oneidensis biofilm structures reproducibly stratify with respect to growth activity and metabolism as a function of size. Within domains of growth-inactive cells, genes typically upregulated under anaerobic conditions are expressed well after growth has ceased. These findings reveal that, far from being dead, the majority of cells in mature S. oneidensis biofilms have actively turned-on metabolic programs appropriate to their local microenvironment and developmental stage.

Anaerobiosis↗

Mining gene expression data by interpreting principal components.

BACKGROUND: There are many methods for analyzing microarray data that group together genes having similar patterns of expression over all conditions tested. However, in many instances the biologically important goal is to identify relatively small sets of genes that share coherent expression across only some conditions, rather than all or most conditions as required in traditional clustering; e.g. genes that are highly up-regulated and/or down-regulated similarly across only a subset of conditions. Equally important is the need to learn which conditions are the decisive ones in forming such gene sets of interest, and how they relate to diverse conditional covariates, such as disease diagnosis or prognosis. RESULTS: We present a method for automatically identifying such candidate sets of biologically relevant genes using a combination of principal components analysis and information theoretic metrics. To enable easy use of our methods, we have developed a data analysis package that facilitates visualization and subsequent data mining of the independent sources of significant variation present in gene microarray expression datasets (or in any other similarly structured high-dimensional dataset). We applied these tools to two public datasets, and highlight sets of genes most affected by specific subsets of conditions (e.g. tissues, treatments, samples, etc.). Statistically significant associations for highlighted gene sets were shown via global analysis for Gene Ontology term enrichment. Together with covariate associations, the tool provides a basis for building testable hypotheses about the biological or experimental causes of observed variation. CONCLUSION: We provide an unsupervised data mining technique for diverse microarray expression datasets that is distinct from major methods now in routine use. In test uses, this method, based on publicly available gene annotations, appears to identify numerous sets of biologically relevant genes. It has proven especially valuable in instances where there are many diverse conditions (10's to hundreds of different tissues or cell types), a situation in which many clustering and ordering algorithms become problematic. This approach also shows promise in other topic domains such as multi-spectral imaging datasets.

Algorithms↗

Genomic DNA as a general cohybridization standard for ratiometric microarrays.

Feature variability on ratiometric microarrays is accommodated by simultaneous cohybridization of a labeled reference standard with a labeled experimental sample. An optimal reference standard would provide full and equal representation for all array features from a given genome so that it would function on any array, would represent all features with similar signal intensity, and would be highly reproducible-both technically and biologically-from preparation to preparation and laboratory to laboratory. A low cost and a good shelf life are also highly desirable. Finally, providing for straightforward recovery of RNA prevalence information and for integration of data across multiple, initially unrelated studies would be significant advances over current methods. For virtually all ratiometric array studies published to date the reference standard has been some kind of RNA sample assembled from a number of different cell lines, tissues, or experimental time points. These RNA references fall short of the desired universality, uniformity, and reproducibility criteria, which then affect data quality and integration across studies. Also, the various mixed RNA standards cannot be used to derive RNA prevalence information from an experimental sample. In contrast, genomic DNA is a natural choice to meet all the criteria, although it has not yet been widely exploited for eukaryotic array experiments. Principal stumbling blocks have been achieving high enough absolute signals for large mammalian and plant genomes and finding a way to stabilize labeled DNA so that it can be stored and used with ease. This chapter describes two genomic DNA-labeling methods that make it possible to use genomic DNA as a universal microarray cohybridization standard. The indirect labeling method permits production of a large quantity of a stable genomic DNA standard that can then be quality tested and stored frozen. This optimizes experimental consistency and significantly improves ease of use. This chapter also shows that the genomic DNA reference standard can deliver RNA prevalence measurements from ratiometric array platforms.

Animals↗

A mathematical and computational framework for quantitative comparison and integration of large-scale gene expression data.

Analysis of large-scale gene expression studies usually begins with gene clustering. A ubiquitous problem is that different algorithms applied to the same data inevitably give different results, and the differences are often substantial, involving a quarter or more of the genes analyzed. This raises a series of important but nettlesome questions: How are different clustering results related to each other and to the underlying data structure? Is one clustering objectively superior to another? Which differences, if any, are likely candidates to be biologically important? A systematic and quantitative way to address these questions is needed, together with an effective way to integrate and leverage expression results with other kinds of large-scale data and annotations. We developed a mathematical and computational framework to help quantify, compare, visualize and interactively mine clusterings. We show that by coupling confusion matrices with appropriate metrics (linear assignment and normalized mutual information scores), one can quantify and map differences between clusterings. A version of receiver operator characteristic analysis proved effective for quantifying and visualizing cluster quality and overlap. These methods, plus a flexible library of clustering algorithms, can be called from a new expandable set of software tools called CompClust 1.0 (http://woldlab.caltech.edu/compClust/). CompClust also makes it possible to relate expression clustering patterns to DNA sequence motif occurrences, protein-DNA interaction measurements and various kinds of functional annotations. Test analyses used yeast cell cycle data and revealed data structure not obvious under all algorithms. These results were then integrated with transcription motif and global protein-DNA interaction data to identify G1 regulatory modules.

Algorithms↗

Reproducibility, fidelity, and discriminant validity of mRNA amplification for microarray analysis from primary hematopoietic cells.

Analysis of gene expression in clinical samples poses special challenges, including limited RNA availability and poor RNA quality. Quantitative information regarding reliability of RNA amplification methodologies applied to primary cells and representativeness of resulting gene expression profiles is limited. We evaluated four protocols for RNA amplification from peripheral blood mononuclear cells. Results obtained with 100 ng or 10 ng of RNA amplified using two rounds of cDNA synthesis and in vitro transcription were compared with control 2.5-microg RNA samples processed using a single round of in vitro transcription. Samples were hybridized to Affymetrix HG-U133A arrays. Considerable differences in results were obtained with different protocols. The optimal protocol resulted in highly reproducible gene expression profiles from amplified samples (r = 0.98) and good correlation between amplified and control samples (r = 0.94). Using the optimal protocol dissimilarities of gene expression between mononuclear cells from a normal individual and a patient with myelodysplastic syndrome were primarily maintained after amplification compared with controls. We conclude that small variations in methodology introduce considerable distortion of gene expression profiles obtained after RNA amplification from clinical samples and too strong a focus on a very small number of genes picked from an array analysis could be unduly influenced by seemingly acceptable methodologies. However, it is possible to obtain reproducible and representative results using optimized protocols.

Gene Expression Profiling↗

Genomic DNA as a cohybridization standard for mammalian microarray measurements.

A persistent design problem for ratiometric microarray studies is selecting the 'denominator' RNA cohybridization standard. The ideal standard should be readily available, inexpensive, invariant over time and from laboratory to laboratory, and should represent all genes with a uniform signal. RNA references (both commercial 'universal' and experiment--specific types), fall short of these goals. We show here that mouse genomic DNA is a reliable microarray cohybridization standard which can meet these criteria. Genomic DNA was superior in universality of coverage (>98% of genes from a 16,000 feature mouse 70mer microarray) to the Stratagene Universal Mouse Reference RNA standard. Ratios for genes in very low abundance in the Stratagene standard were more unstable with the Stratagene standard than with genomic DNA. Genes with mid-range, and therefore presumably optimal RNA denominator values, showed comparable reproducibility with both standards. Inferred ratios made between two different experimental RNAs using a genomic DNA standard were found to correlate well with companion, directly measured ratios (Spearman correlation coefficient = 0.98). The advantage in array feature coverage of genomic DNA will likely increase as newer generation microarrays include genes which are expressed exclusively in minor tissue or developmental domains that are not represented in mixed tissue RNA standards.

Animals↗

Applicability of tandem affinity purification MudPIT to pathway proteomics in yeast.

A combined multidimensional chromatography-mass spectrometry approach known as "MudPIT" enables rapid identification of proteins that interact with a tagged bait while bypassing some of the problems associated with analysis of polypeptides excised from SDS-polyacrylamide gels. However, the reproducibility, success rate, and applicability of MudPIT to the rapid characterization of dozens of proteins have not been reported. We show here that MudPIT reproducibly identified bona fide partners for budding yeast Gcn5p. Additionally, we successfully applied MudPIT to rapidly screen through a collection of tagged polypeptides to identify new protein interactions. Twenty-five proteins involved in transcription and progression through mitosis were modified with a new tandem affinity purification (TAP) tag. TAP-MudPIT analysis of 22 yeast strains that expressed these tagged proteins uncovered known or likely interacting partners for 21 of the baits, a figure that compares favorably with traditional approaches. The proteins identified here comprised 102 previously known and 279 potential physical interactions. Even for the intensively studied Swi2p/Snf2p, the catalytic subunit of the Swi/Snf chromatin remodeling complex, our analysis uncovered a new interacting protein, Rtt102p. Reciprocal tagging and TAP-MudPIT analysis of Rtt102p revealed subunits of both the Swi/Snf and RSC complexes, identifying Rtt102p as a common interactor with, and possible integral component of, these chromatin remodeling machines. Our experience indicates it is feasible for an investigator working with a single ion trap instrument in a conventional molecular/cellular biology laboratory to carry out proteomic characterization of a pathway, organelle, or process (i.e. "pathway proteomics") by systematic application of TAP-MudPIT.

Affinity Labels↗

p62 overexpression in breast tumors and regulation by prostate-derived Ets factor in breast cancer cells.

p62 is a multifunctional cytoplasmic protein able to noncovalently bind ubiquitin and several signaling proteins, suggesting a regulatory role connected to the ubiquitin-proteasome pathway. No studies to date have linked p62 protein expression with pathological states. Here we demonstrate the overabundance of p62 protein in malignant breast tissue relative to normal breast tissue. The proteasome inhibitor PSI increased p62 mRNA and protein; however, PSI treatment of breast epithelial cells transfected with the p62 promoter did not affect promoter activity. High levels of prostate-derived Ets factor (PDEF) mRNA have been identified in breast cancer compared to normal breast. Only the PSA and maspin promoters have been identified as targets of this transcription factor. Here we show that PDEF stimulates the p62 promoter through at least two sites, and likely acts as a coactivator. PSI treatment abrogates the PDEF-stimulated increase of p62 promoter activity by 50%. Thus, multiple mechanisms for the induction of p62 exist. We conclude that (1) p62 protein is overexpressed in breast cancer; (2) p62 mRNA and protein increase in response to PSI, with no change of basal promoter activity; (3) PDEF upregulates p62 promoter activity through at least two sites; and (4) PSI downregulates PDEF-induced p62 promoter activation through one of these sites.

Acetylcysteine↗

Cellerator: extending a computer algebra system to include biochemical arrows for signal transduction simulations.

Cellerator describes single and multi-cellular signal transduction networks (STN) with a compact, optionally palette-driven, arrow-based notation to represent biochemical reactions and transcriptional activation. Multi-compartment systems are represented as graphs with STNs embedded in each node. Interactions include mass-action, enzymatic, allosteric and connectionist models. Reactions are translated into differential equations and can be solved numerically to generate predictive time courses or output as systems of equations that can be read by other programs. Cellerator simulations are fully extensible and portable to any operating system that supports Mathematica, and can be indefinitely nested within larger data structures to produce highly scaleable models.

Computer Graphics↗

Significance and statistical errors in the analysis of DNA microarray data.

DNA microarrays are important devices for high throughput measurements of gene expression, but no rational foundation has been established for understanding the sources of within-chip statistical error. We designed a specialized chip and protocol to investigate the distribution and magnitude of within-chip errors and discovered that, as expected from theoretical expectations, measurement errors follow a Lorentzian-like distribution, which explains the widely observed but unexplained ill-reproducibility in microarray data. Using this specially designed chip, we examined a data set of repeated measurements to extract estimates of the distribution and magnitude of statistical errors in DNA microarray measurements. Using the common "ratio of medians" method, we find that the measurements follow a Lorentzian-like distribution, which is problematic for subsequent analysis. We show that a method of analysis dubbed "median of ratios" yields a more Gaussian-like distribution of errors. Finally, we show that the bootstrap algorithm can be used to extract the best estimates of the error in the measurement. Quantifying the statistical error in such measurements has important applications for estimating significance levels, clustering algorithms, and process optimization.

Algorithms↗

New computational approaches for analysis of cis-regulatory networks.

The investigation and modeling of gene regulatory networks requires computational tools specific to the task. We present several locally developed software tools that have been used in support of our ongoing research into the embryogenesis of the sea urchin. These tools are especially well suited to iterative refinement of models through experimental and computational investigation. They include: BioArray, a macroarray spot processing program; SUGAR, a system to display and correlate large-BAC sequence analyses; SeqComp and FamilyRelations, programs for comparative sequence analysis; and NetBuilder, an environment for creating and analyzing models of gene networks. We also present an overview of the process used to build our model of the Strongylocentrotus purpuratus endomesoderm gene network. Several of the tools discussed in this paper are still in active development and some are available as open source.

Chromosomes, Artificial, Bacterial↗

Identification and confirmation of a module of coexpressed genes.

We synthesize a large gene expression data set using dbEST and UniGene. We use guilt-by-association (GBA) to analyze this data set and identify coexpressed genes. One module, or group of genes, was found to be coexpressed mainly in tissue extracted from breast and ovarian cancers, but also found in tissue from lung cancers, brain cancers, and bone marrow. This module contains at least six members that are believed to be involved in either transcritional regulation (PDEF, H2AFO, NUCKS) or the ubiquitin proteasome pathway (PSMD7, SQSTM1, FLJ10111). We confirm these observations of coexpression by real-time RT-PCR analysis of mRNA extracted from four model breast epithelial cell lines.

Adult↗