PubMed Health⌕ Search

Biomedical subjects

Nir Friedman

Publications and source records attributed to Nir Friedman.

15 recordsLinked to original sources

A module map showing conditional activity of expression modules in cancer.

DNA microarrays are widely used to study changes in gene expression in tumors, but such studies are typically system-specific and do not address the commonalities and variations between different types of tumor. Here we present an integrated analysis of 1,975 published microarrays spanning 22 tumor types. We describe expression profiles in different tumors in terms of the behavior of modules, sets of genes that act in concert to carry out a specific function. Using a simple unified analysis, we extract modules and characterize gene-expression profiles in tumors as a combination of activated and deactivated modules. Activation of some modules is specific to particular types of tumor; for example, a growth-inhibitory module is specifically repressed in acute lymphoblastic leukemias and may underlie the deregulated proliferation in these cancers. Other modules are shared across a diverse set of clinical conditions, suggestive of common tumor progression mechanisms. For example, the bone osteoblastic module spans a variety of tumor types and includes both secreted growth factors and their receptors. Our findings suggest that there is a single mechanism for both primary tumor proliferation and metastasis to bone. Our analysis presents multiple research directions for diagnostic, prognostic and therapeutic studies.

Databases, Genetic↗

Sfp1 is a stress- and nutrient-sensitive regulator of ribosomal protein gene expression.

Yeast cells modulate their protein synthesis capacity in response to physiological needs through the transcriptional control of ribosomal protein (RP) genes. Here we demonstrate that the transcription factor Sfp1, previously shown to play a role in the control of cell size, regulates RP gene expression in response to nutrients and stress. Under optimal growth conditions, Sfp1 is localized to the nucleus, bound to the promoters of RP genes, and helps promote RP gene expression. In response to inhibition of target of rapamycin (TOR) signaling, stress, or changes in nutrient availability, Sfp1 is released from RP gene promoters and leaves the nucleus, and RP gene transcription is down-regulated. Additionally, cells lacking Sfp1 fail to appropriately modulate RP gene expression in response to environmental cues. We conclude that Sfp1 integrates information from nutrient- and stress-responsive signaling pathways to help control RP gene expression.

Cyclic AMP-Dependent Protein Kinases↗

Modulation of DNA conformations through the formation of alternative high-order HU-DNA complexes.

HU is an abundant, highly conserved protein associated with the bacterial chromosome. It belongs to a small class of proteins that includes the eukaryotic proteins TBP, SRY, HMG-I and LEF-I, which bind to DNA non-specifically at the minor groove. HU plays important roles as an accessory architectural factor in a variety of bacterial cellular processes such as DNA compaction, replication, transposition, recombination and gene regulation. In an attempt to unravel the role this protein plays in shaping nucleoid structure, we have carried out fluorescence resonance energy transfer measurements of HU-DNA oligonucleotide complexes, both at the ensemble and single-pair levels. Our results provide direct experimental evidence for concerted DNA bending by HU, and the abrogation of this effect at HU to DNA ratios above about one HU dimer per 10-12 bp. These findings support a model in which a number of HU molecules form an ordered helical scaffold with DNA lying in the periphery. The abrogation of these nucleosome-like structures for high HU to DNA ratios suggests a unique role for HU in the dynamic modulation of bacterial nucleoid structure.

Bacterial Proteins↗

Stress-related genomic responses during the course of heat acclimation and its association with ischemic-reperfusion cross-tolerance.

Acclimation to heat is a biphasic process involving a transient perturbed phase followed by a long lasting period during which acclimatory homeostasis is developed. In this investigation, we used cDNA stress microarray (Clontech Laboratory) to characterize the stress-related genomic response during the course of heat acclimation and to test the hypotheses that 1) heat acclimation influences the threshold of activation of protective molecular signaling, and 2) heat-acclimation-mediated ischemic-reperfusion (I/R) protection is coupled with reprogrammed gene expression leading to altered capacity or responsiveness of protective-signaling pathways shared by heat and I/R cytoprotective systems. Rats were acclimated at 34 degrees C for 0, 2, and 30 days. 32P-labeled RNA samples prepared from the left ventricles of rats before and after subjection to heat stress (HS; 2 h, 41 degrees C) or after I/R insult (ischemia: 75%, 45 min; reperfusion: 30 min) were hybridized onto the array membranes. Confirmatory RT-PCR of selected genes conducted on samples taken at 0, 30, and 60 min after HS or total ischemia was used to assess the promptness of the transcriptional response. Cluster analysis of the expressed genes indicated that acclimation involves a "two-tier" defense strategy: an immediate transient response peaking at the initial acclimating phase to maintain DNA and cellular integrity, and a sustained response, correlated with slowly developed adaptive, long-lasting cytoprotective signaling networks involving genes encoding proteins that are essential for the heat-shock response, antiapoptosis, and antioxidation. Gene activation was stress specific. Faster activation and suppression of signaling pathways shared by HS and I/R stressors probably contribute to heat-acclimation I/R cross-tolerance.

Acclimatization↗

Inferring cellular networks using probabilistic graphical models.

High-throughput genome-wide molecular assays, which probe cellular networks from different perspectives, have become central to molecular biology. Probabilistic graphical models are useful for extracting meaningful biological insights from the resulting data sets. These models provide a concise representation of complex cellular networks by composing simpler submodels. Procedures based on well-understood principles for inferring such models from data facilitate a model-based methodology for analysis and discovery. This methodology and its capabilities are illustrated by several recent applications to gene expression data.

Bayes Theorem↗

Comparative analysis of algorithms for signal quantitation from oligonucleotide microarrays.

MOTIVATION: Recent years' exponential increase in DNA microarrays experiments has motivated the development of many signal quantitation (SQ) algorithms. These algorithms perform various transformations on the actual measurements aimed to enable researchers to compare readings of different genes quantitatively within one experiment and across separate experiments. However, it is relatively unclear whether there is a 'best' algorithm to quantitate microarray data. The ability to compare and assess such algorithms is crucial for any downstream analysis. In this work, we suggest a methodology for comparing different signal quantitation algorithms for gene expression data. Our aim is to enable researchers to compare the effect of different SQ algorithms on the specific dataset they are dealing with. We combine two kinds of tests to assess the effect of an SQ algorithm in terms of signal to noise ratio. To assess noise, we exploit redundancy within the experimental dataset to test the variability of a given SQ algorithm output. For the effect of the SQ on the signal we evaluate the overabundance of differentially expressed genes using various statistical significance tests. RESULTS: We demonstrate our analysis approach with three SQ algorithms for oligonucleotide microarrays. We compare the results of using the dChip software and the RMAExpress software to the ones obtained by using the standard Affymetrix MAS5 on a dataset containing pairs of repeated hybridizations. Our analysis suggests that dChip is more robust and stable than the MAS5 tools for about 60% of the genes while RMAExpress is able to achieve an even greater improvement in terms of signal to noise, for more than 95% of the genes.

Algorithms↗

Blood transcriptional signatures of multiple sclerosis: unique gene expression of disease activity.

Multiple sclerosis (MS) is a central nervous system disease with an unpredictable course and outcome. Peripheral blood mononuclear cells (PBMCs) are involved in the disease pathogenesis and induce active demyelination. Using oligonucleotide microarrays, we identified a statistically significant transcriptional signature of 1,109 genes in PBMCs from 26 MS patients, irrespective of disease activation state or immunomodulatory treatment. This signature contains genes that implicate underlying processes involved in MS pathogenesis including T-cell activation and expansion, inflammation, and apoptosis. Another transcriptional signature of 721 genes involved in cellular recruitment, epitope spreading, and escape from regulatory immune surveillance identified MS patients in acute relapse compared with remission. Our results offer new opportunity for understanding the mechanisms involved in MS and indicate that gene expression patterns in PBMCs contain information about a remote-target disease process that may be useful for diagnosis and future tailoring of therapeutic strategies for MS.

Adjuvants, Immunologic↗

Efficient exact p-value computation for small sample, sparse, and surprising categorical data.

A major obstacle in applying various hypothesis testing procedures to datasets in bioinformatics is the computation of ensuing p-values. In this paper, we define a generic branch-and-bound approach to efficient exact p-value computation and enumerate the required conditions for successful application. Explicit procedures are developed for the entire Cressie-Read family of statistics, which includes the widely used Pearson and likelihood ratio statistics in a one-way frequency table goodness-of-fit test. This new formulation constitutes a first practical exact improvement over the exhaustive enumeration performed by existing statistical software. The general techniques we develop to exploit the convexity of many statistics are also shown to carry over to contingency table tests, suggesting that they are readily extendible to other tests and test statistics of interest. Our empirical results demonstrate a speed-up of orders of magnitude over the exhaustive computation, significantly extending the practical range for performing exact tests. We also show that the relative speed-up gain increases as the null hypothesis becomes sparser, that computation precision increases with increase in speed-up, and that computation time is very moderately affected by the magnitude of the computed p-value. These qualities make our algorithm especially appealing in the regimes of small samples, sparse null distributions, and rare events, compared to the alternative asymptotic approximations and Monte Carlo samplers. We discuss several established bioinformatics applications, where small sample size, small expected counts in one or more categories (sparseness), and very small p-values do occur. Our computational framework could be applied in these, and similar cases, to improve performance.

Computational Biology↗

Module networks: identifying regulatory modules and their condition-specific regulators from gene expression data.

Much of a cell's activity is organized as a network of interacting modules: sets of genes coregulated to respond to different conditions. We present a probabilistic method for identifying regulatory modules from gene expression data. Our procedure identifies modules of coregulated genes, their regulators and the conditions under which regulation occurs, generating testable hypotheses in the form 'regulator X regulates module Y under conditions W'. We applied the method to a Saccharomyces cerevisiae expression data set, showing its ability to identify functionally coherent modules and their correct regulators. We present microarray experiments supporting three novel predictions, suggesting regulatory roles for previously uncharacterized proteins.

Algorithms↗

Human and porcine early kidney precursors as a new source for transplantation.

Kidney transplantation has been one of the major medical advances of the past 30 years. However, tissue availability remains a major obstacle. This can potentially be overcome by the use of undifferentiated or partially developed kidney precursor cells derived from early embryos and fetal tissue. Here, transplantation in mice reveals the earliest gestational time point at which kidney precursor cells, of both human and pig origin, differentiate into functional nephrons and not into other, non-renal professional cell types. Moreover, successful organogenesis is achieved when using the early kidney precursors, but not later-gestation kidneys. The formed, miniature kidneys are functional as evidenced by the dilute urine they produce. In addition, decreased immunogenicity of the transplants of early human and pig kidney precursors compared with adult kidney transplants is demonstrated in vivo. Our data pinpoint a window of human and pig kidney organogenesis that may be optimal for transplantation in humans.

Adult↗

Context-specific Bayesian clustering for gene expression data.

The recent growth in genomic data and measurements of genome-wide expression patterns allows us to apply computational tools to examine gene regulation by transcription factors. In this work, we present a class of mathematical models that help in understanding the connections between transcription factors and functional classes of genes based on genetic and genomic data. Such a model represents the joint distribution of transcription factor binding sites and of expression levels of a gene in a unified probabilistic model. Learning a combined probability model of binding sites and expression patterns enables us to improve the clustering of the genes based on the discovery of putative binding sites and to detect which binding sites and experiments best characterize a cluster. To learn such models from data, we introduce a new search method that rapidly learns a model according to a Bayesian score. We evaluate our method on synthetic data as well as on real life data and analyze the biological insights it provides. Finally, we demonstrate the applicability of the method to other data analysis problems in gene expression data.

Bayes Theorem↗

A structural EM algorithm for phylogenetic inference.

A central task in the study of molecular evolution is the reconstruction of a phylogenetic tree from sequences of current-day taxa. The most established approach to tree reconstruction is maximum likelihood (ML) analysis. Unfortunately, searching for the maximum likelihood phylogenetic tree is computationally prohibitive for large data sets. In this paper, we describe a new algorithm that uses Structural Expectation Maximization (EM) for learning maximum likelihood phylogenetic trees. This algorithm is similar to the standard EM method for edge-length estimation, except that during iterations of the Structural EM algorithm the topology is improved as well as the edge length. Our algorithm performs iterations of two steps. In the E-step, we use the current tree topology and edge lengths to compute expected sufficient statistics, which summarize the data. In the M-Step, we search for a topology that maximizes the likelihood with respect to these expected sufficient statistics. We show that searching for better topologies inside the M-step can be done efficiently, as opposed to standard methods for topology search. We prove that each iteration of this procedure increases the likelihood of the topology, and thus the procedure must converge. This convergence point, however, can be a suboptimal one. To escape from such "local optima," we further enhance our basic EM procedure by incorporating moves in the flavor of simulated annealing. We evaluate these new algorithms on both synthetic and real sequence data and show that for protein sequences even our basic algorithm finds more plausible trees than existing methods for searching maximum likelihood phylogenies. Furthermore, our algorithms are dramatically faster than such methods, enabling, for the first time, phylogenetic analysis of large protein data sets in the maximum likelihood framework.

Algorithms↗

A branch-and-bound algorithm for the inference of ancestral amino-acid sequences when the replacement rate varies among sites: Application to the evolution of five gene families.

MOTIVATION: We developed an algorithm to reconstruct ancestral sequences, taking into account the rate variation among sites of the protein sequences. Our algorithm maximizes the joint probability of the ancestral sequences, assuming that the rate is gamma distributed among sites. Our algorithm probably finds the global maximum. The use of 'joint' reconstruction is motivated by studies that use the sequences at all the internal nodes in a phylogenetic tree, such as, for instance, the inference of patterns of amino-acid replacement, or tracing the biochemical changes that occurred during the evolution of a given protein family. RESULTS: We give an algorithm that guarantees finding the global maximum. The efficient search method makes our method applicable to datasets with large number sequences. We analyze ancestral sequences of five gene families, exploring the effect of the amount of among-site-rate-variation, and the degree of sequence divergence on the resulting ancestral states. AVAILABILITY AND SUPPLEMENTARY INFORMATION: http://evolu3.ism.ac.jp/~tal/ CONTACT: tal@ism.ac.jp

Algorithms↗

Practical approaches to analyzing results of microarray experiments.

Microarray technology is rapidly becoming a standard laboratory technique. The main challenges related to the successful implementation of the technology are analysis-related. In this article we provide a practically oriented review focusing on methods for analysis of large-scale gene expression data in the research laboratory. We describe the various common clustering methods and outline our approach to using them. We discuss methods for scoring genes for their relevance, focusing on the statistical meaning of microarray results, especially with regard to the problem of multiple testing. We also deal with the problem of adding biologic meaning to the results of microarray experiments and describe advanced tools that represent different but valid directions in providing automated solutions to this problem. The tools and approaches described and discussed here should provide the reader with a preliminary understanding of the analysis of the results of microarray experiments. The practical focus of this review should remove the mystery behind the analysis of microarray experiments, thus leading to more productive and efficient use of the technology.

Fibroblasts↗