PubMed Health⌕ Search

Biomedical subjects

Zohar Yakhini

Publications and source records attributed to Zohar Yakhini.

At least 19 recordsLinked to original sources

Polycomb-mediated methylation on Lys27 of histone H3 pre-marks genes for de novo methylation in cancer.

Many genes associated with CpG islands undergo de novo methylation in cancer. Studies have suggested that the pattern of this modification may be partially determined by an instructive mechanism that recognizes specifically marked regions of the genome. Using chromatin immunoprecipitation analysis, here we show that genes methylated in cancer cells are specifically packaged with nucleosomes containing histone H3 trimethylated on Lys27. This chromatin mark is established on these unmethylated CpG island genes early in development and then maintained in differentiated cell types by the presence of an EZH2-containing Polycomb complex. In cancer cells, as opposed to normal cells, the presence of this complex brings about the recruitment of DNA methyl transferases, leading to de novo methylation. These results suggest that tumor-specific targeting of de novo methylation is pre-programmed by an established epigenetic system that normally has a role in marking embryonic genes for repression.

Caco-2 Cells↗

Genetic variation in putative regulatory loci controlling gene expression in breast cancer.

Candidate single-nucleotide polymorphisms (SNPs) were analyzed for associations to an unselected whole genome pool of tumor mRNA transcripts in 50 unrelated patients with breast cancer. SNPs were selected from 203 candidate genes of the reactive oxygen species pathway. We describe a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort, which revealed significant associations between subsets of SNPs and transcripts, shedding light on the underlying biology. We identified SNPs in EGF, IL1A, MAPK8, XPC, SOD2, and ALOX12 that are associated with the expression patterns of a significant number of transcripts, indicating the presence of regulatory SNPs in these genes. SNPs were found to act in trans in a total of 115 genes. SNPs in 43 of these 115 genes were found to act both in cis and in trans. Finally, subsets of SNPs that share significantly many common associations with a set of transcripts (biclusters) were identified. The subsets of transcripts that are significantly associated with the same set of SNPs or to a single SNP were shown to be functionally coherent in Gene Ontology and pathway analyses and coexpressed in other independent data sets, suggesting that many of the observed associations are within the same functional pathways. To our knowledge, this article is the first study to correlate SNP genotype data in the germ line with somatic gene expression data in breast tumors. It provides the statistical framework for further genotype expression correlation studies in cancer data sets.

Breast Neoplasms↗

Molecular signatures determining coronary artery and saphenous vein smooth muscle cell phenotypes: distinct responses to stimuli.

OBJECTIVE: Phenotypic differences between vascular smooth muscle cell (VSMC) subtypes lead to diverse pathological processes including atherosclerosis, postangioplasty restenosis and vein graft disease. To better understand the molecular mechanisms underlying functional differences among distinct SMC subtypes, we compared gene expression profiles and functional responses to oxidized low-density lipoprotein (OxLDL) and platelet-derived growth factor (PDGF) between cultured SMCs from human coronary artery (CASM) and saphenous vein (SVSM). METHODS AND RESULTS: OxLDL and PDGF elicited markedly different functional responses and expression profiles between the 2 SMC subtypes. In CASM, OxLDL inhibited cell proliferation and migration and modified gene expression of chemokines (CXCL10, CXCL11 and CXCL12), proinflammatory cytokines (IL-1, IL-6, and IL-18), insulin-like growth factor binding proteins (IGFBPs), and both endothelial and smooth muscle marker genes. In SVSM, OxLDL promoted proliferation partially via IGF1 signaling, activated NF-kappaB and phosphatidylinositol signaling pathways, and upregulated prostaglandin (PG) receptors and synthases. In untreated cells, alpha-chemokines, proinflammatory cytokines, and genes associated with apoptosis, inflammation, and lipid biosynthesis were higher in CASM, whereas some beta-chemokines, metalloproteinase inhibitors, and IGFBPs were higher in SVSM. Interestingly, the basal expression levels of these genes seemed closely related to their responses to OxLDL and PDGF. In summary, our results suggest dramatic differences in gene expression patterns and functional responses to OxLDL and PDGF between venous and arterial SMCs, with venous SMCs having stronger proliferative/migratory responses to stimuli but also higher expression of atheroprotective genes at baseline. CONCLUSIONS: These results reveal molecular signatures that define the distinct phenotypes characteristics of coronary artery and saphenous vein SMC subtypes.

Atherosclerosis↗

Multilocus analysis of SNP and metabolic data within a given pathway.

BACKGROUND: Complex traits, which are under the influence of multiple and possibly interacting genes, have become a subject of new statistical methodological research. One of the greatest challenges facing human geneticists is the identification and characterization of susceptibility genes for common multifactorial diseases and their association to different quantitative phenotypic traits. RESULTS: Two types of data from the same metabolic pathway were used in the analysis: categorical measurements of 18 SNPs; and quantitative measurements of plasma levels of several steroids and their precursors. Using the combinatorial partitioning method we tested various thresholds for each metabolic trait and each individual SNP locus. One SNP in CYP19, 3UTR, two SNPs in CYP1B1 (R48G and A119S) and one in CYP1A1 (T461N) were significantly differently distributed between the high and low level metabolic groups. The leave one out cross validation method showed that 6 SNPs in concert make 65% correct prediction of phenotype. Further we used pattern recognition, computing the p-value by Monte Carlo simulation to identify sets of SNPs and physiological characteristics such as age and weight that contribute to a given metabolic level. Since the SNPs detected by both methods reside either in the same gene (CYP1B1) or in 3 different genes in immediate vicinity on chromosome 15 (CYP19, CYP11 and CYP1A1) we investigated the possibility that they form intragenic and intergenic haplotypes, which may jointly account for a higher activity in the pathway. We identified such haplotypes associated with metabolic levels. CONCLUSION: The methods reported here may enable to study multiple low-penetrance genetic factors that together determine various quantitative phenotypic traits. Our preliminary data suggest that several genes coding for proteins involved in a common pathway, that happen to be located on common chromosomal areas and may form intragenic haplotypes, together account for a higher activity of the whole pathway.

Aged↗

Efficient calculation of interval scores for DNA copy number data analysis.

DNA amplifications and deletions characterize cancer genome and are often related to disease evolution. Microarray-based techniques for measuring these DNA copy-number changes use fluorescence ratios at arrayed DNA elements (BACs, cDNA, or oligonucleotides) to provide signals at high resolution, in terms of genomic locations. These data are then further analyzed to map aberrations and boundaries and identify biologically significant structures. We develop a statistical framework that enables the casting of several DNA copy number data analysis questions as optimization problems over real-valued vectors of signals. The simplest form of the optimization problem seeks to maximize phi(I) = Sigmanu(i)/radical|I| over all subintervals I in the input vector. We present and prove a linear time approximation scheme for this problem, namely, a process with time complexity O (nepsilon(-2)) that outputs an interval for which phi(I) is at least Opt/alpha(epsilon), where Opt is the actual optimum and alpha(epsilon) --> 1 as epsilon --> 0. We further develop practical implementations that improve the performance of the naive quadratic approach by orders of magnitude. We discuss properties of optimal intervals and how they apply to the algorithm performance. We benchmark our algorithms on synthetic as well as publicly available DNA copy number data. We demonstrate the use of these methods for identifying aberrations in single samples as well as common alterations in fixed sets and subsets of breast cancer samples.

Algorithms↗

A high-throughput approach for associating MicroRNAs with their activity conditions.

Plant microRNAs (miRNAs) are short RNA sequences that bind to target mRNAs and change their expression levels by redirecting their stabilities and marking them for cleavage. In Arabidopsis thaliana, microRNAs have been shown to regulate development and are believed to impact expression both under various conditions, such as stress and stimuli, as well as in specific tissue types. We present a high throughput approach for associating between microRNAs and conditions in which they act, using novel statistical and algorithmic techniques. Our new tool, miRNAXpress, at first computes a (binary) matrix T denoting the potential targets of microRNAs. Then, using T and an additional predefined matrix X indicating expression of genes under various conditions, it produces a new matrix that predicts associations between microRNAs and the conditions in which they act. Thus, the program comprises two main modules that work in tandem to compute the desired output. The first is an efficient target prediction engine that predicts mRNA targets of query microRNAs by evaluating the optimal duplex that could be formed between the two: given a short query RNA, a long target RNA, and a predefined energy cut-off threshold, the program finds and reports all putative binding sites of the query RNA in the target RNA with hybridization energy bounded by the predefined threshold. The second module realizes an association operation that is computed by a method which relies on an efficient t-test to compute the associations. The calculation of the matrix of microRNAs and their potential targets is the computationally intensive part of the work done by miRNAXpress, and therefore an efficient algorithm for this portion facilitates the entire process. Thus, the target prediction engine is based on an efficient approximate hybridization search algorithm whose efficiency is the result of utilizing the sparsity of the search space without sacrificing the optimality of the results. The time complexity of this algorithm is almost linear in the size of a sparse set of locations where base-pairs are stacked at a height of three or more. Thus miRNAXpress is a novel tool for associating between microRNAs and the conditions in which they act. We employed it to conduct a study, using the plant Arabidopsis thaliana as our model organism. By applying miRNAXpress to 98 microRNAs and 380 conditions, some biologically interesting and statistically strong relations were discovered. For example, mir159C activity is possibly a factor in the misresponse of nph4 mutants to phototropic stimulations.

Arabidopsis↗

Analysis of SNP-expression association matrices.

High throughput expression profiling and genotyping technologies provide the means to study the genetic determinants of population variation in gene expression variation. In this paper we present a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort. The framework consists of methods to associate transcripts with SNPs affecting their expression, algorithms to detect subsets of transcripts that share significantly many associations with a subset of SNPs, and methods to visualize the identified relations. We apply our framework to SNP-expression data collected from 50 breast cancer patients. Our results demonstrate an overabundance of transcript-SNP associations in this data, and pinpoint SNPs that are potential master regulators of transcription. We also identify several statistically significant transcript-subsets with common putative regulators that fall into well-defined functional categories.

Algorithms↗

Differences in vascular bed disease susceptibility reflect differences in gene expression response to atherogenic stimuli.

Atherosclerosis occurs predominantly in arteries and only rarely in veins. The goal of this study was to test whether differences in the molecular responses of venous and arterial endothelial cells (ECs) to atherosclerotic stimuli might contribute to vascular bed differences in susceptibility to atherosclerosis. We compared gene expression profiles of primary cultured ECs from human saphenous vein (SVEC) and coronary artery (CAEC) exposed to atherogenic stimuli. In addition to identifying differentially expressed genes, we applied statistical analysis of gene ontology and pathway annotation terms to identify signaling differences related to cell type and stimulus. Differential gene expression of untreated venous and arterial endothelial cells yielded 285 genes more highly expressed in untreated SVEC (P<0.005 and fold change >1.5). These genes represented various atherosclerosis-related pathways including responses to proliferation, oxidoreductase activity, antiinflammatory responses, cell growth, and hemostasis functions. Moreover, stimulation with oxidized LDL induced dramatically greater gene expression responses in CAEC compared with SVEC, relating to adhesion, proliferation, and apoptosis pathways. In contrast, interleukin 1beta and tumor necrosis factor alpha activated similar gene expression responses in both CAEC and SVEC. The differences in functional response and gene expression were further validated by an in vitro proliferation assay and in vivo immunostaining of alphabeta-crystallin protein. Our results strongly suggest that different inherent gene expression programs in arterial versus venous endothelial cells contribute to differences in atherosclerotic disease susceptibility.

Atherosclerosis↗

Marek's disease virus Meq transforms chicken cells via the v-Jun transcriptional cascade: a converging transforming pathway for avian oncoviruses.

Marek's disease virus (MDV) is a highly pathogenic and oncogenic herpesvirus of chickens. MDV encodes a basic leucine zipper (bZIP) protein, Meq (MDV EcoQ). The bZIP domain of Meq shares homology with Jun/Fos, whereas the transactivation/repressor domain is entirely different. Increasing evidence suggests that Meq is the oncoprotein of MDV. Direct evidence that Meq transforms chicken cells and the underlying mechanism, however, remain completely unknown. Taking advantage of the DF-1 chicken embryo fibroblast transformation system, a well established model for studying avian sarcoma and leukemia oncogenes, we probed the transformation properties and pathways of Meq. We found that Meq transforms DF-1, with a cell morphology akin to v-Jun and v-Ski transformed cells, and protects DF-1 from apoptosis, and the transformed cells are tumorigenic in chorioallantoic membrane assay. Significantly, using microarray and RT-PCR analyses, we have identified up-regulated genes such as JTAP-1, JAC, and HB-EGF, which belong to the v-Jun transforming pathway. In addition, c-Jun was found to form stable dimers with Meq and colocalize with it in the transformed cells. RNA interference to Meq and c-Jun down-modulated the expression of these genes and reduced the growth of the transformed DF-1, suggesting that Meq transforms chicken cells by pirating the Jun pathway. These data suggest that avian herpesvirus and retrovirus oncogenes use a similar strategy in transformation and oncogenesis.

Animals↗

Pathway analysis of coronary atherosclerosis.

Large-scale gene expression studies provide significant insight into genes differentially regulated in disease processes such as cancer. However, these investigations offer limited understanding of multisystem, multicellular diseases such as atherosclerosis. A systems biology approach that accounts for gene interactions, incorporates nontranscriptionally regulated genes, and integrates prior knowledge offers many advantages. We performed a comprehensive gene level assessment of coronary atherosclerosis using 51 coronary artery segments isolated from the explanted hearts of 22 cardiac transplant patients. After histological grading of vascular segments according to American Heart Association guidelines, isolated RNA was hybridized onto a customized 22-K oligonucleotide microarray, and significance analysis of microarrays and gene ontology analyses were performed to identify significant gene expression profiles. Our studies revealed that loss of differentiated smooth muscle cell gene expression is the primary expression signature of disease progression in atherosclerosis. Furthermore, we provide insight into the severe form of coronary artery disease associated with diabetes, reporting an overabundance of immune and inflammatory signals in diabetics. We present a novel approach to pathway development based on connectivity, determined by language parsing of the published literature, and ranking, determined by the significance of differentially regulated genes in the network. In doing this, we identify highly connected "nexus" genes that are attractive candidates for therapeutic targeting and followup studies. Our use of pathway techniques to study atherosclerosis as an integrated network of gene interactions expands on traditional microarray analysis methods and emphasizes the significant advantages of a systems-based approach to analyzing complex disease.

Adult↗

Multiplexing schemes for generic SNP genotyping assays.

Association studies in populations relate genomic variation among individuals with medical condition. Key to these studies is the development of efficient and affordable genotyping techniques. Generic genotyping assays are independent of the target SNPs and offer great flexibility in the genotyping process. Efficient use of such assays calls for identifying sets of SNPs that can be interrogated in parallel under constraints imposed by the genotyping technology. In this paper, we study problems arising in the design of genotyping experiments using generic assays. Our problem formulation deals with two main factors that affect the genotyping cost: the number of assays used and the number of PCR reactions required for sample preparation. We prove that the resulting computational problems are hard, but provide approximate and heuristic solutions to these problems. Our algorithmic approach is based on recasting the multiplexing problems as partitioning and packing problems on a bipartite graph. We tested our algorithmic approaches on an extensive collection of synthetic data and on data that was simulated using real SNP sequences. Our results show that the algorithms achieve near-optimal designs in many cases and demonstrate the applicability of generic assays to SNP genotyping.

Algorithms↗

Finding approximate tandem repeats in genomic sequences.

An efficient algorithm is presented for detecting approximate tandem repeats in genomic sequences. The algorithm is based on a flexible statistical model which allows a wide range of definitions of approximate tandem repeats. The ideas and methods underlying the algorithm are described and its effectiveness on genomic data is demonstrated.

Algorithms↗

Analysis of SNP-expression association matrices.

High throughput expression profiling and genotyping technologies provide the means to study the genetic determinants of population variation in gene expression variation. In this paper we present a general statistical framework for the simultaneous analysis of gene expression data and SNP genotype data measured for the same cohort. The framework consists of methods to associate transcripts with SNPs affecting their expression, algorithms to detect subsets of transcripts that share significantly many associations with a subset of SNPs, and methods to visualize the identified relations. We apply our framework to SNP-expression data collected from 49 breast cancer patients. Our results demonstrate an overabundance of transcript-SNP associations in this data, and pinpoint SNPs that are potential master regulators of transcription. We also identify several statistically significant transcript-subsets with common putative regulators that fall into well-defined functional categories.

Algorithms↗

Comparative genomic hybridization using oligonucleotide microarrays and total genomic DNA.

Array-based comparative genomic hybridization (CGH) measures copy-number variations at multiple loci simultaneously, providing an important tool for studying cancer and developmental disorders and for developing diagnostic and therapeutic targets. Arrays for CGH based on PCR products representing assemblies of BAC or cDNA clones typically require maintenance, propagation, replication, and verification of large clone sets. Furthermore, it is difficult to control the specificity of the hybridization to the complex sequences that are present in each feature of such arrays. To develop a more robust and flexible platform, we created probe-design methods and assay protocols that make oligonucleotide microarrays synthesized in situ by inkjet technology compatible with array-based comparative genomic hybridization applications employing samples of total genomic DNA. Hybridization of a series of cell lines with variable numbers of X chromosomes to arrays designed for CGH measurements gave median ratios for X-chromosome probes within 6% of the theoretical values (0.5 for XY/XX, 1.0 for XX/XX, 1.4 for XXX/XX, 2.1 for XXXX/XX, and 2.6 for XXXXX/XX). Furthermore, these arrays detected and mapped regions of single-copy losses, homozygous deletions, and amplicons of various sizes in different model systems, including diploid cells with a chromosomal breakpoint that has been mapped and sequenced to a precise nucleotide and tumor cell lines with highly variable regions of gains and losses. Our results demonstrate that oligonucleotide arrays designed for CGH provide a robust and precise platform for detecting chromosomal alterations throughout a genome with high sensitivity even when using full-complexity genomic samples.

Cell Line↗

Towards optimally multiplexed applications of universal arrays.

We study a design and optimization problem that occurs, for example, when single nucleotide polymorphisms (SNPs) are to be genotyped using a universal DNA tag array. The problem of optimizing the universal array to avoid disruptive cross-hybridization between universal components of the system was addressed in previous work. Cross-hybridization can, however, also occur assay specifically, due to unwanted complementarity involving assay-specific components. Here we examine the problem of identifying the most economic experimental configuration of the assay-specific components that avoids cross-hybridization. Our formalization translates this problem into the problem of covering the vertices of one side of a bipartite graph by a minimum number of balanced subgraphs of maximum degree 1. We show that the general problem is NP-complete. However, in the real biological setting, the vertices that need to be covered have degrees bounded by d. We exploit this restriction and develop an O(d)-approximation algorithm for the problem. We also give an O(d)-approximation for a variant of the problem in which the covering subgraphs are required to be vertex disjoint. In addition, we propose a stochastic model for the input data and use it to prove a lower bound on the cover size. We complement our theoretical analysis by implementing two heuristic approaches and testing their performance on synthetic data as well as on simulated SNP data.

Algorithms↗

Novel role for the potent endogenous inotrope apelin in human cardiac dysfunction.

BACKGROUND: Apelin is among the most potent stimulators of cardiac contractility known. However, no physiological or pathological role for apelin-angiotensin receptor-like 1 (APJ) signaling has ever been described. METHODS AND RESULTS: We performed transcriptional profiling using a spotted cDNA microarray with 12 814 unique clones on paired samples of left ventricle obtained before and after placement of a left ventricular assist device in 11 patients. The significance analysis of microarrays and a novel rank consistency score designed to exploit the paired structure of the data confirmed that natriuretic peptides were among the most significantly downregulated genes after offloading. The most significantly upregulated gene was the G-protein-coupled receptor APJ, the specific receptor for apelin. We demonstrate here using immunoassay and immunohistochemical techniques that apelin is localized primarily in the endothelium of the coronary arteries and is found at a higher concentration in cardiac tissue after mechanical offloading. These findings imply an important paracrine signaling pathway in the heart. We additionally extend the clinical significance of this work by reporting for the first time circulating human apelin levels and demonstrating increases in the plasma level of apelin in patients with left ventricular dysfunction. CONCLUSIONS: The apelin-APJ signaling pathway emerges as an important novel mediator of cardiovascular control.

Adolescent↗

Identification of endothelial cell genes by combined database mining and microarray analysis.

Vascular endothelial cells maintain the interface between the systemic circulation and soft tissues and mediate critical processes such as inflammation in a vascular bed-selective fashion. To expand our understanding of the genetic pathways that underlie these specific functions, we have focused on the identification of novel genes that are differentially expressed in all endothelial cells, as well as restricted groups of this cell type. Virtual subtraction was conducted employing gene expression data deposited in public databases and 384 genes identified. These genes were spotted on custom microarrays, along with 288 genes identified through subtraction cloning from TGF-beta-stimulated endothelial cells. Arrays were evaluated with RNA samples representing endothelial cells cultured from four vascular sources and five non-endothelial cell types. These studies identified 64 pan-endothelial markers that were differentially expressed with at least a threefold difference (range 3- to 55-fold). In addition, differences in gene expression profiles among endothelial cells from different vascular beds were identified. Validation of these findings was performed by RNA blot expression studies, and a number of the novel genes were shown to be expressed under angiogenic conditions in the developing mouse embryo. The combined tools of database mining and transcriptional profiling thus provide expanded knowledge of endothelial cell gene expression and endothelial cell biology.

Adult↗

Molecular classification of familial non-BRCA1/BRCA2 breast cancer.

In the decade since their discovery, the two major breast cancer susceptibility genes BRCA1 and BRCA2, have been shown conclusively to be involved in a significant fraction of families segregating breast and ovarian cancer. However, it has become equally clear that a large proportion of families segregating breast cancer alone are not caused by mutations in BRCA1 or BRCA2. Unfortunately, despite intensive effort, the identification of additional breast cancer predisposition genes has so far been unsuccessful, presumably because of genetic heterogeneity, low penetrance, or recessive/polygenic mechanisms. These non-BRCA1/2 breast cancer families (termed BRCAx families) comprise a histopathologically heterogeneous group, further supporting their origin from multiple genetic events. Accordingly, the identification of a method to successfully subdivide BRCAx families into recognizable groups could be of considerable value to further genetic analysis. We have previously shown that global gene expression analysis can identify unique and distinct expression profiles in breast tumors from BRCA1 and BRCA2 mutation carriers. Here we show that gene expression profiling can discover novel classes among BRCAx tumors, and differentiate them from BRCA1 and BRCA2 tumors. Moreover, microarray-based comparative genomic hybridization (CGH) to cDNA arrays revealed specific somatic genetic alterations within the BRCAx subgroups. These findings illustrate that, when gene expression-based classifications are used, BRCAx families can be grouped into homogeneous subsets, thereby potentially increasing the power of conventional genetic analysis.

Adult↗