PubMed Health⌕ Search

Biomedical subjects

Michael Shmoish

Publications and source records attributed to Michael Shmoish.

9 recordsLinked to original sources

An integrative method for accurate comparative genome mapping.

We present MAGIC, an integrative and accurate method for comparative genome mapping. Our method consists of two phases: preprocessing for identifying "maximal similar segments," and mapping for clustering and classifying these segments. MAGIC's main novelty lies in its biologically intuitive clustering approach, which aims towards both calculating reorder-free segments and identifying orthologous segments. In the process, MAGIC efficiently handles ambiguities resulting from duplications that occurred before the speciation of the considered organisms from their most recent common ancestor. We demonstrate both MAGIC's robustness and scalability: the former is asserted with respect to its initial input and with respect to its parameters' values. The latter is asserted by applying MAGIC to distantly related organisms and to large genomes. We compare MAGIC to other comparative mapping methods and provide detailed analysis of the differences between them. Our improvements allow a comprehensive study of the diversity of genetic repertoires resulting from large-scale mutations, such as indels and duplications, including explicitly transposable and phagic elements. The strength of our method is demonstrated by detailed statistics computed for each type of these large-scale mutations. MAGIC enabled us to conduct a comprehensive analysis of the different forces shaping prokaryotic genomes from different clades, and to quantify the importance of novel gene content introduced by horizontal gene transfer relative to gene duplication in bacterial genome evolution. We use these results to investigate the breakpoint distribution in several prokaryotic genomes.

Algorithms↗

Annotation of androgen dependence to human prostate cancer-associated genes by microarray analysis of mouse prostate.

In silico methods and array technologies have identified genes differentially expressed in prostate cancer. Biological functions of the identified genes are often unclear. Considering the biological significance of androgens in prostate cancer, we profiled the prostate transcripts of congenital androgen-deficient mice with or without androgen replacement in vivo using murine gene expression array. In parallel genes differentially expressed in human prostate cancer were identified by Digital Differential Display and the Serial Analysis of Gene Expression. Androgen dependence of the identified genes was then determined by the steady-state mRNA levels of the murine orthologs in response to androgen treatment. The annotation is supported by the finding that some of the androgen target genes have been reported previously with independent experiments.

Androgens↗

GeneTide--Terra Incognita Discovery Endeavor: a new transcriptome focused member of the GeneCards/GeneNote suite of databases.

GeneCards is an automatically mined database of human genes that strives to create, along with its auxiliary databases--GeneLoc, GeneNote and GeneAnnot--the most inclusive resource of gene-centered information of the human genome. GeneTide, the Gene Terra Incognita Discovery Endeavor (http://genecards.weizmann.ac.il/genetide/), the newest addition to this family, is a transcriptome-focused database which aims to enhance GeneCards with additional expressed sequence tag (EST)-based genes. This is achieved by comprehensively mapping >85% of the approximately 5.6 million human ESTs currently available at dbEST to known genes by means of data mining and integration of genomic resources including UniGene, DoTS, AceView and in-house resources. GeneTide thus creates comprehensive links between ESTs and GeneCards genes. Furthermore, groups of unassociated transcripts serve as a basis for defining novel EST-based GeneCards Candidates (EGCs). These EGCs, nearly 25,000 of which were defined in version 0.3 of GeneTide, are further annotated with various parameters, including splicing evidence and expression data extracted from the GeneNote database, to determine their validity as possible de novo genes.

Databases, Genetic↗

Potential photosynthesis gene recombination between Prochlorococcus and Synechococcus via viral intermediates.

Genes (psbA and psbD) encoding for photosynthetically important proteins were recently found in a number of cultured cyanophage genomes. This phenomenon may be a beneficial trait to the viruses or their photosynthetic cyanobacterial hosts, or may represent an untapped pool of genes involved in the formation of the photosynthetic apparatus that are prone to lateral gene transfer. Here we show analyses of psbA genes from uncultured environmental viruses and prophage populations. We observe a statistically significant separation between viral genes and their potential Synechococcus hosts' genes, and statistical analyses under models of codon evolution indicate that the psbA genes of viruses are evolving under levels of purifying selection that are virtually indistinguishable from their hosts. Furthermore, our data also indicate the possible exchange and reshuffling of psbA genes between Synechococcus and Prochlorococcus via phage intermediates. Overall, these observations raise the possibility that marine viruses serve as a potential genetic pool in shaping the evolution of cyanobacterial photosynthesis.

Bacterial Proteins↗

Genome-wide midrange transcription profiles reveal expression level relationships in human tissue specification.

MOTIVATION: Genes are often characterized dichotomously as either housekeeping or single-tissue specific. We conjectured that crucial functional information resides in genes with midrange profiles of expression. RESULTS: To obtain such novel information genome-wide, we have determined the mRNA expression levels for one of the largest hitherto analyzed set of 62 839 probesets in 12 representative normal human tissues. Indeed, when using a newly defined graded tissue specificity index tau, valued between 0 for housekeeping genes and 1 for tissue-specific genes, genes with midrange profiles having 0.15< tau<0.85 were found to constitute >50% of all expression patterns. We developed a binary classification, indicating for every gene the I(B) tissues in which it is overly expressed, and the 12-I(B) tissues in which it shows low expression. The 85 dominant midrange patterns with I(B)=2-11 were found to be bimodally distributed, and to contribute most significantly to the definition of tissue specification dendrograms. Our analyses provide a novel route to infer expression profiles for presumed ancestral nodes in the tissue dendrogram. Such definition has uncovered an unsuspected correlation, whereby de novo enhancement and diminution of gene expression go hand in hand. These findings highlight the importance of gene suppression events, with implications to the course of tissue specification in ontogeny and phylogeny. AVAILABILITY: All data and analyses are publically available at the GeneNote website, http://genecards.weizmann.ac.il/genenote/ and, GEO accession GSE803. CONTACT: doron.lancet@weizmann.ac.il SUPPLEMENTARY INFORMATION: Four tables available at the above site.

Algorithms↗

GeneAnnot: comprehensive two-way linking between oligonucleotide array probesets and GeneCards genes.

MOTIVATION: High density oligonucleotide arrays are usually annotated in a one-to-one fashion, with each probeset assigned to one gene. However, in reality, subsets of oligonucleotides in a probeset may match sequences within more than one gene, potentially leading to misinterpretations. Moreover, a gene is often represented by more than one probeset, and analyzing probe matches at the mRNA level can help one deduce whether these probesets are derived from the same or different splice variants. RESULTS: The GeneAnnot system comprehensively documents the many-to-many relationship between oligonucleotide array probesets and annotated genes in GeneCards. It performs pairwise alignments between the probe sequences and gene transcripts, and assigns sensitivity and specificity scores to each probeset/gene pair. AVAILABILITY: http://genecards.weizmann.ac.il/geneannot/ SUPPLEMENTARY INFORMATION: Program description and statistics http://genecards.weizmann.ac.il/geneannot/DOC/index.html

Algorithms↗

Human Gene-Centric Databases at the Weizmann Institute of Science: GeneCards, UDB, CroW 21 and HORDE.

Recent enhancements and current research in the GeneCards (GC) (http://bioinfo.weizmann.ac.il/cards/) project are described, including the addition of gene expression profiles and integrated gene locations. Also highlighted are the contributions of specialized associated human gene-centric databases developed at the Weizmann Institute. These include the Unified Database (UDB) (http://bioinfo.weizmann.ac.il/udb) for human genome mapping, the human Chromosome 21 database at the Weizmann Insti-tute (CroW 21) (http://bioinfo.weizmann.ac.il/crow21), and the Human Olfactory Receptor Data Explora-torium (HORDE) (http://bioinfo.weizmann.ac.il/HORDE). The synergistic relationships amongst these efforts have positively impacted the quality, quantity and usefulness of the GeneCards gene compendium.

Algorithms↗

GeneAnnot: interfacing GeneCards with high-throughput gene expression compendia.

The interpretation of microarray expression results often includes extensive efforts to identify and annotate the gene representatives immobilised on the arrays. In this paper we describe the usage of our automatic GeneAnnot system, which links between Affymetrix arrays and the rich human gene annotations available in GeneCards. We explain GeneCards search options and results display; elaborate on the presentation of expression information in GeneCards, including both our whole-genome GeneNote project and external expression resources; describe the various parameters and displays used by GeneAnnot to assess the annotation quality and probeset specificity; and show how to search GeneAnnot and GeneNote websites directly.

Data Interpretation, Statistical↗

GeneNote: whole genome expression profiles in normal human tissues.

A novel data set, GeneNote (Gene Normal Tissue Expression), was produced to portray complete gene expression profiles in healthy human tissues using the Affymetrix GeneChip HG-U95 set, which includes 62 839 probe-sets. The hybridization intensities of two replicates were processed and analyzed to yield the complete transcriptome for twelve human tissues. Abundant novel information on tissue specificity provides a baseline for past and future expression studies related to diseases. The data is posted in GeneNote (http://genecards.weizmann.ac.il/genenote/), a widely used compendium of human genes (http://bioinfo.weizmann.ac.il/genecards).

Gene Expression↗