PubMed Health⌕ Search

Biomedical subjects

Voichita D Marinescu

Publications and source records attributed to Voichita D Marinescu.

6 recordsLinked to original sources

START: an automated tool for serial analysis of chromatin occupancy data.

UNLABELLED: The serial analysis of chromatin occupancy technique (SACO) promises to become a widely used method for the unbiased genome-wide experimental identification of loci bound by a transcription factor of interest. We describe the first web-based automatic tool, termed sequence tag analysis and reporting tool (START), for processing SACO data generated by experiments performed for the yeast, fruit fly, mouse, rat or human genomes. The program uses as input sequences of inserts from a SACO library from which it extracts all SACO tags, maps them to genomic locations and annotates them. START returns detailed information about these tags including the genes, the genomic elements and the miRNA precursors found in their vicinity, and makes use of the MAPPER database to identify putative transcription factor binding sites located close to the tags. AVAILABILITY: The program is available at http://bio.chip.org/start/. SUPPLEMENTARY INFORMATION: SUPPLEMENTARY INFORMATION is available at http://bio.chip.org/doc/start/START-supplementary.pdf

Artificial Intelligence↗

MAPPER: a search engine for the computational identification of putative transcription factor binding sites in multiple genomes.

BACKGROUND: Cis-regulatory modules are combinations of regulatory elements occurring in close proximity to each other that control the spatial and temporal expression of genes. The ability to identify them in a genome-wide manner depends on the availability of accurate models and of search methods able to detect putative regulatory elements with enhanced sensitivity and specificity. RESULTS: We describe the implementation of a search method for putative transcription factor binding sites (TFBSs) based on hidden Markov models built from alignments of known sites. We built 1,079 models of TFBSs using experimentally determined sequence alignments of sites provided by the TRANSFAC and JASPAR databases and used them to scan sequences of the human, mouse, fly, worm and yeast genomes. In several cases tested the method identified correctly experimentally characterized sites, with better specificity and sensitivity than other similar computational methods. Moreover, a large-scale comparison using synthetic data showed that in the majority of cases our method performed significantly better than a nucleotide weight matrix-based method. CONCLUSION: The search engine, available at http://mapper.chip.org, allows the identification, visualization and selection of putative TFBSs occurring in the promoter or other regions of a gene from the human, mouse, fly, worm and yeast genomes. In addition it allows the user to upload a sequence to query and to build a model by supplying a multiple sequence alignment of binding sites for a transcription factor of interest. Due to its extensive database of models, powerful search engine and flexible interface, MAPPER represents an effective resource for the large-scale computational analysis of transcriptional regulation.

Algorithms↗

The MAPPER database: a multi-genome catalog of putative transcription factor binding sites.

We describe a comprehensive map of putative transcription factor binding sites (TFBSs) across multiple genomes created using a search method that relies on hidden Markov models built from experimentally determined TFBSs. Using the information in the TRANSFAC and JASPAR databases, we built 1134 models for TFBSs and used them to scan regions 10 kb upstream of the start of the transcript for all known genes in the human, mouse and Drosophila melanogaster genomes. The results, together with homology information on clusters of ortholog genes across the three genomes, were used to create a multi-organism catalog of annotated TFBSs. The catalog can be queried through a web interface accessible at http://bio.chip.org/mapper that allows the identification, visualization and selection of TFBSs occurring in the promoter of a gene of interest and also the common factors predicted to bind across the cluster of orthologs that includes that gene. Alternatively, the interface allows the user to retrieve binding sites for a single transcription factor of interest in a single gene or in all genes of the human, mouse or fruit fly genomes.

Animals↗

Bayesian approach to discovering pathogenic SNPs in conserved protein domains.

The success rate of association studies can be improved by selecting better genetic markers for genotyping or by providing better leads for identifying pathogenic single nucleotide polymorphisms (SNPs) in the regions of linkage disequilibrium with positive disease associations. We have developed a novel algorithm to predict pathogenic single amino acid changes, either nonsynonymous SNPs (nsSNPs) or missense mutations, in conserved protein domains. Using a Bayesian framework, we found that the probability of a microbial missense mutation causing a significant change in phenotype depended on how much difference it made in several phylogenetic, biochemical, and structural features related to the single amino acid substitution. We tested our model on pathogenic allelic variants (missense mutations or nsSNPs) included in OMIM, and on the other nsSNPs in the same genes (from dbSNP) as the nonpathogenic variants. As a result, our model predicted pathogenic variants with a 10% false-positive rate. The high specificity of our prediction algorithm should make it valuable in genetic association studies aimed at identifying pathogenic SNPs.

Algorithms↗

Expression profiling and identification of novel genes involved in myogenic differentiation.

Skeletal muscle differentiation is a complex, highly coordinated process that relies on precise temporal gene expression patterns. To better understand this cascade of transcriptional events, we used expression profiling to analyze gene expression in a 12-day time course of differentiating C2C12 myoblasts. Cluster analysis specific for time-ordered microarray experiments classified 2895 genes and ESTs with variable expression levels between proliferating and differentiating cells into 22 clusters with distinct expression patterns during myogenesis. Expression patterns for several known and novel genes were independently confirmed by real-time quantitative RT-PCR and/or Western blotting and immunofluorescence. MyoD and MEF family members exhibited unique expression kinetics that were highly coordinated with cell-cycle withdrawal regulators. Among genes with peak expression levels during cell cycle withdrawal were Vcam1, Itgb3, Itga5, Vcl, as well as Ptger4, a gene not previously associated with the process of myogenesis. One interesting uncharacterized transcript that is highly induced during myogenesis encodes several immunoglobulin repeats with sequence similarity to titin, a large sarcomeric protein. These data sets identify many additional uncharacterized transcripts that may play important functions in muscle cell proliferation and differentiation and provide a baseline for comparison with C2C12 cells expressing various mutant genes involved in myopathic disorders.

Animals↗

Bayesian estimation of transcript levels using a general model of array measurement noise.

Gene arrays demonstrate a promising ability to characterize expression levels across the entire genome but suffer from significant levels of measurement noise. We present a rigorous new approach to estimate transcript levels and ratios from one or more gene array experiments, given a model of measurement noise and available prior information. The Bayesian estimation of array measurements (BEAM) technique provides a principled method to identify changes in expression level, combine repeated measurements, or deal with negative expression level measurements. BEAM is more flexible than existing techniques, because it does not assume a specific functional form for noise and prior models. Instead, it relies on computational techniques that apply to a broad range of models. We use Affymetrix yeast chip data to illustrate the process of developing accurate noise and prior models from existing experimental data. The resulting noise model includes novel features such as heavy-tailed additive noise and a gene-specific bias term. We also verify that the resulting noise and prior models fit data from an Affymetrix human chip set.

Bayes Theorem↗