PubMed Health⌕ Search

Biomedical subjects

Ivo Grosse

Publications and source records attributed to Ivo Grosse.

13 recordsLinked to original sources

An expanded codebook of human transcription factor DNA-binding specificity.

Gene expression is regulated by transcription factors (TFs), which recognize specific DNA sequence motifs. Several hundred putative human TFs, identified mainly by an apparent DNA-binding domain, lack known binding motifs1. Furthermore, even for well-characterized TFs, it remains controversial the degree to which motifs accurately reflect binding sites in living cells2. Here we describe a systematic effort ('Codebook') to determine the sequence specificity of 332 putative and poorly characterized human TFs. More than 4,000 independent experiments, encompassing multiple in vitro and in vivo assays, produced motifs for just over half (177; 53%) of the TFs, of which most are associated with only a single protein. These results extend the vocabulary of sequence recognition encoded by human TFs by around 130 distinct motifs. Moreover, binding motifs identified in vitro are strongly enriched in cellular binding sites. Collectively, the data reveal tens of thousands of previously unknown, conserved and direct TF-binding sites across the human genome. These sites are concentrated in promoter regions and are predictive of gene expression. In summary, this new codebook provides an important step forward in decoding the human genome.

Humans↗

A 1,000-loci transcript map of the barley genome: new anchoring points for integrative grass genomics.

An integrated barley transcript map (consensus map) comprising 1,032 expressed sequence tag (EST)-based markers (total 1,055 loci: 607 RFLP, 190 SSR, and 258 SNP), and 200 anchor markers from previously published data, has been generated by mapping in three doubled haploid (DH) populations. Between 107 and 179 EST-based markers were allocated to the seven individual barley linkage groups. The map covers 1118.3 cM with individual linkage groups ranging from 130 cM (chromosome 4H) to 199 cM (chromosome 3H), yielding an average marker interval distance of 0.9 cM. 475 EST-based markers showed a syntenic organisation to known colinear linkage groups of the rice genome, providing an extended insight into the status of barley/rice genome colinearity as well as ancient genome duplications predating the divergence of rice and barley. The presented barley transcript map is a valuable resource for targeted marker saturation and identification of candidate genes at agronomically important loci. It provides new anchor points for detailed studies in comparative grass genomics and will support future attempts towards the integration of genetic and physical mapping information.

Chromosome Mapping↗

Meta-All: a system for managing metabolic pathway information.

BACKGROUND: Many attempts are being made to understand biological subjects at a systems level. A major resource for these approaches are biological databases, storing manifold information about DNA, RNA and protein sequences including their functional and structural motifs, molecular markers, mRNA expression levels, metabolite concentrations, protein-protein interactions, phenotypic traits or taxonomic relationships. The use of these databases is often hampered by the fact that they are designed for special application areas and thus lack universality. Databases on metabolic pathways, which provide an increasingly important foundation for many analyses of biochemical processes at a systems level, are no exception from the rule. Data stored in central databases such as KEGG, BRENDA or SABIO-RK is often limited to read-only access. If experimentalists want to store their own data, possibly still under investigation, there are two possibilities. They can either develop their own information system for managing that own data, which is very time-consuming and costly, or they can try to store their data in existing systems, which is often restricted. Hence, an out-of-the-box information system for managing metabolic pathway data is needed. RESULTS: We have designed META-ALL, an information system that allows the management of metabolic pathways, including reaction kinetics, detailed locations, environmental factors and taxonomic information. Data can be stored together with quality tags and in different parallel versions. META-ALL uses Oracle DBMS and Oracle Application Express. We provide the META-ALL information system for download and use. In this paper, we describe the database structure and give information about the tools for submitting and accessing the data. As a first application of META-ALL, we show how the information contained in a detailed kinetic model can be stored and accessed. CONCLUSION: META-ALL is a system for managing information about metabolic pathways. It facilitates the handling of pathway-related data and is designed to help biochemists and molecular biologists in their daily research. It is available on the Web at http://bic-gh.de/meta-all and can be downloaded free of charge and installed locally.

Database Management Systems↗

VOMBAT: prediction of transcription factor binding sites using variable order Bayesian trees.

Variable order Markov models and variable order Bayesian trees have been proposed for the recognition of transcription factor binding sites, and it could be demonstrated that they outperform traditional models, such as position weight matrices, Markov models and Bayesian trees. We develop a web server for the recognition of DNA binding sites based on variable order Markov models and variable order Bayesian trees offering the following functionality: (i) given datasets with annotated binding sites and genomic background sequences, variable order Markov models and variable order Bayesian trees can be trained; (ii) given a set of trained models, putative DNA binding sites can be predicted in a given set of genomic sequences and (iii) given a dataset with annotated binding sites and a dataset with genomic background sequences, cross-validation experiments for different model combinations with different parameter settings can be performed. Several of the offered services are computationally demanding, such as genome-wide predictions of DNA binding sites in mammalian genomes or sets of 10(4)-fold cross-validation experiments for different model combinations based on problem-specific data sets. In order to execute these jobs, and in order to serve multiple users at the same time, the web server is attached to a Linux cluster with 150 processors. VOMBAT is available at http://pdw-24.ipk-gatersleben.de:8080/VOMBAT/.

Algorithms↗

Fractionally integrated process with power-law correlations in variables and magnitudes.

Motivated by the fact that many empirical time series--including changes of heartbeat intervals, physical activity levels, intertrade times in finance, and river flux values--exhibit power-law anticorrelations in the variables and power-law correlations in their magnitudes, we propose a simple stochastic process that can account for both types of correlations. The process depends on only two parameters, where one controls the correlations in the variables and the other controls the correlations in their magnitudes. We apply the process to time series of heartbeat interval changes and air temperature changes and find that the statistical properties of the modeled time series are in agreement with those observed in the data.

Algorithms↗

Power-law correlated processes with asymmetric distributions.

Motivated by the fact that many physical systems display (i) power-law correlations together with (ii) an asymmetry in the probability distribution, we propose a stochastic process that can model both properties. The process depends on only two parameters, where one controls the scaling exponent of the power-law correlations, and the other controls the degree of asymmetry in the distributions leaving the correlations unaffected. We apply the process to air humidity data and find that the statistical properties of the process are in a good agreement with those observed in the data.

Algorithms↗

Analysis of sequence, map position, and gene expression reveals conserved essential genes for iron uptake in Arabidopsis and tomato.

Arabidopsis (Arabidopsis thaliana) and tomato (Lycopersicon esculentum) show similar physiological responses to iron deficiency, suggesting that homologous genes are involved. Essential gene functions are generally considered to be carried out by orthologs that have remained conserved in sequence and map position in evolutionarily related species. This assumption has not yet been proven for plant genomes that underwent large genome rearrangements. We addressed this question in an attempt to deduce functional gene pairs for iron reduction, iron transport, and iron regulation between Arabidopsis and tomato. Iron uptake processes are essential for plant growth. We investigated iron uptake gene pairs from tomato and Arabidopsis, namely sequence, conserved gene content of the regions containing iron uptake homologs based on conserved orthologous set marker analysis, gene expression patterns, and, in two cases, genetic data. Compared to tomato, the Arabidopsis genome revealed more and larger gene families coding for the iron uptake functions. The number of possible homologous pairs was reduced if functional expression data were taken into account in addition to sequence and map position. We predict novel homologous as well as partially redundant functions of ferric reductase-like and iron-regulated transporter-like genes in Arabidopsis and tomato. Arabidopsis nicotianamine synthase genes encode a partially redundant family. In this study, Arabidopsis gene redundancy generally reflected the presumed genome duplication structure. In some cases, statistical analysis of conserved gene regions between tomato and Arabidopsis suggested a common evolutionary origin. Although involvement of conserved genes in iron uptake was found, these essential genes seem to be of paralogous rather than orthologous origin in tomato and Arabidopsis.

Amino Acid Sequence↗

Coding and non-coding DNA thermal stability differences in eukaryotes studied by melting simulation, base shuffling and DNA nearest neighbor frequency analysis.

The melting of the coding and non-coding classes of natural DNA sequences was investigated using a program, MELTSIM, which simulates DNA melting based upon an empirically parameterized nearest neighbor thermodynamic model. We calculated T(m) results of 8144 natural sequences from 28 eukaryotic organisms of varying F(GC) (mole fraction of G and C) and of 3775 coding and 3297 non-coding sequences derived from those natural sequences. These data demonstrated that the T(m) vs. F(GC) relationships in coding and non-coding DNAs are both linear but have a statistically significant difference (6.6%) in their slopes. These relationships are significantly different from the T(m) vs. F(GC) relationship embodied in the classical Marmur-Schildkraut-Doty (MSD) equation for the intact long natural sequences. By analyzing the simulation results from various base shufflings of the original DNAs and the average nearest neighbor frequencies of those natural sequences across the F(GC) range, we showed that these differences in the T(m) vs. F(GC) relationships are largely a direct result of systematic F(GC)-dependent biases in nearest neighbor frequencies for those two different DNA classes. Those differences in the T(m) vs. F(GC) relationships and biases in nearest neighbor frequencies also appear between the sequences from multicellular and unicellular organisms in the same coding or non-coding classes, albeit of smaller but significant magnitudes.

Base Composition↗

SNP2CAPS: a SNP and INDEL analysis tool for CAPS marker development.

With the influx of various SNP genotyping assays in recent years, there has been a need for an assay that is robust, yet cost effective, and could be performed using standard gel-based procedures. In this context, CAPS markers have been shown to meet these criteria. However, converting SNPs to CAPS markers can be a difficult process if done manually. In order to address this problem, we describe a computer program, SNP2CAPS, that facilitates the computational conversion of SNP markers into CAPS markers. 413 multiple aligned sequences derived from barley ESTs were analysed for the presence of polymorphisms in 235 distinct restriction sites. 282 (90%) of 314 alignments that contain sequence variation due to SNPs and InDels revealed at least one polymorphic restriction site. After reducing the number of restriction enzymes from 235 to 10, 31% of the polymorphic sites could still be detected. In order to demonstrate the usefulness of this tool for marker development, we experimentally validated some of the results predicted by SNP2CAPS.

Base Sequence↗

Extreme value distribution based gene selection criteria for discriminant microarray data analysis using logistic regression.

One important issue commonly encountered in the analysis of microarray data is to decide which and how many genes should be selected for further studies. For discriminant microarray data analyses based on statistical models, such as the logistic regression models, gene selection can be accomplished by a comparison of the maximum likelihood of the model given the real data, L(D|M), and the expected maximum likelihood of the model given an ensemble of surrogate data with randomly permuted label, L(D(0)|M). Typically, the computational burden for obtaining L(D(0)M) is immense, often exceeding the limits of available computing resources by orders of magnitude. Here, we propose an approach that circumvents such heavy computations by mapping the simulation problem to an extreme-value problem. We present the derivation of an asymptotic distribution of the extreme-value as well as its mean, median, and variance. Using this distribution, we propose two gene selection criteria, and we apply them to two microarray datasets and three classification tasks for illustration.

Chromosome Mapping↗

Repeats and correlations in human DNA sequences.

We study the nucleotide-nucleotide mutual information function I(k) of the DNA sequences of the three completely sequenced human chromosomes 20, 21, and 22. We find in each human chromosome (i) the absence of the k=3 base pair (bp) sequence periodicity characteristic for protein coding regions, (ii) the absence of the k=10-11 bp sequence periodicity characteristic for both protein secondary structure and DNA bendability, and (iii) the presence of significant statistical dependencies at about k=135 bp and at about k=165 bp. We investigate to which degree the density and composition of interspersed repeats might explain these observed statistical patterns in all three human chromosomes. We use simple stochastic models to substitute known interspersed repeats and find by numerical studies that (iv) the presence of interspersed repeats dominates short-range correlations as measured by I(k) on the scale of several hundred base pairs in human chromosomes 20, 21, and 22. On the other hand, we find that (v) interspersed repeats contribute only weakly to long-range correlations due to the clustering of highly abundant Alu repeats.

Alu Elements↗

Analysis of symbolic sequences using the Jensen-Shannon divergence.

We study statistical properties of the Jensen-Shannon divergence D, which quantifies the difference between probability distributions, and which has been widely applied to analyses of symbolic sequences. We present three interpretations of D in the framework of statistical physics, information theory, and mathematical statistics, and obtain approximations of the mean, the variance, and the probability distribution of D in random, uncorrelated sequences. We present a segmentation method based on D that is able to segment a nonstationary symbolic sequence into stationary subsequences, and apply this method to DNA sequences, which are known to be nonstationary on a wide range of different length scales.

Computational Biology↗

Applications of recursive segmentation to the analysis of DNA sequences.

Recursive segmentation is a procedure that partitions a DNA sequence into domains with a homogeneous composition of the four nucleotides A, C, G and T. This procedure can also be applied to any sequence converted from a DNA sequence, such as to a binary strong(G + C)/weak(A + T) sequence, to a binary sequence indicating the presence or absence of the dinucleotide CpG, or to a sequence indicating both the base and the codon position information. We apply various conversion schemes in order to address the following five DNA sequence analysis problems: isochore mapping, CpG island detection, locating the origin and terminus of replication in bacterial genomes, finding complex repeats in telomere sequences, and delineating coding and noncoding regions. We find that the recursive segmentation procedure can successfully detect isochore borders, CpG islands, and the origin and terminus of replication, but it needs improvement for detecting complex repeats as well as borders between coding and noncoding regions.

Algorithms↗