PubMed Health⌕ Search

Biomedical subjects

Jens Hollunder

Publications and source records attributed to Jens Hollunder.

5 recordsLinked to original sources

DASS: efficient discovery and p-value calculation of substructures in unordered data.

MOTIVATION: Pattern identification in biological sequence data is one of the main objectives of bioinformatics research. However, few methods are available for detecting patterns (substructures) in unordered datasets. Data mining algorithms mainly developed outside the realm of bioinformatics have been adapted for that purpose, but typically do not determine the statistical significance of the identified patterns. Moreover, these algorithms do not exploit the often modular structure of biological data. RESULTS: We present the algorithm DASS (Discovery of All Significant Substructures) that first identifies all substructures in unordered data (DASS(Sub)) in a manner that is especially efficient for modular data. In addition, DASS calculates the statistical significance of the identified substructures, for sets with at most one element of each type (DASS(P(set))), or for sets with multiple occurrence of elements (DASS(P(mset))). The power and versatility of DASS is demonstrated by four examples: combinations of protein domains in multi-domain proteins, combinations of proteins in protein complexes (protein subcomplexes), combinations of transcription factor target sites in promoter regions and evolutionarily conserved protein interaction subnetworks. AVAILABILITY: The program code and additional data are available at http://www.fli-leibniz.de/tsb/DASS

Algorithms↗

Integrated assessment and prediction of transcription factor binding.

Systematic chromatin immunoprecipitation (chIP-chip) experiments have become a central technique for mapping transcriptional interactions in model organisms and humans. However, measurement of chromatin binding does not necessarily imply regulation, and binding may be difficult to detect if it is condition or cofactor dependent. To address these challenges, we present an approach for reliably assigning transcription factors (TFs) to target genes that integrates many lines of direct and indirect evidence into a single probabilistic model. Using this approach, we analyze publicly available chIP-chip binding profiles measured for yeast TFs in standard conditions, showing that our model interprets these data with significantly higher accuracy than previous methods. Pooling the high-confidence interactions reveals a large network containing 363 significant sets of factors (TF modules) that cooperate to regulate common target genes. In addition, the method predicts 980 novel binding interactions with high confidence that are likely to occur in so-far untested conditions. Indeed, using new chIP-chip experiments we show that predicted interactions for the factors Rpn4p and Pdr1p are observed only after treatment of cells with methyl-methanesulfonate, a DNA-damaging agent. We outline the first approach for consistently integrating all available evidences for TF-target interactions and we comprehensively identify the resulting TF module hierarchy. Prioritizing experimental conditions for each factor will be especially important as increasing numbers of chIP-chip assays are performed in complex organisms such as humans, for which "standard conditions" are ill defined.

Algorithms↗

Common patterns in type II restriction enzyme binding sites.

Restriction enzymes are among the best studied examples of DNA binding proteins. In order to find general patterns in DNA recognition sites, which may reflect important properties of protein-DNA interaction, we analyse the binding sites of all known type II restriction endonucleases. We find a significantly enhanced GC content and discuss three explanations for this phenomenon. Moreover, we study patterns of nucleotide order in recognition sites. Our analysis reveals a striking accumulation of adjacent purines (R) or pyrimidines (Y). We discuss three possible reasons: RR/YY dinucleotides are characterized by (i) stronger H-bond donor and acceptor clusters, (ii) specific geometrical properties and (iii) a low stacking energy. These features make RR/YY steps particularly accessible for specific protein-DNA interactions. Finally, we show that the recognition sites of type II restriction enzymes are underrepresented in host genomes and in phage genomes.

Bacteriophages↗

Identification and characterization of protein subcomplexes in yeast.

Protein complexes are major components of cellular organization. Based on large-scale protein complex data, we present the first statistical procedure to find insightful substructures in protein complexes: we identify protein subcomplexes (SCs), i.e., multiprotein assemblies residing in different protein complexes. Four protein complex datasets with different origins and variable reliability are separately analyzed. Our method identifies well-characterized protein assemblies with known functions, thereby confirming the utility of the procedure. In addition, we also identify hitherto unknown functional entities consisting of either functionally unknown proteins or proteins with different functional annotation. We show that SCs represent more reliable protein assemblies than the original complexes. Finally, we demonstrate unique properties of subcomplex proteins that underline the distinct roles of SCs: (i) SCs are functionally and spatially more homogeneous than complete protein complexes (this fact is utilized to predict functional roles and subcellular localizations for so far unannotated proteins); (ii) the abundance of subcomplex proteins is less variable than the abundance of other proteins; (iii) SCs are enriched with essential and synthetic lethal proteins; and (iv) mutations in SC-proteins have higher fitness effects than mutations in other proteins.

Gene Deletion↗

Post-transcriptional expression regulation in the yeast Saccharomyces cerevisiae on a genomic scale.

Based on large-scale data for the yeast Saccharomyces cerevisiae (protein and mRNA abundance, translational status, transcript length), we investigate the relation of transcription, translation, and protein turnover on a genome-wide scale. We elucidate variations between different spatial cell compartments and functional modules by comparing protein-to-mRNA ratios, translational activity, and a novel descriptor for protein-specific degradation (protein half-life descriptor). This analysis helps to understand the cell's strategy to use transcriptional and post-transcriptional regulation mechanisms for managing protein levels. For instance, it is possible to identify modules that are subject to suppressed translation under normal conditions ("translation on demand"). In order to reduce inconsistencies between the datasets, we compiled a new reference mRNA abundance dataset and we present a novel approach to correct large microarray signals for a saturation bias. Accounting for ribosome density based on transcript length rather than ORF length improves the correlation of observed protein levels to translational activity. We discuss potential causes for the deviations of these correlations. Finally, we introduce a quantitative descriptor for protein degradation (protein half-life descriptor) and compare it to measured half-lives. The study demonstrates significant post-transcriptional control of protein levels for a number of different compartments and functional modules, which is missed when exclusively focusing on transcript levels.

Cell Compartmentation↗