PubMed Health⌕ Search

Biomedical subjects

Markus Krupp

Publications and source records attributed to Markus Krupp.

4 recordsLinked to original sources

Genome-wide analysis of factors regulating gene expression in liver.

In recent decades, multiple individual genes have been studied with respect to their level of expression in liver tissue and in many cases substantial progress has been made in identifying individual factors promoting gene expression in liver. However, the overall picture is still undefined and general rules or factors regulating gene expression in liver have not yet been established. Thus, a genome-wide screen for factors regulating gene expression in liver is of high interest, as it may reveal common regulatory mechanisms for most genes highly expressed in liver. These factors represent potential new targets in liver disease associated with differential gene expression. Using a novel bioinformatics approach, we have performed a genome-wide, bioinformatic screen to identify genetic factors regulating gene expression in liver. As the expression of an individual gene is generally driven by its promoter activity, we compared the level of expression to individual promoter sequences. Expression data of 15,704 individual genes in 12 tissues were obtained from a normal tissue microarray dataset. The genes were subsequently divided into two groups, according to whether or not their highest expression level was found in liver tissue. Scanning 1000 bp upstream of the transcription start of each individual gene and using the PromoterScan algorithm, we were able to identify a total of 7042 promoters containing a total of 241,984 transcription factors. To eliminate the possibility that currently unknown transcription factors may be crucial to liver expression regulation, we investigated all possible nucleotide combinations of 8 bp and 10 bp which we reasoned may serve as novel binding sites for transcription factors. In both screens we did not detect any significant, biologically relevant differences in numbers of transcription factors and binding sites between the two groups. Furthermore, we excluded possible differences in distribution of TATA-boxes or CpG islands as well as differences in nucleotide composition of RNA sequences or amino acid composition of transcribed protein sequences. We conclude that the existence of central, superordinated regulatory factors in liver gene expression is unlikely and that expression of individual genes in liver is more likely to be dependent on individual combinations of regulating factors for each gene.

Algorithms↗

Actin binding LIM protein 3 (abLIM3).

LIM domain proteins were demonstrated to play key roles in various biological processes such as embryonic development, cell lineage determination, and cancer differentiation. Actin binding LIM protein 1 (abLIM1) was reported to be localized in a genomic region often deleted in human cancers and suggested to be involved in axon guidance. Recently, existence of a second family member was reported, actin binding LIM protein 2. By means of computational biology and comparative genomics, we now characterized an additional, third member of the actin binding LIM protein subgroup, actin binding LIM protein 3 (abLIM3). The human mRNA sequence was previously annotated as differentially regulated in hepatoblastoma compared to normal livers. Conservation of key structural features of abLIM1 and abLIM2, four LIM domains and a VHD domain, suggested comparable biological function of abLIM3 as a linker between actin cytoskeleton and cell signaling pathways. AbLIM3 was found to be conserved in vertebrates, as orthologous sequences were characterized for mouse, fish, and frog. In addition, we report the existence of abLIM2 orthologs in fish and frog, suggesting a similar degree of evolutionary conservation. The intracellular localization of the abLIM3 protein was predicted to be nuclear by means of Reinhardt's neural network and the k-nearest neighbor algorithm. The corresponding abLIM3 gene was localized to chromosome 5q32 and spanned 119 kb, organized in 24 exons. An RT-PCR based expression profile available from the human unidentified gene-encoded (HUGE) database demonstrated highest expression for abLIM3 in heart, lung, liver, and brain/cerebellum accompanied by lower expression in multiple other tissues. Furthermore, abLIM3 was expressed in fetal liver, CNS, and spinal cord.

Amino Acid Sequence↗

Current bioinformatics tools in genomic biomedical research (Review).

On the advent of a completely assembled human genome, modern biology and molecular medicine stepped into an era of increasingly rich sequence database information and high-throughput genomic analysis. However, as sequence entries in the major genomic databases currently rise exponentially, the gap between available, deposited sequence data and analysis by means of conventional molecular biology is rapidly widening, making new approaches of high-throughput genomic analysis necessary. At present, the only effective way to keep abreast of the dramatic increase in sequence and related information is to apply biocomputational approaches. Thus, over recent years, the field of bioinformatics has rapidly developed into an essential aid for genomic data analysis and powerful bioinformatics tools have been developed, many of them publicly available through the World Wide Web. In this review, we summarize and describe the basic bioinformatics tools for genomic research such as: genomic databases, genome browsers, tools for sequence alignment, single nucleotide polymorphism (SNP) databases, tools for ab initio gene prediction, expression databases, and algorithms for promoter prediction.

Computational Biology↗

STRING: known and predicted protein-protein associations, integrated and transferred across organisms.

A full description of a protein's function requires knowledge of all partner proteins with which it specifically associates. From a functional perspective, 'association' can mean direct physical binding, but can also mean indirect interaction such as participation in the same metabolic pathway or cellular process. Currently, information about protein association is scattered over a wide variety of resources and model organisms. STRING aims to simplify access to this information by providing a comprehensive, yet quality-controlled collection of protein-protein associations for a large number of organisms. The associations are derived from high-throughput experimental data, from the mining of databases and literature, and from predictions based on genomic context analysis. STRING integrates and ranks these associations by benchmarking them against a common reference set, and presents evidence in a consistent and intuitive web interface. Importantly, the associations are extended beyond the organism in which they were originally described, by automatic transfer to orthologous protein pairs in other organisms, where applicable. STRING currently holds 730,000 proteins in 180 fully sequenced organisms, and is available at http://string.embl.de/.

Databases, Protein↗