PubMed Health⌕ Search

Biomedical subjects

Stefan Wiemann

Publications and source records attributed to Stefan Wiemann.

At least 19 recordsLinked to original sources

High-content microscopy identifies new neurite outgrowth regulators.

Neurons, with their long axons and elaborate dendritic arbour, establish the complex circuitry that is essential for the proper functioning of the nervous system. Whereas a catalogue of structural, molecular, and functional differences between axons and dendrites is accumulating, the mechanisms involved in early events of neuronal differentiation, such as neurite initiation and elongation, are less well understood, mainly because the key molecules involved remain elusive. Here we describe the establishment and application of a microscopy-based approach designed to identify novel proteins involved in neurite initiation and/or elongation. We identified 21 proteins that affected neurite outgrowth when ectopically expressed in cells. Complementary time-lapse microscopy allowed us to discriminate between early and late effector proteins. Localization experiments with GFP-tagged proteins in fixed and living cells revealed a further 14 proteins that associated with neurite tips either early or late during neurite outgrowth. Coexpression experiments of the new effector proteins provide a first glimpse on a possible functional relationship of these proteins during neurite outgrowth. Altogether, we demonstrate the potential of the systematic microscope-based screening approaches described here to tackle the complex biological process of neurite outgrowth regulation.

Animals↗

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

Statistical methods and software for the analysis of highthroughput reverse genetic assays using flow cytometry readouts.

Highthroughput cell-based assays with flow cytometric readout provide a powerful technique for identifying components of biologic pathways and their interactors. Interpretation of these large datasets requires effective computational methods. We present a new approach that includes data pre-processing, visualization, quality assessment, and statistical inference. The software is freely available in the Bioconductor package prada. The method permits analysis of large screens to detect the effects of molecular interventions in cellular systems.

Databases, Factual↗

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗

An anthropoid-specific segmental duplication on human chromosome 1q22.

Segmental duplications (SDs) play a key role in genome evolution by providing material for gene diversification and creation of variant or novel functions. They also mediate recombinations, resulting in microdeletions, which have occasionally been associated with human genetic diseases. Here, we present a detailed analysis of a large genomic region (about 240 kb), located on human chromosome 1q22, that contains a tandem SD, SD1q22. This duplication occurred about 37 million years ago in a lineage leading to anthropoid primates, after their separation from prosimians but before the Old and New World monkey split. We reconstructed the hypothetical unduplicated ancestral locus and compared it with the extant SD1q22 region. Our data demonstrate that, as a consequence of the duplication, new anthropoid-specific genetic material has evolved in the resulting paralogous segments. We describe the emergence of two new genes, whose new functions could contribute to the speciation of anthropoid primates. Moreover, we provide detailed information regarding structure and evolution of the SD1q22 region that is a prerequisite for future studies of its anthropoid-specific functions and possible linkage to human genetic disorders.

Animals↗

The 3of5 web application for complex and comprehensive pattern matching in protein sequences.

BACKGROUND: The identification of patterns in biological sequences is a key challenge in genome analysis and in proteomics. Frequently such patterns are complex and highly variable, especially in protein sequences. They are frequently described using terms of regular expressions (RegEx) because of the user-friendly terminology. Limitations arise for queries with the increasing complexity of patterns and are accompanied by requirements for enhanced capabilities. This is especially true for patterns containing ambiguous characters and positions and/or length ambiguities. RESULTS: We have implemented the 3of5 web application in order to enable complex pattern matching in protein sequences. 3of5 is named after a special use of its main feature, the novel n-of-m pattern type. This feature allows for an extensive specification of variable patterns where the individual elements may vary in their position, order, and content within a defined stretch of sequence. The number of distinct elements can be constrained by operators, and individual characters may be excluded. The n-of-m pattern type can be combined with common regular expression terms and thus also allows for a comprehensive description of complex patterns. 3of5 increases the fidelity of pattern matching and finds ALL possible solutions in protein sequences in cases of length-ambiguous patterns instead of simply reporting the longest or shortest hits. Grouping and combined search for patterns provides a hierarchical arrangement of larger patterns sets. The algorithm is implemented as internet application and freely accessible. The application is available at http://dkfz.de/mga2/3of5/3of5.html. CONCLUSION: The 3of5 application offers an extended vocabulary for the definition of search patterns and thus allows the user to comprehensively specify and identify peptide patterns with variable elements. The n-of-m pattern type offers an improved accuracy for pattern matching in combination with the ability to find all solutions, without compromising the user friendliness of regular expression terms.

Algorithms↗

The systematic functional characterisation of Xq28 genes prioritises candidate disease genes.

BACKGROUND: Well known for its gene density and the large number of mapped diseases, the human sub-chromosomal region Xq28 has long been a focus of genome research. Over 40 of approximately 300 X-linked diseases map to this region, and systematic mapping, transcript identification, and mutation analysis has led to the identification of causative genes for 26 of these diseases, leaving another 17 diseases mapped to Xq28, where the causative gene is still unknown. To expedite disease gene identification, we have initiated the functional characterisation of all known Xq28 genes. RESULTS: By using a systematic approach, we describe the Xq28 genes by RNA in situ hybridisation and Northern blotting of the mouse orthologs, as well as subcellular localisation and data mining of the human genes. We have developed a relational web-accessible database with comprehensive query options integrating all experimental data. Using this database, we matched gene expression patterns with affected tissues for 16 of the 17 remaining Xq28 linked diseases, where the causative gene is unknown. CONCLUSION: By using this systematic approach, we have prioritised genes in linkage regions of Xq28-mapped diseases to an amenable number for mutational screens. Our database can be queried by any researcher performing highly specified searches including diseases not listed in OMIM or diseases that might be linked to Xq28 in the future.

Animals↗

PML-associated repressor of transcription (PAROT), a novel KRAB-zinc finger repressor, is regulated through association with PML nuclear bodies.

Promyelocytic leukemia nuclear bodies (PML-NBs) are implicated in transcriptional regulation. Here we identify a novel transcriptional repressor, PML-associated repressor of transcription (PAROT), which is regulated in its repressor activity through recruitment to PML-NBs. PAROT is a Krüppel-associated box ( KRAB) zinc-finger (ZNF) protein, which comprises an amino terminal KRAB-A and KRAB-B box, a linker domain and 8 tandemly repeated C(2)H(2)-ZNF motifs at its carboxy terminus. Consistent with its domain structure, when tethered to DNA, PAROT represses transcription, and this is partially released by the HDAC inhibitor trichostatin A. PAROT colocalizes with members of the heterochromatin protein 1 (HP1) family and with transcriptional intermediary factor-1beta/KRAB-associated protein 1 (TIF-1beta/KAP1), a transcriptional corepressor for the KRAB-ZNF family. Interestingly, PML isoform IV, in contrast to PML-III, efficiently recruits PAROT and TIF-1beta from heterochromatin to PML-NBs. PML-NB recruitment of PAROT partially releases its transcriptional repressor activity, indicating that PAROT can be regulated through subnuclear compartmentalization. Taken together, our data identify a novel transcriptional repressor and provide evidence for its regulation through association with PML-NBs.

Amino Acid Sequence↗

The LIFEdb database in 2006.

LIFEdb (http://www.LIFEdb.de) integrates data from large-scale functional genomics assays and manual cDNA annotation with bioinformatics gene expression and protein analysis. New features of LIFEdb include (i) an updated user interface with enhanced query capabilities, (ii) a configurable output table and the option to download search results in XML, (iii) the integration of data from cell-based screening assays addressing the influence of protein-overexpression on cell proliferation and (iv) the display of the relative expression ('Electronic Northern') of the genes under investigation using curated gene expression ontology information. LIFEdb enables researchers to systematically select and characterize genes and proteins of interest, and presents data and information via its user-friendly web-based interface.

Cell Proliferation↗

Proteomics and Beyond: a report on the 3rd Annual Spring Workshop of the HUPO-PSI 21-23 April 2006, San Francisco, CA, USA.

The theme of the third annual Spring workshop of the HUPO-PSI was "proteomics and beyond" and its underlying goal was to reach beyond the boundaries of the proteomics community to interact with groups working on the similar issues of developing interchange standards and minimal reporting requirements. Significant developments in many of the HUPO-PSI XML interchange formats, minimal reporting requirements and accompanying controlled vocabularies were reported, with many of these now feeding into the broader efforts of the Functional Genomics Experiment (FuGE) data model and Functional Genomics Ontology (FuGO) ontologies.

Animals↗

Functional profiling: from microarrays via cell-based assays to novel tumor relevant modulators of the cell cycle.

Cancer transcription microarray studies commonly deliver long lists of "candidate" genes that are putatively associated with the respective disease. For many of these genes, no functional information, even less their relevance in pathologic conditions, is established as they were identified in large-scale genomics approaches. Strategies and tools are thus needed to distinguish genes and proteins with mere tumor association from those causally related to cancer. Here, we describe a functional profiling approach, where we analyzed 103 previously uncharacterized genes in cancer relevant assays that probed their effects on DNA replication (cell proliferation). The genes had previously been identified as differentially expressed in genome-wide microarray studies of tumors. Using an automated high-throughput assay with single-cell resolution, we discovered seven activators and nine repressors of DNA replication. These were further characterized for effects on extracellular signal-regulated kinase 1/2 (ERK1/2) signaling (G1-S transition) and anchorage-independent growth (tumorigenicity). One activator and one inhibitor protein of ERK1/2 activation and three repressors of anchorage-independent growth were identified. Data from tumor and functional profiling make these proteins novel prime candidates for further in-depth study of their roles in cancer development and progression. We have established a novel functional profiling strategy that links genomics to cell biology and showed its potential for discerning cancer relevant modulators of the cell cycle in the candidate lists from microarray studies.

Animals↗

Alternative pre-mRNA processing regulates cell-type specific expression of the IL4l1 and NUP62 genes.

BACKGROUND: Given the complexity of higher organisms, the number of genes encoded by their genomes is surprisingly small. Tissue specific regulation of expression and splicing are major factors enhancing the number of the encoded products. Commonly these mechanisms are intragenic and affect only one gene. RESULTS: Here we provide evidence that the IL4I1 gene is specifically transcribed from the apparent promoter of the upstream NUP62 gene, and that the first two exons of NUP62 are also contained in the novel IL4I1_2 variant. While expression of IL4I1 driven from its previously described promoter is found mostly in B cells, the expression driven by the NUP62 promoter is restricted to cells in testis (Sertoli cells) and in the brain (e.g., Purkinje cells). Since NUP62 is itself ubiquitously expressed, the IL4I1_2 variant likely derives from cell type specific alternative pre-mRNA processing. CONCLUSION: Comparative genomics suggest that the promoter upstream of the NUP62 gene originally belonged to the IL4I1 gene and was later acquired by NUP62 via insertion of a retroposon. Since both genes are apparently essential, the promoter had to serve two genes afterwards. Expression of the IL4I1 gene from the "NUP62" promoter and the tissue specific involvement of the pre-mRNA processing machinery to regulate expression of two unrelated proteins indicate a novel mechanism of gene regulation.

Alternative Splicing↗

Gamma-BAR, a novel AP-1-interacting protein involved in post-Golgi trafficking.

A novel peripheral membrane protein (2c18) that interacts directly with the gamma 'ear' domain of the adaptor protein complex 1 (AP-1) in vitro and in vivo is described. Ultrastructural analysis demonstrates a colocalization of 2c18 and gamma1-adaptin at the trans-Golgi network (TGN) and on vesicular profiles. Overexpression of 2c18 increases the fraction of membrane-bound gamma1-adaptin and inhibits its release from membranes in response to brefeldin A. Knockdown of 2c18 reduces the steady-state levels of gamma1-adaptin on membranes. Overexpression or downregulation of 2c18 leads to an increased secretion of the lysosomal hydrolase cathepsin D, which is sorted by the mannose-6-phosphate receptor at the TGN, which itself involves AP-1 function for trafficking between the TGN and endosomes. This suggests that the direct interaction of 2c18 and gamma1-adaptin is crucial for membrane association and thus the function of the AP-1 complex in living cells. We propose to name this protein gamma-BAR.

Adaptor Protein Complex 1↗

Large-scale protein expression for proteome research.

Access to pure and soluble recombinant proteins is essential for numerous applications in proteome research, such as the production of antibodies, structural characterization of proteins, and protein microarrays. Through the German cDNA Consortium we have access to more than 1500 ORFs encoding uncharacterized proteins. Preparing a large number of recombinant proteins calls for the careful refinement and re-evaluation of protein purification tools. The expression and purification strategy should result in mg quantities of protein that can be employed in microarray-based assays. In addition, the experimental set-up should be robust enough to allow both automated protein expression screening and the production of the proteins on a mg scale. These requirements are best fulfilled by a bacterial expression system such as Escherichia coli. To develop an efficient expression strategy, 75 different ORFs were transferred into suitable expression vectors using the Gateway cloning system. Four different fusion tags (E. coli transcription-termination anti-termination factor (NusA), hexahistidine tag (6xHis), maltose binding protein (MBP) and GST) were analyzed for their effect on yield of induced fusion protein and its solubility, as determined at two different induction temperatures. Affinity-purified fusion proteins were confirmed by MALDI-TOF MS.

Amino Acid Sequence↗

Systematic comparison of surface coatings for protein microarrays.

To process large numbers of samples in parallel is one potential of protein microarrays for research and diagnostics. However, the application of protein arrays is currently hampered by the lack of comprehensive technological knowledge about the suitability of 2-D and 3-D slide surface coatings. We have performed a systematic study to analyze how both surface types perform in combination with different fluorescent dyes to generate significant and reproducible data. In total, we analyzed more than 100 slides containing 1152 spots each. Slides were probed against different monoclonal antibodies (mAbs) and recombinant fusion proteins. We found two surface coatings to be most suitable for protein and antibody (Ab) immobilization. These were further subjected to quantitative analyses by evaluating intraslide and slide-to-slide reproducibilities, and the linear range of target detection. In summary, we demonstrate that only suitable combinations of surface and fluorescent dyes allow the generation of highly reproducible data.

Antibodies↗

Protein microarrays as a discovery tool for studying protein-protein interactions.

Exploring the function of the genome and the encoded proteins has emerged as a new and exciting challenge in the postgenomic era. Novel technologies come into view that promise to be valuable for the investigation not only of single proteins, but of entire protein networks. Protein microarrays are the innovative assay platform for highly parallel in vitro studies of protein-protein interactions. Due to their flexibility and multiplexing capacity, protein microarrays benefit basic research, diagnosis and biomedicine. This review provides an overview on the basic principles of protein microarrays and their potential to multiplex protein-protein interaction studies.

DNA Replication↗

SMART amplification combined with cDNA size fractionation in order to obtain large full-length clones.

BACKGROUND: cDNA libraries are widely used to identify genes and splice variants, and as a physical resource for full-length clones. Conventionally-generated cDNA libraries contain a high percentage of 5'-truncated clones. Current library construction methods that enrich for full-length mRNA are laborious, and involve several enzymatic steps performed on mRNA, which renders them sensitive to RNA degradation. The SMART technique for full-length enrichment is robust but results in limited cDNA insert size of the library. RESULTS: We describe a method to construct SMART full-length enriched cDNA libraries with large insert sizes. Sub-libraries were generated from size-fractionated cDNA with an average insert size of up to seven kb. The percentage of full-length clones was calculated for different size ranges from BLAST results of over 12,000 5'ESTs. CONCLUSIONS: The presented technique is suitable to generate full-length enriched cDNA libraries with large average insert sizes in a straightforward and robust way. The representation of full-coding clones is high also for large cDNAs (70%, 4-10 kb), when high-quality starting mRNA is used.

Cell Line, Tumor↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗