PubMed Health⌕ Search

Biomedical subjects

Markus Seiler

Publications and source records attributed to Markus Seiler.

3 recordsLinked to original sources

The 3of5 web application for complex and comprehensive pattern matching in protein sequences.

BACKGROUND: The identification of patterns in biological sequences is a key challenge in genome analysis and in proteomics. Frequently such patterns are complex and highly variable, especially in protein sequences. They are frequently described using terms of regular expressions (RegEx) because of the user-friendly terminology. Limitations arise for queries with the increasing complexity of patterns and are accompanied by requirements for enhanced capabilities. This is especially true for patterns containing ambiguous characters and positions and/or length ambiguities. RESULTS: We have implemented the 3of5 web application in order to enable complex pattern matching in protein sequences. 3of5 is named after a special use of its main feature, the novel n-of-m pattern type. This feature allows for an extensive specification of variable patterns where the individual elements may vary in their position, order, and content within a defined stretch of sequence. The number of distinct elements can be constrained by operators, and individual characters may be excluded. The n-of-m pattern type can be combined with common regular expression terms and thus also allows for a comprehensive description of complex patterns. 3of5 increases the fidelity of pattern matching and finds ALL possible solutions in protein sequences in cases of length-ambiguous patterns instead of simply reporting the longest or shortest hits. Grouping and combined search for patterns provides a hierarchical arrangement of larger patterns sets. The algorithm is implemented as internet application and freely accessible. The application is available at http://dkfz.de/mga2/3of5/3of5.html. CONCLUSION: The 3of5 application offers an extended vocabulary for the definition of search patterns and thus allows the user to comprehensively specify and identify peptide patterns with variable elements. The n-of-m pattern type offers an improved accuracy for pattern matching in combination with the ability to find all solutions, without compromising the user friendliness of regular expression terms.

Algorithms↗

Functional profiling: from microarrays via cell-based assays to novel tumor relevant modulators of the cell cycle.

Cancer transcription microarray studies commonly deliver long lists of "candidate" genes that are putatively associated with the respective disease. For many of these genes, no functional information, even less their relevance in pathologic conditions, is established as they were identified in large-scale genomics approaches. Strategies and tools are thus needed to distinguish genes and proteins with mere tumor association from those causally related to cancer. Here, we describe a functional profiling approach, where we analyzed 103 previously uncharacterized genes in cancer relevant assays that probed their effects on DNA replication (cell proliferation). The genes had previously been identified as differentially expressed in genome-wide microarray studies of tumors. Using an automated high-throughput assay with single-cell resolution, we discovered seven activators and nine repressors of DNA replication. These were further characterized for effects on extracellular signal-regulated kinase 1/2 (ERK1/2) signaling (G1-S transition) and anchorage-independent growth (tumorigenicity). One activator and one inhibitor protein of ERK1/2 activation and three repressors of anchorage-independent growth were identified. Data from tumor and functional profiling make these proteins novel prime candidates for further in-depth study of their roles in cancer development and progression. We have established a novel functional profiling strategy that links genomics to cell biology and showed its potential for discerning cancer relevant modulators of the cell cycle in the candidate lists from microarray studies.

Animals↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗