PubMed Health⌕ Search

Biomedical subjects

J Schuchhardt

Publications and source records attributed to J Schuchhardt.

9 recordsLinked to original sources

Combining frequency and positional information to predict transcription factor binding sites.

MOTIVATION: Even though a number of genome projects have been finished on the sequence level, still only a small proportion of DNA regulatory elements have been identified. Growing amounts of gene expression data provide the possibility of finding coregulated genes by clustering methods. By analysis of the promoter regions of those genes, rather weak signals of transcription factor binding sites may be detected. RESULTS: We introduce the new algorithm ITB, an Integrated Tool for Box finding, which combines frequency and positional information to predict transcription factor binding sites in upstream regions of coregulated genes. Motifs are extracted by exhaustive analysis of regular expression-like patterns and by estimating probabilities of positional clusters of motifs. ITB detects consensus sequences of experimentally verified transcription factor binding sites of the yeast Saccharomyces cerevisiae. Moreover, a number of new binding site candidates with significant scores are predicted. Besides applying ITB on yeast upstream regions, the program is run on human promoter sequences. AVAILABILITY: ITB is available upon request.

Algorithms↗

Normalization strategies for cDNA microarrays.

Multiple Arabidopsis thaliana clones from an experimental series of cDNA microarrays are evaluated in order to identify essential sources of noise in the spotting and hybridization process. Theoretical and experimental strategies for an improved quantitative evaluation of cDNA microarrays are proposed and tested on a series of differently diluted control clones. Several sources of noise are identified from the data. Systematic and stochastic fluctuations in the spotting process are reduced by control spots and statistical techniques. The reliability of slide to slide comparison is critically assessed within the statistical framework of pattern matching and classification.

Arabidopsis↗

Adaptive encoding neural networks for the recognition of human signal peptide cleavage sites.

MOTIVATION: Data representation and encoding are essential for classification of protein sequences with artificial neural networks (ANN). Biophysical properties are appropriate for low dimensional encoding of protein sequence data. However, in general there is no a priori knowledge of the relevant properties for extraction of representative features. RESULTS: An adaptive encoding artificial neural network (ACN) for recognition of sequence patterns is described. In this approach parameters for sequence encoding are optimized within the same process as the weight vectors by an evolutionary algorithm. The method is applied to the prediction of signal peptide cleavage sites in human secretory proteins and compared with an established predictor for signal peptides. CONCLUSION: Knowledge of physico-chemical properties is not necessary for training an ACN. The advantage is a low dimensional data representation leading to computational efficiency, easy evaluation of the detected features, and high prediction accuracy. AVAILABILITY: A cleavage site prediction server is located at the Humboldt University http://itb.biologie.hu-berlin.de/ approximately jo/sig-cleave/ACNpredictor.cgi CONTACT: jo@itb.hu-berlin.de; berndj@zedat.fu-berlin.de

Binding Sites↗

Tissue gene expression analysis using arrayed normalized cDNA libraries.

We have used oligonucleotide-fingerprinting data on 60,000 cDNA clones from two different mouse embryonic stages to establish a normalized cDNA clone set. The normalized set of 5,376 clones represents different clusters and therefore, in almost all cases, different genes. The inserts of the cDNA clones were amplified by PCR and spotted on glass slides. The resulting arrays were hybridized with mRNA probes prepared from six different adult mouse tissues. Expression profiles were analyzed by hierarchical clustering techniques. We have chosen radioactive detection because it combines robustness with sensitivity and allows the comparison of multiple normalized experiments. Sensitive detection combined with highly effective clustering algorithms allowed the identification of tissue-specific expression profiles and the detection of genes specifically expressed in the tissues investigated. The obtained results are publicly available (http://www.rzpd.de) and can be used by other researchers as a digital expression reference.

Algorithms↗

Local structural motifs of protein backbones are classified by self-organizing neural networks.

Important and relevant information is expected to be encoded in local structural elements of proteins. An unsupervised learning algorithm (Kohonen algorithm) was applied to the representation and unbiased classification of local backbone structures contained in a set of proteins. Training yielded a two-dimensional Kohonen feature map with 100 different structural motifs including certain helical and strand structures. All motifs were represented in a phi-psi-plot and some of them as a three-dimensional model. The course of structural motifs along the backbone of four selected proteins (cytochrome b5, cytochrome b562, lysozyme, gamma crystallin) was investigated in detail. Trajectories and histograms visualizing the abundance of characteristic motifs allowed for the distinction between different types of protein overall folds. It is demonstrated how the histograms may be used to construct a structural similarity matrix for proteins. The Kohonen algorithm provides a simple procedure for classification of local protein structures independent of any a priori knowledge of leading structural motifs. Training of the Kohonen network leads to the generation of "consensus structures' serving for the task of classification.

Algorithms↗

Development of simple fitness landscapes for peptides by artificial neural filter systems.

The applicability of artificial neural filter systems as fitness functions for sequence-oriented peptide design was evaluated. Two example applications were selected: classification of dipeptides according to their hydrophobicity and classification of proteolytic cleavage-sites of protein precursor sequences according to their mean hydrophobicities and mean side-chain volumes. The cleavage-sites covered 12 residues. In the dipeptide experiments the objective was to separate a selected set of molecules from all other possible dipeptide sequences. Perceptrons, feedforward networks with one hidden layer, and a hybrid network were applied. The filters were trained by a (1, lambda) evolution strategy. Two types of network units employing either a sigmoidal or a unimodal transfer function were used in the feedforward filters, and their influence on classification was investigated. The two-layer hybrid network employed gaussian activation functions. To analyze classification of the different filter systems, their output was plotted in the two-dimensional sequence space. The diagrams were interpreted as fitness landscapes qualifying the markedness of a characteristic peptide feature which can be used as a guide through sequence space for rational peptide design. It is demonstrated that the applicability of neural filter systems as a heuristic method for sequence optimization depends on both the appropriate network architecture and selection of representative sequence data. The networks with unimodal activation functions and the hybrid networks both led to a number of local optima. However, the hybrid networks produced the best prediction results. In contrast, the filters with sigmoidal activation produced good reclassification results leading to fitness landscapes lacking unreasonable local optima. Similar results were obtained for classification of both dipeptides and cleavage-site sequences.

Amino Acid Sequence↗

Peptide design in machina: development of artificial mitochondrial protein precursor cleavage sites by simulated molecular evolution.

Artificial neural networks were used for extraction of characteristic physiochemical features from mitochondrial matrix metalloprotease target sequences. The amino acid properties hydrophobicity and volume were used for sequence encoding. A window of 12 residues was employed, encompassing positions -7 to +5 of precursors with cleavage sites. Two sets of noncleavage site examples were selected for network training which was performed by an evolution strategy. The weight vectors of the optimized networks were visualized and interpreted by Hinton diagrams. A neural filter system consisting of 13 perceptron-type networks accurately classified the data. It served as the fitness function in a simulated molecular evolution procedure for sequence-oriented de novo design of idealized cleavage sites. A detailed description of the strategy is given. Several putative high-quality cleavage sites were obtained revealing the critical nature of the residues in the positions -2 and -5. Charged residues seem to have a major influence on cleavage site function.

Amino Acid Sequence↗

Artificial neural networks and simulated molecular evolution are potential tools for sequence-oriented protein design.

The potential of artificial neural filter systems for feature extraction from amino acid sequences is discussed. Analysis of signal peptidase I cleavage-sites in protein precursor sequences serves as an example application. Trained neural networks can be used as the fitness function in an evolutionary protein design cycle termed 'simulated molecular evolution' which is an entirely computer-based method for the rational design of locally encoded amino acid sequence features. The design procedure itself is regarded as an optimization process which can follow several schemes. Gradient search, diffusive search, and evolution strategy have been compared with regard to their usefulness for optimization. It turns out that gradient search is well suited for optimization in smooth fitness landscapes without local minima, whereas evolution strategy seems to be a method of choice for optimization in a high-dimensional multimodal search space. This is concluded from optimization experiments using a multimodal example function.

Amino Acid Sequence↗