PubMed Health⌕ Search

Biomedical subjects

Michael Beckstette

Publications and source records attributed to Michael Beckstette.

3 recordsLinked to original sources

Fast index based algorithms and software for matching position specific scoring matrices.

BACKGROUND: In biological sequence analysis, position specific scoring matrices (PSSMs) are widely used to represent sequence motifs in nucleotide as well as amino acid sequences. Searching with PSSMs in complete genomes or large sequence databases is a common, but computationally expensive task. RESULTS: We present a new non-heuristic algorithm, called ESAsearch, to efficiently find matches of PSSMs in large databases. Our approach preprocesses the search space, e.g., a complete genome or a set of protein sequences, and builds an enhanced suffix array that is stored on file. This allows the searching of a database with a PSSM in sublinear expected time. Since ESAsearch benefits from small alphabets, we present a variant operating on sequences recoded according to a reduced alphabet. We also address the problem of non-comparable PSSM-scores by developing a method which allows the efficient computation of a matrix similarity threshold for a PSSM, given an E-value or a p-value. Our method is based on dynamic programming and, in contrast to other methods, it employs lazy evaluation of the dynamic programming matrix. We evaluated algorithm ESAsearch with nucleotide PSSMs and with amino acid PSSMs. Compared to the best previous methods, ESAsearch shows speedups of a factor between 17 and 275 for nucleotide PSSMs, and speedups up to factor 1.8 for amino acid PSSMs. Comparisons with the most widely used programs even show speedups by a factor of at least 3.8. Alphabet reduction yields an additional speedup factor of 2 on amino acid sequences compared to results achieved with the 20 symbol standard alphabet. The lazy evaluation method is also much faster than previous methods, with speedups of a factor between 3 and 330. CONCLUSION: Our analysis of ESAsearch reveals sublinear runtime in the expected case, and linear runtime in the worst case for sequences not shorter than the absolute value of A(m) + m - 1, where m is the length of the PSSM and A a finite alphabet. In practice, ESAsearch shows superior performance over the most widely used programs, especially for DNA sequences. The new algorithm for accurate on-the-fly calculations of thresholds has the potential to replace formerly used approximation approaches. Beyond the algorithmic contributions, we provide a robust, well documented, and easy to use software package, implementing the ideas and algorithms presented in this manuscript.

Algorithms↗

XenDB: full length cDNA prediction and cross species mapping in Xenopus laevis.

BACKGROUND: Research using the model system Xenopus laevis has provided critical insights into the mechanisms of early vertebrate development and cell biology. Large scale sequencing efforts have provided an increasingly important resource for researchers. To provide full advantage of the available sequence, we have analyzed 350,468 Xenopus laevis Expressed Sequence Tags (ESTs) both to identify full length protein encoding sequences and to develop a unique database system to support comparative approaches between X. laevis and other model systems. DESCRIPTION: Using a suffix array based clustering approach, we have identified 25,971 clusters and 40,877 singleton sequences. Generation of a consensus sequence for each cluster resulted in 31,353 tentative contig and 4,801 singleton sequences. Using both BLASTX and FASTY comparison to five model organisms and the NR protein database, more than 15,000 sequences are predicted to encode full length proteins and these have been matched to publicly available IMAGE clones when available. Each sequence has been compared to the KOG database and approximately 67% of the sequences have been assigned a putative functional category. Based on sequence homology to mouse and human, putative GO annotations have been determined. CONCLUSION: The results of the analysis have been stored in a publicly available database XenDB http://bibiserv.techfak.uni-bielefeld.de/xendb/. A unique capability of the database is the ability to batch upload cross species queries to identify potential Xenopus homologues and their associated full length clones. Examples are provided including mapping of microarray results and application of 'in silico' analysis. The ability to quickly translate the results of various species into 'Xenopus-centric' information should greatly enhance comparative embryological approaches.

Animals↗

The maize Single myb histone 1 gene, Smh1, belongs to a novel gene family and encodes a protein that binds telomere DNA repeats in vitro.

We screened maize (Zea mays) cDNAs for sequences similar to the single myb-like DNA-binding domain of known telomeric complex proteins. We identified, cloned, and sequenced five full-length cDNAs representing a novel gene family, and we describe the analysis of one of them, the gene Single myb histone 1 (Smh1). The Smh1 gene encodes a small, basic protein with a unique triple motif structure of (a) an N-terminal SANT/myb-like domain of the homeodomain-like superfamily of 3-helical-bundle-fold proteins, (b) a central region with homology to the conserved H1 globular domain found in the linker histones H1/H5, and (c) a coiled-coil domain near the C terminus. The Smh-type genes are plant specific and include a gene family in Arabidopsis and the PcMYB1 gene of parsley (Petroselinum crispum) but are distinct from those (AtTRP1, AtTBP1, and OsRTBP1) recently shown to encode in vitro telomere-repeat DNA-binding activity. The Smh1 gene is expressed in leaf tissue and maps to chromosome 8 (bin 8.05), with a duplicate locus on chromosome 3 (bin 3.09). A recombinant full-length SMH1, rSMH1, was found by band-shift assays to bind double-stranded oligonucleotide probes with at least two internal tandem copies of the maize telomere repeat, TTTAGGG. Point mutations in the telomere repeat residues reduced or abolished the binding, whereas rSMH1 bound nonspecifically to single-stranded DNA probes. The two DNA-binding motifs in SMH proteins may provide a link between sequence recognition and chromatin dynamics and may function at telomeres or other sites in the nucleus.

Amino Acid Sequence↗