PubMed Health⌕ Search

Biomedical subjects

Volker Brendel

Publications and source records attributed to Volker Brendel.

12 recordsLinked to original sources

Identification, characterization and molecular phylogeny of U12-dependent introns in the Arabidopsis thaliana genome.

U12-dependent introns are spliced by the minor U12-type spliceosome and occur in a variety of eukaryotic organisms, including Arabidopsis. In this study, a set of putative U12-dependent introns was compiled from a large collection of cDNA/EST- confirmed introns in the Arabidopsis thaliana genome by means of high-throughput bioinformatic analysis combined with manual scrutiny. A total of 165 U12-type introns were identified based upon stringent criteria. This number of sequences well exceeds the total number of U12-type introns previously reported for plants and allows a more thorough statistical analysis of U12-type signals. Of particular note is the discovery that the distance between the branch site adenosine and the acceptor site ranges from 10 to 39 nt, significantly longer than the previously postulated limit of 21 bp. Further analysis indicates that, in addition to the spacing constraint, the sequence context of the potential acceptor site may have an important role in 3' splice site selection. Several alternative splicing events involving U12-type introns were also captured in this study, providing evidence that U12-dependent acceptor sites can also be recognized by the U2-type spliceosome. Furthermore, phylogenetic analysis suggests that both U12-type AT-AC and U12-type GT-AG introns occurred in Na+/H+ antiporters in a progenitor of animals and plants.

Alternative Splicing↗

GeneSeqer@PlantGDB: Gene structure prediction in plant genomes.

The GeneSeqer@PlantGDB Web server (http://www.plantgdb.org/cgi-bin/GeneSeqer.cgi) provides a gene structure prediction tool tailored for applications to plant genomic sequences. Predictions are based on spliced alignment with source-native ESTs and full-length cDNAs or non-native probes derived from putative homologous genes. The tool is illustrated with applications to refinement of current gene structure annotation and de novo annotation of draft genomic sequences. The service should facilitate expert annotation as a community effort by providing convenient access to all public plant sequences via the PlantGDB database, a simple four-step protocol for spliced alignment and visually appealing displays of the predicted gene structures in addition to detailed sequence alignments.

Arabidopsis↗

Efficient clustering of large EST data sets on parallel computers.

Clustering expressed sequence tags (ESTs) is a powerful strategy for gene identification, gene expression studies and identifying important genetic variations such as single nucleotide polymorphisms. To enable fast clustering of large-scale EST data, we developed PaCE (for Parallel Clustering of ESTs), a software program for EST clustering on parallel computers. In this paper, we report on the design and development of PaCE and its evaluation using Arabidopsis ESTs. The novel features of our approach include: (i) design of memory efficient algorithms to reduce the memory required to linear in the size of the input, (ii) a combination of algorithmic techniques to reduce the computational work without sacrificing the quality of clustering, and (iii) use of parallel processing to reduce run-time and facilitate clustering of larger data sets. Using a combination of these techniques, we report the clustering of 168 200 Arabidopsis ESTs in 15 min on an IBM xSeries cluster with 30 dual-processor nodes. We also clustered 327 632 rat ESTs in 47 min and 420 694 Triticum aestivum ESTs in 3 h and 15 min. We demonstrate the quality of our software using benchmark Arabidopsis EST data, and by comparing it with CAP3, a software widely used for EST assembly. Our software allows clustering of much larger EST data sets than is possible with current software. Because of its speed, it also facilitates multiple runs with different parameters, providing biologists a tool to better analyze EST sequence data. Using PaCE, we clustered EST data from 23 plant species and the results are available at the PlantGDB website.

Algorithms↗

ZmDB, an integrated database for maize genome research.

Zea mays DataBase (ZmDB) seeks to provide a comprehensive view of maize (corn) genetics by linking genomic sequence data with gene expression analysis and phenotypes of mutant plants. ZmDB originated in 1999 as the Web portal for a large project of maize gene discovery, sequencing and phenotypic analysis using a transposon tagging strategy and expressed sequence tag (EST) sequencing. Recently, ZmDB has broadened its scope to include all public maize ESTs, genome survey sequences (GSSs), and protein sequences. More than 170 000 ESTs are currently clustered into approximately 20 000 contigs and about an equal number of apparent singlets. These clusters are continuously updated and annotated with respect to potential encoded protein products. More than 100 000 GSSs are similarly assembled and annotated by spliced alignment with EST and protein sequences. The ZmDB interface provides quick access to analytical tools for further sequence analysis. Every sequence record is linked to several display options and similarity search tools, including services for multiple sequence alignment, protein domain determination and spliced alignment. Furthermore, ZmDB provides web-based ordering of materials generated in the project, including ESTs, ordered collections of genomic sequences tagged with the RescueMu transposon and microarrays of amplified ESTs. ZmDB can be accessed at http://zmdb.iastate.edu/.

DNA Transposable Elements↗

Refined annotation of the Arabidopsis genome by complete expressed sequence tag mapping.

Expressed sequence tags (ESTs) currently encompass more entries in the public databases than any other form of sequence data. Thus, EST data sets provide a vast resource for gene identification and expression profiling. We have mapped the complete set of 176,915 publicly available Arabidopsis EST sequences onto the Arabidopsis genome using GeneSeqer, a spliced alignment program incorporating sequence similarity and splice site scoring. About 96% of the available ESTs could be properly aligned with a genomic locus, with the remaining ESTs deriving from organelle genomes and non-Arabidopsis sources or displaying insufficient sequence quality for alignment. The mapping provides verified sets of EST clusters for evaluation of EST clustering programs. Analysis of the spliced alignments suggests corrections to current gene structure annotation and provides examples of alternative and non-canonical pre-mRNA splicing. All results of this study were parsed into a database and are accessible via a flexible Web interface at http://www.plantgdb.org/AtGDB/.

Alternative Splicing↗

The maize genome contains a helitron insertion.

The maize mutation sh2-7527 was isolated in a conventional maize breeding program in the 1970s. Although the mutant contains foreign sequences within the gene, the mutation is not attributable to an interchromosomal exchange or to a chromosomal inversion. Hence, the mutation was caused by an insertion. Sequences at the two Sh2 borders have not been scrambled or mutated, suggesting that the insertion is not caused by a catastrophic reshuffling of the maize genome. The insertion is large, at least 12 kb, and is highly repetitive in maize. As judged by hybridization, sorghum contains only one or a few copies of the element, whereas no hybridization was seen to the Arabidopsis genome. The insertion acts from a distance to alter the splicing of the sh2 pre-mRNA. Three distinct intron-bearing maize genes were found in the insertion. Of most significance, the insertion bears striking similarity to the recently described DNA helicase-bearing transposable elements termed HELITRONS: Like Helitrons, the inserted sequence of sh2-7527 is large, lacks terminal repeats, does not duplicate host sequences, and was inserted between a host dinucleotide AT. Like Helitrons, the maize element contains 5' TC and 3' CTRR termini as well as two short palindromic sequences near the 3' terminus that potentially can form a 20-bp hairpin. Although the maize element lacks sequence information for a DNA helicase, it does contain four exons with similarity to a plant DEAD box RNA helicase. A second Helitron insertion was found in the maize genomic database. These data strongly suggest an active Helitron in the present-day maize genome.

Alternative Splicing↗

Comparative genomics of Arabidopsis and maize: prospects and limitations.

The completed Arabidopsis genome seems to be of limited value as a model for maize genomics. In addition to the expansion of repetitive sequences in maize and the lack of genomic micro-colinearity, maize-specific or highly-diverged proteins contribute to a predicted maize proteome of about 50,000 proteins, twice the size of that of Arabidopsis.

Arabidopsis↗

Gene structure identification with MyGV using cDNA evidence and protein homologs to improve ab initio predictions.

UNLABELLED: MyGV is an application to visualize (potentially genome-scale) gene structure annotation and prediction. The output of any external gene prediction program can be easily converted to a generalized format for input into MyGV. The application displays all input simultaneously in graphical representation, with a toggle option for a text-based view. Zooming capabilities allow detailed comparisons for specific genome locations. The tool is particularly helpful for refinement of ab initio predicted gene structures by spliced alignment with cDNA or protein homologs. AVAILABILITY: The program was written in Java and is freely available to non-commercial users by electronic download from http://bioinformatics.iastate.edu/bioinformatics2go/MyGV.

Animals↗

Comparison of RNA expression profiles based on maize expressed sequence tag frequency analysis and micro-array hybridization.

Assembly of 73,000 expressed sequence tags (ESTs) representing multiple organs and developmental stages of maize (Zea mays) identified approximately 22,000 tentative unique genes (TUGs) at the criterion of 95% identity. Based on sequence similarity, overlap between any two of nine libraries with more than 3,000 ESTs ranged from 4% to 20% of the constituent TUGs. The most abundant ESTs were recovered from only one or a minority of the libraries, and only 26 EST contigs had members from all nine EST sets (presumably representing ubiquitously expressed genes). For several examples, ESTs for different members of gene families were detected in distinct organs. To study this further, two types of micro-array slides were fabricated, one containing 5,534 ESTs from 10- to 14-d-old endosperm, and the other 4,844 ESTs from immature ear, estimated to represent about 2,800 and 2,500 unique genes, respectively. Each array type was hybridized with fluorescent cDNA targets prepared from endosperm and immature ear poly(A(+)) RNA. Although the 10- to 14-d-old postpollination endosperm TUGs showed only 12% overlap with immature ear TUGs, endosperm target hybridized with 94% of the ear TUGs, and ear target hybridized with 57% of the endosperm TUGs. Incomplete EST sampling of low-abundance transcripts contributes to an underestimate of shared gene expression profiles. Reassembly of ESTs at the criterion of 90% identity suggests how cross hybridization among gene family members can overestimate the overlap in genes expressed in micro-array hybridization experiments.

Contig Mapping↗

A compilation of soybean ESTs: generation and analysis.

Whole-genome sequencing is fundamental to understanding the genetic composition of an organism. Given the size and complexity of the soybean genome, an alternative approach is targeted random-gene sequencing, which provides an immediate and productive method of gene discovery. In this study, more than 120000 soybean expressed sequence tags (ESTs) generated from more than 50 cDNA libraries were evaluated. These ESTs coalesced into 16928 contigs and 17336 singletons. On average, each contig was composed of 6 ESTs and spanned 788 bases. The average sequence length submitted to dbEST was 414 bases. Using only those libraries generating more than 800 ESTs each and only those contigs with 10 or more ESTs each, correlated patterns of gene expression among libraries and genes were discerned. Two-dimensional qualitative representations of contig and library similarities were generated based on expression profiles. Genes with similar expression patterns and, potentially, similar functions were identified. These studies provide a rich source of publicly available gene sequences as well as valuable insight into the structure, function, and evolution of a model crop legume genome.

Contig Mapping↗

Computational modeling of gene structure in Arabidopsis thaliana.

Computational gene identification by sequence inspection remains a challenging problem. For a typical Arabidopsis thaliana gene with five exons, at least one of the exons is expected to have at least one of its borders predicted incorrectly by ab initio gene finding programs. More detailed analysis for individual genomic loci can often resolve the uncertainty on the basis of EST evidence or similarity to potential protein homologues. Such methods are part of the routine annotation process. However, because the EST and protein databases are constantly growing, in many cases original annotation must be re-evaluated, extended, and corrected on the basis of the latest evidence. The Arabidopsis Genome Initiative is undertaking this task on the whole-genome scale via its participating genome centers. The current Arabidopsis genome annotation provides an excellent starting point for assessing the protein repertoire of a flowering plant. More accurate whole-genome annotation will require the combination of high-throughput and individual gene experimental approaches and computational methods. The purpose of this article is to discuss tools available to an individual researcher to evaluate gene structure prediction for a particular locus.

Algorithms↗