PubMed Health⌕ Search

Biomedical subjects

Hank Wu

Publications and source records attributed to Hank Wu.

3 recordsLinked to original sources

The TIGR Plant Transcript Assemblies database.

The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.

DNA, Complementary↗

Development of Arabidopsis whole-genome microarrays and their application to the discovery of binding sites for the TGA2 transcription factor in salicylic acid-treated plants.

We have developed two long-oligonucleotide microarrays for the analysis of genome features in Arabidopsis thaliana, in particular for the high-throughput identification of transcription factor-binding sites. The first platform contains 190,000 probes representing the 2-kb regions upstream of all annotated genes at a density of seven probes per promoter. The second platform is divided into three chips, each of over 390,000 features, and represents the entire Arabidopsis genome at a density of one probe per 90 bases. Protein-DNA complexes resulting from the formaldehyde fixation of leaves of plants 2 h after exposure to 1 mm salicylic acid (SA) were immunoprecipitated using antibodies against the TGA2 transcription factor. After reversal of the cross-links and amplification, the resulting ChIP sample was hybridized to both platforms. High signal ratios of the ChIP sample versus raw chromatin for clusters of neighboring probes provided evidence for 51 putative binding sites for TGA2, including the only previously confirmed site in the promoter of PR-1 (At2g14610). Enrichment of several regions was confirmed by quantitative real-time PCR. Motif search revealed that the palindromic octamer TGACGTCA was found in 55% of the enriched regions. Interestingly, 15 of the putative binding sites for TGA2 lie outside the presumptive promoter regions. The effect of the 2-h SA treatment on gene expression was measured using Affymetrix ATH1 arrays, and SA-induced genes were found to be significantly over-represented among genes neighboring putative TGA2-binding sites.

Arabidopsis↗

Whole genome shotgun sequencing of Brassica oleracea and its application to gene discovery and annotation in Arabidopsis.

Through comparative studies of the model organism Arabidopsis thaliana and its close relative Brassica oleracea, we have identified conserved regions that represent potentially functional sequences overlooked by previous Arabidopsis genome annotation methods. A total of 454,274 whole genome shotgun sequences covering 283 Mb (0.44 x) of the estimated 650 Mb Brassica genome were searched against the Arabidopsis genome, and conserved Arabidopsis genome sequences (CAGSs) were identified. Of these 229,735 conserved regions, 167,357 fell within or intersected existing gene models, while 60,378 were located in previously unannotated regions. After removal of sequences matching known proteins, CAGSs that were close to one another were chained together as potentially comprising portions of the same functional unit. This resulted in 27,347 chains of which 15,686 were sufficiently distant from existing gene annotations to be considered a novel conserved unit. Of 192 conserved regions examined, 58 were found to be expressed in our cDNA populations. Rapid amplification of cDNA ends (RACE) was used to obtain potentially full-length transcripts from these 58 regions. The resulting sequences led to the creation of 21 gene models at 17 new Arabidopsis loci and the addition of splice variants or updates to another 19 gene structures. In addition, CAGSs overlapping already annotated genes in Arabidopsis can provide guidance for manual improvement of existing gene models. Published genome-wide expression data based on whole genome tiling arrays and massively parallel signature sequencing were overlaid on the Brassica-Arabidopsis conserved sequences, and 1399 regions of intersection were identified. Collectively our results and these data sets suggest that several thousand new Arabidopsis genes remain to be identified and annotated.

Arabidopsis↗