PubMed Health⌕ Search

Biomedical subjects

Hwan-Gue Cho

Publications and source records attributed to Hwan-Gue Cho.

4 recordsLinked to original sources

GAME: a simple and efficient whole genome alignment method using maximal exact match filtering.

In this paper, we present a simple and efficient whole genome alignment method using maximal exact match (MEM). The major problem with the use of MEM anchor is that the number of hits in non-homologous regions increases exponentially when shorter MEM anchors are used to detect more homologous regions. To deal with this problem, we have developed a fast and accurate anchor filtering scheme based on simple match extension with minimum percent identity and extension length criteria. Due to its simplicity and accuracy, all MEM anchors in a pair of genomes can be exhaustively tested and filtered. In addition, by incorporating the translation technique, the alignment quality and speed of our genome alignment algorithm have been further improved. As a result, our genome alignment algorithm, GAME (Genome Alignment by Match Extension), performs competitively over existing algorithms and can align large whole genomes, e.g., A. thaliana, without the requirement of typical large memory and parallel processors. This is shown using an experiment which compares the performance of BLAST, BLASTZ, PatternHunter, MUMmer and our algorithm in aligning all 45 pairs of 10 microbial genomes. The scalability of our algorithm is shown in another experiment where all pairs of five chromosomes in A. thaliana were compared.

Algorithms↗

AngioDB: database of angiogenesis and angiogenesis-related molecules.

Angiogenesis is the formation of new capillaries sprouting from pre-existing vessels. Angiogenesis occurs in a variety of normal physiological and pathological conditions and is regulated by a balance of stimulatory and inhibitory angiogenic factors. The control of this balance may fail and result in the formation of a pathologic capillary network during the development of many diseases. Therefore, we developed the angiogenesis database (AngioDB), which can provide a signaling network of angiogenesis-related biomolecules in human. Each record of AngioDB consisted of 12 fields and was developed by using a relational database management system. For the retrieval of data, Active Server Page (ASP) technology was integrated in this system. Users can access the database by a query or imagemap browsing program. The retrieving system also provides a list of angiogenesis-related molecules classified by three categories, and the database has an external link to NCBI databases. AngioDB is available via the Internet at http://angiodb.snu.ac.kr/.

Amino Acid Sequence↗

An automatic block and spot indexing with k-nearest neighbors graph for microarray image analysis.

MOTIVATION: In this paper, we propose a fully automatic block and spot indexing algorithm for microarray image analysis. A microarray is a device which enables a parallel experiment of ten to hundreds of thousands of test genes in order to measure gene expression. Due to this huge size of experimental data, automated image analysis is gaining importance in microarray image processing systems. Currently, most of the automated microarray image processing systems require manual block indexing and, in some cases, spot indexing. If the microarray image is large and contains a lot of noise, it is very troublesome work. In this paper, we show it is possible to locate the addresses of blocks and spots by applying the Nearest Neighbors Graph Model. Also, we propose an analytic model for the feasibility of block addressing. Our analytic model is validated by a large body of experimental results. RESULTS: We demonstrate the features of automatic block detection, automatic spot addressing, and correction of the distortion and skewedness of each microarray image.

Algorithms↗

Analysis of common k-mers for whole genome sequences using SSB-tree.

As sequenced genomes become larger and sequencing process becomes faster, there is a need to develop a tool to analyze sequences in the whole genomic scale. However, on-memory algorithms such as suffix tree and suffix array are not applicable to the analysis of whole genome sequence set, since the size of individual whole genome ranges from several million base pairs to hundreds billion base pairs. In order to effectively manipulate the huge sequence data, it is necessary to use the indexed data structure for external memory. In this paper, we introduce a workbench called SequeX for the analysis and visualization of whole genome sequences using SSB-tree (Static SB-tree). It consists of two parts: the analysis query subsystem and the visualization subsystem. The query subsystem supports various transactions such as pattern matching, k-occurrence, and k-mer analysis. The visualization subsystem helps biologists to easily understand whole genome structure and feature by sequence viewer, annotation viewer, CGR (Chaos Game Representation) viewer, and k-mer viewer. The system also supports a user-friendly programming interface based on Java script for batch processing and the extension for a specific purpose of a user. SequeX can be used to identify conserved genes or sequences by the analysis of the common k-mers and annotation. We analyze the common k-mer for 72 microbial genomes announced by Entrez, and find an interesting biological fact that the longest common k-mer for 72 sequences is 11-mer, and only 11 such sequences exist. Finally we note that many common k-mers occur in conserved region such as CDS, rRNA, and tRNA.

Archaea↗