PubMed Health⌕ Search

Biomedical subjects

Yoonsoo Hahn

Publications and source records attributed to Yoonsoo Hahn.

14 recordsLinked to original sources

Transcriptome mining and comparative genomics reveal 36 putative novel marafivirus species and conserved evolution of the marafibox regulatory element.

BACKGROUND: Marafiviruses are plant-infecting RNA viruses associated with several economically important crops, but their genomic diversity remains incompletely characterized. OBJECTIVE: This study aimed to identify previously unrecognized marafivirus genomes and investigate their genomic features and evolutionary relationships. METHODS: Publicly available plant transcriptome datasets were systematically mined to detect marafivirus-like sequences. Recovered genomes were analyzed using comparative sequence analysis, phylogenetic reconstruction, and genome organization characterization. RESULTS: A total of 62 marafivirus-like genomes were recovered from 33 independent sources representing diverse plant hosts. Polyprotein-based comparative and phylogenetic analyses grouped these genomes into 36 lineages likely representing novel species. All newly identified viruses clustered within the Marafivirus clade. Genome organization analysis revealed conserved polyprotein architecture and widespread presence of the marafibox promoter element. Conservation of additional open reading frames among closely related isolates aided identification of potentially functional genes. CONCLUSION: These findings substantially expand the known diversity of marafiviruses and demonstrate the effectiveness of transcriptome mining for discovering previously unrecognized plant viruses.

Phylogeny↗

Evolution and expression of chimeric POTE-actin genes in the human genome.

We previously described a primate-specific gene family, POTE, that is expressed in many cancers but in a limited number of normal organs. The 13 POTE genes are dispersed among eight different chromosomes and evolved by duplications and remodeling of the human genome from an ancestral gene, ANKRD26. Based on sequence similarity, the POTE gene family members can be divided into three groups. By genome database searches, we identified an actin retroposon insertion at the carboxyl terminus of one of the ancestral POTE paralogs. By Northern blot analysis, we identified the expected 7.5-kb POTE-actin chimeric transcript in a breast cancer cell line. The protein encoded by the POTE-actin transcript is predicted to be 120 kDa in size. Using anti-POTE mAbs that recognize the amino-terminal portion of the POTE protein, we detected the 120-kDa POTE-actin fusion protein in breast cancer cell lines known to express the fusion transcript. These data demonstrate that insertion of a retroposon produced an altered functional POTE gene. This example indicates that new functional human genes can evolve by insertion of retroposons.

Actins↗

Human-specific nonsense mutations identified by genome sequence comparisons.

The comparative study of the human and chimpanzee genomes may shed light on the genetic ingredients for the evolution of the unique traits of humans. Here, we present a simple procedure to identify human-specific nonsense mutations that might have arisen since the human-chimpanzee divergence. The procedure involves collecting orthologous sequences in which a stop codon of the human sequence is aligned to a non-stop codon in the chimpanzee sequence and verifying that the latter is ancestral by finding homologs in other species without a stop codon. Using this procedure, we identify nine genes (CML2, FLJ14640, MT1L, NPPA, PDE3B, SERPINA13, TAP2, UIP1, and ZNF277) that would produce human-specific truncated proteins resulting in a loss or modification of the function. The premature terminations of CML2, MT1L, and SERPINA13 genes appear to abolish the original function of the encoded protein because the mutation removes a major part of the known active site in each case. The other six mutated genes are either known or presumed to produce functionally modified proteins. The mutations of five genes (CML2, FLJ14640, MT1L, NPPA, TAP2) are known or predicted to be polymorphic in humans. In these cases, the stop codon alleles are more prevalent than the ancestral allele, suggesting that the mutant alleles are approaching fixation since their emergence during the human evolution. The findings support the notion that functional modification or inactivation of genes by nonsense mutation is a part of the process of adaptive evolution and acquisition of species-specific features.

Animals↗

POTE paralogs are induced and differentially expressed in many cancers.

To identify new antigens that are targets for the immunotherapy of prostate and breast cancer, we used expressed sequence tag and genomic databases and discovered POTE, a new primate-specific gene family. Each POTE gene encodes a protein that contains three domains, although the proteins vary greatly in size. The NH2-terminal domain is novel and has properties of an extracellular domain but does not contain a signal sequence. The second and third domains are rich in ankyrin repeats and spectrin-like helices, respectively. The protein encoded by POTE-21, the first family member discovered, is localized on the plasma membrane of the cell. In humans, 13 highly homologous paralogs are dispersed among eight chromosomes. The expression of POTE genes in normal tissues is restricted to prostate, ovary, testis, and placenta. A survey of several cancer samples showed that POTE was expressed in 6 of 6 prostate, 12 of 13 breast, 5 of 5 colon, 5 of 6 lung, and 4 of 5 ovarian cancers. To determine the relative expression of each POTE paralog in cancer and normal samples, we employed a PCR-based cloning and analysis method. We found that POTE-2alpha, POTE-2beta, POTE-2gamma, and POTE-22 are predominantly expressed in cancers whereas POTE expression in normal tissues is somewhat more diverse. Because POTE is primate specific and is expressed in testis and many cancers but only in a few normal tissues, we conclude POTE is a new primate-specific member of the cancer-testis antigen family. It is likely that POTE has a unique role in primate biology.

Base Sequence↗

Duplication and extensive remodeling shaped POTE family genes encoding proteins containing ankyrin repeat and coiled coil domains.

The POTE family genes encode a highly homologous group of primate-specific proteins that contain ankyrin repeats and coiled coil domains. At least 13 paralogous POTE family genes are found on 8 human chromosomes (2, 8, 13, 14, 15, 18, 21 and 22), which can be sorted into 3 groups based on sequence similarity. We identified by a database search a group of additional human ankyrin repeat domain proteins, of which ANKRD26 and ANKRD30A are the best characterized; these are more distant homologs of POTE family proteins. A comprehensive comparison of the genomic organization indicates that ANKRD26 has the genomic structure of the possible ancestor of ANKRD30A and all POTE family genes. Extensive remodeling involving segmental loss and internal duplication appears to have reshaped the ANKRD30A and POTE family genes after the primal duplication of the ancestor gene. We also identified a mouse homolog of human ANKRD26, but failed to find a mouse homolog that bears the structural characteristics of any of the POTE family of proteins. The mouse Ankrd26 may serve as a useful model for the study of the function of human ANKRD26, ANKRD30A and POTE family proteins.

Animals↗

Transcriptome analysis of human gastric cancer.

To elucidate the genetic events associated with gastric cancer, 124,704 cDNA clones were collected from 37 human gastric cDNA libraries, including 20 full-length enriched cDNA libraries of gastric cancer cell lines and tissues from Korean patients. An analysis of the collected ESTs revealed that 97,930 high-quality ESTs coalesced into 13,001 clusters, of which 11,135 clusters (85.6%) were annotated to known ESTs. The analysis of the full-length cDNAs also revealed that 4862 clusters (51.7%) contained at least one putative full-length cDNA clone with an initiation codon, with the average length of the 5' UTR of 140 bp. A large number appear to have a diverse transcription start site (TSS). An examination of the TSS of some genes, such as TEGT and GAPD, using 5' RACE revealed that the predicted TSSs are actually found in human gastric cancer cells and that several TSSs differ depending on the specific gastric cell line. Furthermore, of the human gastric ESTs, 766 genes (9.5%) were present as putative alternatively spliced variants. Confirmation of the predicted spliced isoforms using RT-PCR showed that the predicted isoforms exist in gastric cancer cells and some isoforms coexist in gastric cell lines. These results provide potentially useful information for elucidating the molecular mechanisms associated with gastric oncogenesis.

5' Untranslated Regions↗

Structure and expression of the zebrafish mest gene, an ortholog of mammalian imprinted gene PEG1/MEST.

PEG1/MEST is a paternally expressed gene in placental mammals. Here, we report identification of zebrafish (Danio rerio) gene mest, an ortholog of mammalian PEG1/MEST. Zebrafish mest encodes a polypeptide of 344 amino acids and shows a significant similarity to mammalian orthologs. Zebrafish mest is present as a single copy in the zebrafish genome and is closely linked to copg2 as in mammals. It is notable that 10 of 11 intron positions in mest are conserved among mammalian PEG1/MEST genes, indicating that the genomic organization and linkage between mest and copg2 loci was established in ancient vertebrates. Zebrafish mest is expressed in blastula, segmentation, and larval stages, exhibiting gradually increased expression as the development proceeds. Allelic expression analysis in hybrid larvae shows that both parental alleles are transcribed. We also observed one-codon alternative splicing involving an alternative usage of the two consecutive splice acceptors of intron 1, generating two protein isoforms with different lengths of a single amino acid.

Alternative Splicing↗

Identification of nine human-specific frameshift mutations by comparative analysis of the human and the chimpanzee genome sequences.

MOTIVATION: The recent release of the draft sequence of the chimpanzee genome is an invaluable resource for finding genome-wide genetic differences that might explain phenotypic differences between humans and chimpanzees. AVAILABILITY: In this paper, we describe a simple procedure to identify potential human-specific frameshift mutations that occurred after the divergence of human and chimpanzee. The procedure involves collecting human coding exons bearing insertions or deletions compared with the chimpanzee genome and identification of homologs from other species, in support of the mutations being human-specific. Using this procedure, we identified nine genes, BASE, DNAJB3, FLJ33674, HEJ1, NTSR2, RPL13AP, SCGB1D4, WBSCR27 and ZCCHC13, that show human-specific alterations including truncations of the C-terminus. In some cases, the frameshift mutation results in gene inactivation or decay. In other cases, the altered protein seems to be functional. This study demonstrates that even the unfinished chimpanzee genome sequence can be useful in identifying modification of genes that are specific to the human lineage and, therefore, could potentially be relevant to the study of the acquisition of human-specific traits.

Amino Acid Sequence↗

Finding fusion genes resulting from chromosome rearrangement by analyzing the expressed sequence databases.

Chromosomal rearrangements resulting in gene fusions are frequently involved in carcinogenesis. Here, we describe a semiautomatic procedure for identifying fusion gene transcripts by using publicly available mRNA and EST databases. With this procedure, we have identified 96 transcript sequences that are derived from 60 known fusion genes. Also, 47 or more additional sequences appear to be derived from 20 or more previously unknown putative fusion genes. We have experimentally verified the presence of a previously unknown IRA1/RGS17 fusion in the breast cancer cell line MCF7. The fusion gene encodes the full-length RGS17 protein, a regulator of G protein-coupled signaling, under the control of the IRA1 gene promoter. This study demonstrates that databases of ESTs can be used to discover fusion genes resulting from structural rearrangement of chromosomes.

Cell Line, Tumor↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

NGEP, a gene encoding a membrane protein detected only in prostate cancer and normal prostate.

We identified a gene (NGEP) that is expressed only in prostate cancer and normal prostate. The two NGEP transcripts are 0.9 kb and 3.5 kb in size and are generated by a differential splicing event. The short variant (NGEP-S) is derived from four exons and encodes a 20-kDa intracellular protein. The long form (NGEP-L) is derived from 18 exons and encodes a 95-kDa protein that is predicted to contain seven-membrane-spanning regions. In situ hybridization shows that NGEP mRNA is localized in epithelial cells of normal prostate and prostate cancers. Immunocytochemical analysis of cells transfected with NGEP cDNAs containing a Myc epitope tag at the carboxyl terminus shows that the protein encoded by the short transcript is localized in the cytoplasm, whereas the protein encoded by the long transcript is present on the plasma membrane. Because of its selective expression in prostate cancer and its presence on the cell surface, NGEP-L is a promising target for the antibody-based therapies of prostate cancer.

Animals↗

Rapid grouping of monoclonal antibodies based on their topographical epitopes by a label-free competitive immunoassay.

Topography of epitopes of monoclonal antibodies (MAbs), identified as the mutual competition of the MAbs, can be valuable indicators for the biological functions of MAbs. However, the determination of topographical epitopes is not performed before the functional screening of MAbs, because the requirement for purifying and labeling of MAbs makes the mapping experiment difficult, particularly in the early stage of MAb production. Here we describe a new label-free competitive enzyme-linked immunosorbent assay (LFC-ELISA) for the rapid grouping of MAbs based on the topography of their epitopes. In the LFC-ELISA, the immune complex formed by a competitor, MAb#2, and an antigen is challenged by an indicator, MAb#1 that had been captured on the ELISA plate through a secondary antibody. The MAb#2-antigen immune complex is trapped by MAb#1 only if MAb#1 reacts with an epitope different from that of MAb#2. The immune complex (MAb#2-antigen-MAb#1) is detected with an enzyme-labeled reagent specific to a tag on the antigen. Our experiments using different anti-CD30 MAbs and a CD30-Fc fusion protein as the antigen revealed that the LFC-ELISA performed well with MAbs of different isotypes (IgG1, IgG2a, and IgG2b), and in a practical range of MAb concentrations (0.3-10 microg/ml) and affinities (0.9-13 nM of Kd). We obtained pairwise competition data from all 26 anti-CD30 MAbs. We then utilized a cluster analysis and a bootstrap method to analyze the competition data for grouping of the MAbs. This objective and automated analysis identified eight distinct topographical epitopes on CD30. The reactivity of the anti-CD30 MAbs in immunoblot, and their inhibiting activity on CD30-CD30-ligand binding correlated with the topographical epitopes. The results show that the LFC-ELISA combined with cluster analysis is a useful new method for grouping MAbs based on their topographical epitopes and can be used in the early stage of MAb production. One useful application is to identify MAbs reacting with different epitopes from a large number of MAbs so that the most appropriate MAbs can be selected for therapeutic use.

Antibodies, Monoclonal↗

Gene cataloging and expression profiling in human gastric cancer cells by expressed sequence tags.

To understand the molecular mechanism associated with gastric carcinogenesis, we identified genes expressed in gastric cancer cell lines and tissues. Of 97,609 high-quality ESTs sequenced from 36 cDNA libraries, 92,545 were coalesced into 10,418 human Unigene clusters (Build 151). The gene expression profile was produced by counting the cluster frequencies in each library. Although the profiles of highly expressed genes varied greatly from library to library, those genes related to cell structure formation, heat shock proteins, the glycolysis pathway, and the signaling pathway were highly represented in human gastric cancer cell lines and in primary tumors. Conversely, the genes encoding immunoglobulins, ribosomal proteins, and digestive proteins were down-regulated in gastric cancer cell lines and tissues compared to normal tissues. The transcription levels of some of these genes were confirmed by RT-PCR. We found that genes related to cell adhesion, apoptosis, and cytoskeleton formation were particularly up-regulated in the gastric cancer cell lines established from malignant ascites compared to those from primary tumors. This comprehensive molecular profiling of human gastric cancer should be useful for elucidating the genetic events associated with human gastric cancer.

Base Sequence↗

TEPP, a new gene specifically expressed in testis, prostate, and placenta and well conserved in chordates.

We have combined computer-based screening and experimental expression analysis to identify genes that are expressed in normal prostate and/or prostate cancer but not in essential human tissues. Using this approach we identified a new gene that is specifically expressed in testis, prostate, and placenta. The gene has one major transcript of 1.0kb in size and encodes for a protein of 30.7kDa molecular weight. We named this gene TEPP (expressed in testis, prostate, and placenta). The amino acid sequence analysis of TEPP using SignalP program shows that it has a signal peptide with a predicted cleavage site between amino acids 19 and 20, indicating that it might be a secreted protein. Analysis of the predicted TEPP orthologs from different species shows that these proteins are highly conserved in chordates. In addition we have identified a splice variant of TEPP, which encodes a 37kDa protein. In conclusion, a combination of bioinformatic and molecular approaches is useful in the identification of genes expressed in specific tissues. Selective expression of TEPP in testis, prostate, and in placenta and its high conservation among different species indicate that TEPP might have a role in reproductive biology.

Amino Acid Sequence↗