PubMed Health⌕ Search

Biomedical subjects

Sumio Sugano

Publications and source records attributed to Sumio Sugano.

At least 19 recordsLinked to original sources

Large-scale identification and characterization of alternative splicing variants of human gene transcripts using 56,419 completely sequenced and manually annotated full-length cDNAs.

We report the first genome-wide identification and characterization of alternative splicing in human gene transcripts based on analysis of the full-length cDNAs. Applying both manual and computational analyses for 56,419 completely sequenced and precisely annotated full-length cDNAs selected for the H-Invitational human transcriptome annotation meetings, we identified 6877 alternative splicing genes with 18 297 different alternative splicing variants. A total of 37,670 exons were involved in these alternative splicing events. The encoded protein sequences were affected in 6005 of the 6877 genes. Notably, alternative splicing affected protein motifs in 3015 genes, subcellular localizations in 2982 genes and transmembrane domains in 1348 genes. We also identified interesting patterns of alternative splicing, in which two distinct genes seemed to be bridged, nested or having overlapping protein coding sequences (CDSs) of different reading frames (multiple CDS). In these cases, completely unrelated proteins are encoded by a single locus. Genome-wide annotations of alternative splicing, relying on full-length cDNAs, should lay firm groundwork for exploring in detail the diversification of protein function, which is mediated by the fast expanding universe of alternative splicing variants.

Alternative Splicing↗

Common inheritance of chromosome Ia associated with clonal expansion of Toxoplasma gondii.

Toxoplasma gondii is a globally distributed protozoan parasite that can infect virtually all warm-blooded animals and humans. Despite the existence of a sexual phase in the life cycle, T. gondii has an unusual population structure dominated by three clonal lineages that predominate in North America and Europe, (Types I, II, and III). These lineages were founded by common ancestors approximately10,000 yr ago. The recent origin and widespread distribution of the clonal lineages is attributed to the circumvention of the sexual cycle by a new mode of transmission-asexual transmission between intermediate hosts. Asexual transmission appears to be multigenic and although the specific genes mediating this trait are unknown, it is predicted that all members of the clonal lineages should share the same alleles. Genetic mapping studies suggested that chromosome Ia was unusually monomorphic compared with the rest of the genome. To investigate this further, we sequenced chromosome Ia and chromosome Ib in the Type I strain, RH, and the Type II strain, ME49. Comparative genome analyses of the two chromosomal sequences revealed that the same copy of chromosome Ia was inherited in each lineage, whereas chromosome Ib maintained the same high frequency of between-strain polymorphism as the rest of the genome. Sampling of chromosome Ia sequence in seven additional representative strains from the three clonal lineages supports a monomorphic inheritance, which is unique within the genome. Taken together, our observations implicate a specific combination of alleles on chromosome Ia in the recent origin and widespread success of the clonal lineages of T. gondii.

Animals↗

Solution structure of the antifreeze-like domain of human sialic acid synthase.

The structure of the C-terminal antifreeze-like (AFL) domain of human sialic acid synthase was determined by NMR spectroscopy. The structure comprises one alpha- and two single-turn 3(10)-helices and two beta-strands, and is similar to those of the type III antifreeze proteins. Evolutionary trace analyses of the type III antifreeze protein family suggested that the class-specific residues in the human and bacterial AFL domains are important for their substrate binding, while the class-specific residues of the fish antifreeze proteins are gathered on the ice-binding surface.

Amino Acid Sequence↗

Activation of genes for growth factor and cytokine pathways late in chondrogenic differentiation of ATDC5 cells.

The mouse embryonal carcinoma cell line ATDC5 provides an excellent model system for chondrogenesis in vitro. To understand better the molecular mechanisms of endochondral bone formation, we investigated gene expression profiles during the differentiation course of ATDC5 cells, using an in-house microarray harboring full-length-enriched cDNAs. For 28 days following chondrogenic induction, 507 genes were up- or down-regulated at least 1.5-fold. These genes were classified into five clusters based on their expression patterns. Genes for growth factor and cytokine pathways were significantly enriched in the cluster characterized by increases in expression during late stages of chondrocyte differentiation. mRNAs for decorin and osteoglycin, which have been shown to bind to transforming growth factors-beta and bone morphogenetic proteins, respectively, were found in this cluster and were detected in hypertrophic chondrocytes of developing mouse bones by in situ hybridization analysis. Taken together with assigned functions of individual genes in the cluster, interdigitated interaction between a number of intercellular signaling molecules is likely to take place in the late chondrogenic stage for autocrine and paracrine regulation among chondrocytes, as well as for chemoattraction and stimulation of progenitor cells of other lineages.

Animals↗

Deletion of angiotensin-converting enzyme 2 accelerates pressure overload-induced cardiac dysfunction by increasing local angiotensin II.

Angiotensin-converting enzyme 2 (ACE2) is a carboxypeptidase that cleaves angiotensin II to angiotensin 1-7. Recently, it was reported that mice lacking ACE2 (ACE2(-/y) mice) exhibited reduced cardiac contractility. Because mechanical pressure overload activates the cardiac renin-angiotensin system, we used ACE2(-/y) mice to analyze the role of ACE2 in the response to pressure overload. Twelve-week-old ACE2(-/y) mice and wild-type (WT) mice received transverse aortic constriction (TAC) or sham operation. Sham-operated ACE2(-/y) mice exhibited normal cardiac function and had morphologically normal hearts. In response to TAC, ACE2(-/y) mice developed cardiac hypertrophy and dilatation. Furthermore, their hearts displayed decreased cardiac contractility and increased fetal cardiac gene induction, compared with WT mice. In response to chronic pressure overload, ACE2(-/y) mice developed pulmonary congestion and increased incidence of cardiac death compared with WT mice. On a biochemical level, cardiac angiotensin II concentration and activity of mitogen-activated protein (MAP) kinases were markedly increased in ACE2(-/y) mice in response to TAC. Administration of candesartan, an AT1 subtype angiotensin receptor blocker, attenuated the hypertrophic response and suppressed the activation of MAP kinases in ACE2(-/y) mice. Activation of MAP kinases in response to angiotensin II was greater in cardiomyocytes isolated from ACE2(-/y) mice than in those isolated from WT mice. ACE2 plays an important role in dampening the hypertrophic response to pressure overload mediated by angiotensin II. Disruption of this regulatory function may accelerate cardiac hypertrophy and shorten the transition period from compensated hypertrophy to cardiac failure.

Angiotensin II↗

DBTSS: DataBase of Human Transcription Start Sites, progress report 2006.

DBTSS was first constructed in 2002 based on precise, experimentally determined 5' end clones. Several major updates and additions have been made since the last report. First, the number of human clones has drastically increased, going from 190,964 to 1,359,000. Second, information about potential alternative promoters is presented because the number of 5' end clones is now sufficient to determine several promoters for one gene. Namely, we defined putative promoter groups by clustering transcription start sites (TSSs) separated by <500 bases. A total of 8308 human genes and 4276 mouse genes were found to have putative multiple promoters. Third, DBTSS provides detailed sequence comparisons of user-specified TSSs. Finally, we have added TSS information for zebrafish, malaria and schyzon (a red algae model organism). DBTSS is accessible at http://dbtss.hgc.jp.

Animals↗

Gene expression analysis of human hepatocellular carcinoma by using full-length cDNA library.

Hepatocellular carcinoma (HCC) is a leading cause of death worldwide. Hepatitis B virus (HBV) or hepatitis C virus (HCV) infection has been shown to cause hepatic carcinogenesis. A total 58,251 of cDNA clones of full-length cDNA libraries of HBV and HCV-infected HCC and their surrounding non-tumor tissues, respectively, were sequenced and analyzed by blasting against GENEBANK maintained by NCBI. About 180 and 279 of genes were shown an obviously increased and decreased expression patterns between HCC tissue and its adjacent non-tumor tissue. The candidate genes consisted of the genes encoded liver specific metabolism enzymes, secretory functional proteins, proteases and their inhibitors, protein chaperon, cell cycle components, apoptosis-related proteins, transcriptional factors, and DNA binding proteins. Several genes were further investigated by using real-time PCR to confirm the gene expression levels in at least 24 pairs of HCC tissues and adjacent non-tumor tissues. The results showed that genes encoded reticulon 4, RGS-1, antiplasmin, and kallikrein B were down-regulated with the average of 2.8, 8.5, 3.2, and 10.5-fold, respectively. Our results provide crucial candidate genes to develop clinical diagnosis and gene therapy of HCC.

Carcinoma, Hepatocellular↗

Transcriptome analyses of human genes and applications for proteome analyses.

By utilizing recently developed full-length cDNA technologies, large-scale cDNA sequencing was carried out by several cDNA projects. Now full-length cDNA resources cover the major part of the protein-coding human genes. Comprehensive analyses of the collected full-length cDNA data revealed not only the complete sequences of thousands of novel gene transcripts but also novel alternatively spliced isoforms of hitherto identified genes. However, it was not as easy as expected to deduce their encoded amino acid sequences based solely on the full-length cDNA sequences. It was neither always the case that the longest open reading frame corresponded to the real protein coding region nor that the first ATG was the translation initiator codon. Also, proteome-wide mass-spectrometry analysis has shown that there is an unexpectedly large population of small proteins, encoded by so-called upstream open reading frames, within the cell. Since sound manual annotations by experts were still indispensable to address these problems, an international meeting to make transcriptome-wide functional annotations of cDNAs was held, namely the H-invitational. In this meeting, functional annotations were made both manually and computationally for most of the pre-existing full-length cDNAs collected from world-wide cDNA projects. The achieved integrated information for each of the cDNAs was published as a database. It was also shown that the full-length cDNA data were useful for identifying alternative splicing variants, exact transcriptional start sites of the mRNAs and the adjacent promoter regions. Rapidly accumulating genome data as well as versatile use of the transcriptome information will shortly lay a firm foundation for proteome-level understanding of human gene networks.

Alternative Splicing↗

Diversification of transcriptional modulation: large-scale identification and characterization of putative alternative promoters of human genes.

By analyzing 1,780,295 5'-end sequences of human full-length cDNAs derived from 164 kinds of oligo-cap cDNA libraries, we identified 269,774 independent positions of transcriptional start sites (TSSs) for 14,628 human RefSeq genes. These TSSs were clustered into 30,964 clusters that were separated from each other by more than 500 bp and thus are very likely to constitute mutually distinct alternative promoters. To our surprise, at least 7674 (52%) human RefSeq genes were subject to regulation by putative alternative promoters (PAPs). On average, there were 3.1 PAPs per gene, with the composition of one CpG-island-containing promoter per 2.6 CpG-less promoters. In 17% of the PAP-containing loci, tissue-specific use of the PAPs was observed. The richest tissue sources of the tissue-specific PAPs were testis and brain. It was also intriguing that the PAP-containing promoters were enriched in the genes encoding signal transduction-related proteins and were rarer in the genes encoding extracellular proteins, possibly reflecting the varied functional requirement for and the restricted expression of those categories of genes, respectively. The patterns of the first exons were highly diverse as well. On average, there were 7.7 different splicing types of first exons per locus partly produced by the PAPs, suggesting that a wide variety of transcripts can be achieved by this mechanism. Our findings suggest that use of alternate promoters and consequent alternative use of first exons should play a pivotal role in generating the complexity required for the highly elaborated molecular systems in humans.

Base Sequence↗

Centaurin-alpha1 is a phosphatidylinositol 3-kinase-dependent activator of ERK1/2 mitogen-activated protein kinases.

Centaurin-alpha1 is known to be a phosphatidylinositol 3,4,5-triphosphate (PIP3)-binding protein that has two pleckstrin homology domains and a putative ADP ribosylation factor GTPase-activating protein domain. However, the physiological function of centaurin-alpha1 is still not understood. Here we have shown that transient expression of centaurin-alpha1 in COS-7 cells results in specific activation of ERK, and the activation is inhibited by co-expression of a dominant negative form of Ras. We have also found that a mutant form of centaurin-alpha1 that is unable to bind PIP3 fails to induce ERK activation and that a phosphatidylinositol 3-kinase inhibitor LY294002 inhibits centaurin-alpha1-dependent ERK activation. Furthermore, transient knockdown of centaurin-alpha1 by small interfering RNAs results in reduced ERK activation after epidermal growth factor stimulation in T-REx 293 cells. These results suggest that centaurin-alpha1 contributes to ERK activation in growth factor signaling, linking the PI3K pathway to the ERK mitogen-activated protein kinase pathway through its ability to interact with PIP3.

Adaptor Proteins, Signal Transducing↗

Investigation of protein functions through data-mining on integrated human transcriptome database, H-Invitational database (H-InvDB).

H-Invitational Database (H-InvDB; ) is a human transcriptome database, containing integrative annotation of 41,118 full-length cDNA clones originated from 21,037 loci. H-InvDB is a product of the H-Invitational project, an international collaboration to systematically and functionally validate human genes by analysis of a unique set of high quality full-length cDNA clones using automatic annotation and human curation under unified criteria. Here, 19,574 proteins encoded by these cDNAs were classified into 11,709 function-known and 7865 function-unknown hypothetical proteins by similarity with protein databases and motif prediction (InterProScan). The proportion of "hypothetical proteins" in H-InvDB was as high as 40.4%. In this study, we thus conducted data-mining in H-InvDB with the aim of assigning advanced functional annotations to those hypothetical proteins. First, by data-mining in the H-InvDB version of GTOP, we identified 337 SCOP domains within 7865 H-Inv hypothetical proteins. Second, by data-mining of predicted subcellular localization by SOSUI and TMHMM in H-InvDB, we found 1032 transmembrane proteins within H-Inv hypothetical proteins. These results clearly demonstrate that structural prediction is effective for functional annotation of proteins with unknown functions. All the data in H-InvDB are shown in two main views, the cDNA view and the Locus view, and five auxiliary databases with web-based viewers; DiseaseInfo Viewer, H-ANGEL, Clustering Viewer, G-integra and TOPO Viewer; the data also are provided as flat files and XML files. The data consists of descriptions of their gene structures, novel alternative splicing isoforms, functional RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs in relation with orphan diseases, gene expression profiling, and comparisons with mouse full-length cDNAs in the context of molecular evolution. This unique integrative platform for conducting in silico data-mining represents a substantial contribution to resources required for the exploration of human biology and pathology.

Amino Acid Sequence↗

Cell surface labeling and mass spectrometry reveal diversity of cell surface markers and signaling molecules expressed in undifferentiated mouse embryonic stem cells.

Although interactions between cell surface proteins and extracellular ligands are key to initiating embryonic stem cell differentiation to specific cell lineages, the plasma membrane protein components of these cells are largely unknown. We describe here a group of proteins expressed on the surface of the undifferentiated mouse embryonic stem cell line D3. These proteins were identified using a combination of cell surface labeling with biotin, subcellular fractionation of plasma membranes, and mass spectrometry-based protein identification technology. From 965 unique peptides carrying biotin labels, we assigned 324 proteins including 235 proteins that have putative signal sequences and/or transmembrane segments. Receptors, transporters, and cell adhesion molecules were the major classes of proteins identified. Besides known cell surface markers of embryonic stem cells, such as alkaline phosphatase, the analysis identified 59 clusters of differentiation-related molecules and more than 80 components of multiple cell signaling pathways that are characteristic of a number of different cell lineages. We identified receptors for leukemia-inhibitory factor, interleukin 6, and bone morphogenetic protein, which play critical roles in the maintenance of undifferentiated mouse embryonic stem cells. We also identified receptors for growth factors/cytokines, such as fibroblast growth factor, platelet-derived growth factor, ephrin, Hedgehog, and Wnt, which transduce signals for cell differentiation and embryonic development. Finally we identified a variety of integrins, cell adhesion molecules, and matrix metalloproteases. These results suggest that D3 cells express diverse cell surface proteins that function to maintain pluripotency, enabling cells to respond to various external signals that initiate differentiation into a variety of cell types.

Amino Acid Sequence↗

A novel method for development of malaria vaccines using full-length cDNA libraries.

We describe a novel method to screen malaria DNA vaccine candidates using a full-length cDNA library and a murine malaria infection model. For the development of effective malaria vaccines, much effort has been made with meager success. The completion of genome sequencing of Plasmodium falciparum has provided invaluable information for achieving this goal. We have been studying full-length cDNA libraries of malaria parasites as a part of genome analysis. Mice vaccinated with a DNA vaccine consisting of 2000 pooled clones showed significantly prolonged survival after challenge infection. In addition, spleen cells of vaccinated mice produced augmented levels of IL-2 and IFN-gamma when incubated with the crude parasite antigens, indicating that cellular immunity plays an important role in the protection. This approach will not only form the basis for development of malaria vaccines but will also be applicable to other parasites and pathogenic microorganisms.

Animals↗

Substitution rate and structural divergence of 5'UTR evolution: comparative analysis between human and cynomolgus monkey cDNAs.

The substitution rate and structural divergence in the 5'-untranslated region (UTR) were investigated by using human and cynomolgus monkey cDNA sequences. Due to the weaker functional constraint in the UTR than in the coding sequence, the divergence between humans and macaques would provide a good estimate of the nucleotide substitution rate and structural divergence in the 5'UTR. We found that the substitution rate in the 5'UTR (K5UTR) averaged approximately 10%-20% lower than the synonymous substitution rate (Ks). However, both the K5UTR and nonsynonymous substitution rate (Ka) were significantly higher in the testicular cDNAs than in the brain cDNAs, whereas the Ks did not differ. Further, an in silico analysis revealed that 27% (169/622) of macaque testicular cDNAs had an altered exon-intron structure in the 5'UTR compared with the human cDNAs. The fraction of cDNAs with an exon alteration was significantly higher in the testicular cDNAs than in the brain cDNAs. We confirmed by using reverse transcriptase-polymerase chain reaction that about one-third (6/16) of in silico "macaque-specific" exons in the 5'UTR were actually macaque specific in the testis. The results imply that positive selection increased K5UTR and structural alteration rate of a certain fraction of genes as well as Ka. We found that both positive and negative selection can act on the 5'UTR sequences.

Animals↗

Genome-wide analysis reveals strong correlation between CpG islands with nearby transcription start sites of genes and their tissue specificity.

It has been envisaged that CpG islands are often observed near the transcriptional start sites (TSS) of housekeeping genes. However, neither the precise positions of CpG islands relative to TSS of genes nor the correlation between the presence of the CpG islands and the expression specificity of these genes is well-understood. Using thousands of sequences with known TSS in human and mouse, we found that there is a clear peak in the distribution of CpG islands around TSS in the genes of these two species. Thus, we classified human (mouse) genes into 6600 (2948) CpG+ genes and 2619 (1830) CpG- ones, based on the presence of a CpG island within the -100: +100 region. We estimated the degree of each gene being a housekeeper by the number of cDNA libraries where its ESTs were detected. Then, the tendency that a gene lacking CpG islands around its TSS is expressed with a higher degree of tissue specificity turned out to be evolutionarily conserved. We also confirmed this tendency by analyzing the gene ontology annotation of classified genes. Since no such clear correlation was found in the control data (mRNAs, pre-mRNAs, and chromosome banding pattern), we concluded that the effect of a CpG island near the TSS should be more important than the global GC content of the region where the gene resides.

Animals↗

5'SAGE: 5'-end Serial Analysis of Gene Expression database.

To comprehensively identify transcription start sites and the frequencies of individual mRNAs in human cell libraries, a method of 5' end Serial Analysis of Gene Expression (SAGE) was developed recently, which makes it possible to collect a large amount of start site information, and subsequently, we have established a related database server called 5'SAGE. This database displays the observed frequencies of individual 5' end SAGE tags and previously unknown transcription start sites in the promoter regions, introns and intergenic regions of known genes. 5'SAGE will be useful for analyzing promoter regions and start site variation in different tissues, and is freely available at http://5sage.gi.k.u-tokyo.ac.jp/.

5' Flanking Region↗

dbQSNP: a database of SNPs in human promoter regions with allele frequency information determined by single-strand conformation polymorphism-based methods.

We present a database, dbQSNP (http://qsnp.gen.kyushu-u.ac.jp/), that provides sequence and allele frequency information for single-nucleotide polymorphisms (SNPs) located in the promoter regions of human genes, which were defined by the 5' ends of full-length cDNA clones. We searched for the SNPs in these regions by sequencing or single-strand conformation polymorphism (SSCP) analysis. The allele frequencies of the identified SNPs in two ethnic groups were quantified by SSCP analyses of pooled DNA samples. The accuracy of our estimation is supported by strong correlations between the frequencies in our data and those in other databases for the same ethnic groups. The frequencies vary considerably between the two ethnic groups studied, suggesting the need for population-based collections and allele frequency determination of SNPs, in, e.g., association studies of diseases. We show profiles of SNP densities that are characteristic of transcription start site regions. A fraction of the SNPs revealed a significantly different allele frequency between the groups, suggesting differential selection of the genes involved.

Alleles↗