PubMed Health⌕ Search

Biomedical subjects

Satoshi Fukuchi

Publications and source records attributed to Satoshi Fukuchi.

15 recordsLinked to original sources

Gene cluster analysis method identifies horizontally transferred genes with high reliability and indicates that they provide the main mechanism of operon gain in 8 species of gamma-Proteobacteria.

The formation mechanism of operons remains unresolved: operons may form by rearrangements within a genome or by acquisition of genes from other species, that is, horizontal gene transfer (HGT). One hindrance to its elucidation is the unavailability of a method to accurately identify HGT, although it is generally considered to occur. It is critically important first to select horizontally transferred (HT) genes reliably and then to determine the extent to which HGT is involved in operon formation. For this purpose, we considered indels in terms of gene clusters instead of individual genes and chose candidates of HT genes in 8 species of Escherichia, Shigella, and Salmonella based on the minimization of indels. To select a benchmark set of positively HT genes against which we can evaluate the candidate set, we devised another procedure using intergenetic alignments. Comparison with the benchmark set demonstrated the absence of a significant number of false positives in the candidate set, showing the high reliability of the method. Analyses of Escherichia coli K-12 operons revealed that although approximately 20 operons were probably gained from the last common ancestor of the 8 gamma-proteobacteria, deletion of intervening genes accounts for the formation of no operons, whereas horizontal transfer expanded 2 operons and introduced 4 entire operons. Based on these observations and reasoning, we suggest that the main mechanism of operon gain is HGT rather than intragenomic rearrangements. We propose that genes with related essential functions tend to reside in conserved operons, whereas genes in nonconserved operons mostly confer slight advantage to the organisms and frequently undergo horizontal transfer and decay. HT genes constitute at least 5.5% of the genes in the 8 species and approximately 45% of which originate from other gamma-proteobacteria. Genes involved in viral functions and mobile and extrachromosomal element functions are HT more often than expected. This finding indicates frequent mediation of HGT by bacteriophages. On the other hand, not only informational genes (those involved in transcription, translation, and related processes) but also operational genes (those involved in housekeeping) are HT less frequently than expected.

Cluster Analysis↗

Exploration and grading of possible genes from 183 bacterial strains by a common protocol to identification of new genes: Gene Trek in Prokaryote Space (GTPS).

A large number of complete microorganism genomes has been sequenced and submitted to the public database and then incorporated into our complete genome database, Genome Information Broker (GIB, http://gib.genes.nig.ac.jp/). However, when comparative genomics is carried out, researchers must be aware that there are protein-coding genes not confirmed by homology or motif search and that reliable protein-coding genes are missing. Therefore, we developed a protocol (Gene Trek in Prokaryote Space, GTPS) for finding possible protein-coding genes in bacterial genomes. GTPS assigns a degree of reliability to predicted protein-coding genes. We first systematically applied the protocol to the complete genomes of all 123 bacterial species and strains that were publicly available as of July 2003, and then to those of 183 species and strains available as of September 2004. We found a number of incorrect genes and several new ones in the genome data in question. We also found a way to estimate the total number of orthologous genes in the bacterial world.

Bacteria↗

Sterol regulatory element binding protein (SREBP)-1 expression in brain is affected by age but not by hormones or metabolic changes.

Sterol regulatory element binding protein (SREBP)-1 is a membrane-bound transcription factor that regulates the expression of several genes involved in cellular fatty acid synthesis in the peripheral tissues, including liver. Although SREBP-1 is expressed in brain, little is known about its function. The aim of the present study was to clarify the characteristics of SREBP-1 mRNA expression in rat brain under various nutritional and hormonal conditions. In genetically obese (fa/fa) Zucker rats, expression of SREBP-1 mRNA was greater in liver than in hypothalamus or cerebrum compared to the lean littermates of these rats. Fasting for 45 h and refeeding for 3 h did not affect expression in brains of Wistar rats of SREBP-1 mRNA or the mRNAs of lipogenic enzymes that are targets of SREBP-1, i.e., fatty acid synthase (FAS) and acetyl-CoA carboxylase (ACC). Infusion of 2.0 mIU insulin or 3.0 microg leptin into the third cerebroventricle did not affect SREBP-1 mRNA expression in either hypothalamus or cerebrum. SREBP-1 mRNA expression in brains of transgenic mice that overexpressed leptin did not differ from that of wild-type mice. However, we observed a unique age-related alteration in SREBP-1 mRNA expression in brains of Sprague-Dawley rats. Specifically, SREBP-1 mRNA expression increased between 1 and 20 months of age, while there was no such change in the expression of FAS or ACC. This raises the possibility that increased SREBP-1 expression secondary to aging-related decline of polyunsaturated fatty acid (PUFA) might compensate for the reduction of FAS expression in brain. These findings suggest that the expression of SREBP-1 and downstream lipogenic enzymes in brain is probably not regulated by peripheral nutritional conditions or humoral factors. Aging-related changes in SREBP-1 mRNA expression may be involved in developmental changes in brain lipid metabolism.

Acetyl-CoA Carboxylase↗

Intrinsically disordered loops inserted into the structural domains of human proteins.

Much attention has been paid recently to proteins with partially or fully disordered structures, which are found to exist mostly in eukaryotes and are involved mainly in pivotal cellular processes such as transcriptional regulation, translation and cellular signal transduction. Long disordered sequences are sometimes inserted within the single structural domains of proteins, forming loops from the molecular surface. Such intrinsically disordered loops (IDLs) either are invisible in X-ray crystallography, or hamper protein crystallization itself due to great flexibility. Perhaps because of this, such long disordered sequences have not been characterized adequately. Here, we propose an informational method that stringently identifies IDLs in the structural domains of proteins using the amino acid sequence alone. A genome-wide survey of human proteins conducted with the method identified 50 IDL-containing proteins, several of which have experimentally determined 3D structures. Similar searches in other entirely sequenced organisms revealed that IDLs are prevalent in eukaryotes, while they are much less so in prokaryotes. As there is a statistically significant coincidence between the boundaries of IDLs and those of exons, we suggest that IDLs were produced mainly by exon addition in eukaryotes. IDLs are almost always located at the surface of proteins and are enriched with hydrophilic residues, and IDL-containing proteins tend to be intracellular. Some of the well-characterized proteins with IDLs illustrate that IDLs play pivotal roles in the switching of intracellular signaling or regulatory functions, suggesting that IDL insertion is an effective way to create functionally different domain variants.

Amino Acid Sequence↗

Investigation of protein functions through data-mining on integrated human transcriptome database, H-Invitational database (H-InvDB).

H-Invitational Database (H-InvDB; ) is a human transcriptome database, containing integrative annotation of 41,118 full-length cDNA clones originated from 21,037 loci. H-InvDB is a product of the H-Invitational project, an international collaboration to systematically and functionally validate human genes by analysis of a unique set of high quality full-length cDNA clones using automatic annotation and human curation under unified criteria. Here, 19,574 proteins encoded by these cDNAs were classified into 11,709 function-known and 7865 function-unknown hypothetical proteins by similarity with protein databases and motif prediction (InterProScan). The proportion of "hypothetical proteins" in H-InvDB was as high as 40.4%. In this study, we thus conducted data-mining in H-InvDB with the aim of assigning advanced functional annotations to those hypothetical proteins. First, by data-mining in the H-InvDB version of GTOP, we identified 337 SCOP domains within 7865 H-Inv hypothetical proteins. Second, by data-mining of predicted subcellular localization by SOSUI and TMHMM in H-InvDB, we found 1032 transmembrane proteins within H-Inv hypothetical proteins. These results clearly demonstrate that structural prediction is effective for functional annotation of proteins with unknown functions. All the data in H-InvDB are shown in two main views, the cDNA view and the Locus view, and five auxiliary databases with web-based viewers; DiseaseInfo Viewer, H-ANGEL, Clustering Viewer, G-integra and TOPO Viewer; the data also are provided as flat files and XML files. The data consists of descriptions of their gene structures, novel alternative splicing isoforms, functional RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs in relation with orphan diseases, gene expression profiling, and comparisons with mouse full-length cDNAs in the context of molecular evolution. This unique integrative platform for conducting in silico data-mining represents a substantial contribution to resources required for the exploration of human biology and pathology.

Amino Acid Sequence↗

Estimation of the number of authentic orphan genes in bacterial genomes.

Genome annotation produces a considerable number of putative proteins lacking sequence similarity to known proteins. These are referred to as "orphans." The proportion of orphan genes varies among genomes, and is independent of genome size. In the present study, we show that the proportion of orphan genes roughly correlates with the isolation index of organisms (IIO), an indicator introduced in the present study, which represents the degree of isolation of a given genome as measured by sequence similarity. However, there are outlier genomes with respect to the linear correlation, consisting of those genomes that may contain excess amounts of orphan genes. Comparisons of genome sequences among closely related strains revealed that some of the annotated genes are not conserved, suggesting that they are ORFs occurring by chance. Exclusion of these non-conserved ORFs within closely related genomes improved the correlation between the proportion of orphan genes and the IIO values. Assuming that the correlation holds in general, this relationship was used to estimate the number of "authentic" orphan genes in a genome. Using this definition of authentic orphan genes, the anomalies arising from over-assignments, e.g., the percentages of structural annotations, were corrected for 16 genomes, including those of five archaea.

Amino Acid Sequence↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Role of fatty acid composition in the development of metabolic disorders in sucrose-induced obese rats.

Fatty acids have been shown to be involved in the development of insulin resistance associated with obesity. We used sucrose loading in rats to analyze changes in fatty acid composition in the progression of obesity and the related metabolic disorder. Although rats fed a sucrose diet for 4 weeks had body weights similar to those of control animals, their visceral fat pads were significantly larger, and serum triglyceride levels were higher; however, neither plasma glucose nor insulin levels were significantly higher. After 20 weeks of sucrose loading, body weight and visceral and subcutaneous fat pads had increased significantly compared with those in control rats. Moreover, plasma glucose, insulin, and triglyceride levels were significantly higher. An analysis of individual fatty acid components in the blood and peripheral tissues demonstrated phase- and tissue-dependent changes. After 20 weeks of sucrose loading, palmitoleic acid (16:1 n-7) and oleic acid (18:1 n-9), the major components of monounsaturated fatty acid, showed a ubiquitous increase in plasma and all tissues analyzed. In contrast, linoleic acid (18:2 n-6) and arachidonic acid (20:4 n-6), the major components of polyunsaturated fatty acid in the n-6 family, decreased in plasma and all tissues analyzed. After 4 weeks of sucrose loading, these changes in fatty acid composition were observed only in the liver and plasma and not in fat and muscle. This led us to conclude that elevation of plasma glucose and insulin develop at the late phase of sucrose-induced obesity, when changes in fatty acid composition appear in fat and muscle. Furthermore, changes in fatty acid composition in liver seen after 4 weeks of sucrose loading, when increases in neither plasma glucose nor insulin were detected, suggest that liver may be the initial site of fatty acid imbalance and that aberrations in hepatic fatty acid composition may lead to fatty acid imbalances in other tissues.

Adipose Tissue↗

Unique amino acid composition of proteins in halophilic bacteria.

The amino acid compositions of proteins from halophilic archaea were compared with those from non-halophilic mesophiles and thermophiles, in terms of the protein surface and interior, on a genome-wide scale. As we previously reported for proteins from thermophiles, a biased amino acid composition also exists in halophiles, in which an abundance of acidic residues was found on the protein surface as compared to the interior. This general feature did not seem to depend on the individual protein structures, but was applicable to all proteins encoded within the entire genome. Unique protein surface compositions are common in both halophiles and thermophiles. Statistical tests have shown that significant surface compositional differences exist among halophiles, non-halophiles, and thermophiles, while the interior composition within each of the three types of organisms does not significantly differ. Although thermophilic proteins have an almost equal abundance of both acidic and basic residues, a large excess of acidic residues in halophilic proteins seems to be compensated by fewer basic residues. Aspartic acid, lysine, asparagine, alanine, and threonine significantly contributed to the compositional differences of halophiles from meso- and thermophiles. Among them, however, only aspartic acid deviated largely from the expected amount estimated from the dinucleotide composition of the genomic DNA sequence of the halophile, which has an extremely high G+C content (68%). Thus, the other residues with large deviations (Lys, Ala, etc.) from their non-halophilic frequencies could have arisen merely as "dragging effects" caused by the compositional shift of the DNA, which would have changed to increase principally the fraction of aspartic acid alone.

Amino Acids↗

Compositional changes in RNA, DNA and proteins for bacterial adaptation to higher and lower temperatures.

It is known that in thermophiles the G+C content of ribosomal RNA linearly correlates with growth temperature, while that of genomic DNA does not. Although the G+C contents (singlet) of the genomic DNAs of thermophiles and methophiles do not differ significantly, the dinucleotide (doublet) compositions of the two bacterial groups clearly do. The average amino acid compositions of proteins of the two groups are also distinct. Based on these facts, we here analyzed the DNA and protein compositions of various bacteria in terms of the optimal growth temperature (OGT). Regression analyses of the sequence data for thermophilic, mesophilic and psychrophilic bacteria revealed good linear relationships between OGT and the dinucleotide compositions of DNA, and between OGT and the amino acid compositions of proteins. Together with the above-mentioned linear relationship between ribosomal RNA and OGT, the DNA and protein compositions can be regarded as thermostability measures for RNA, DNA and proteins, covering a wide range of temperatures. Both the DNA and proteins of psychrophiles apparently exhibit characteristics diametrically opposite to those of thermophiles. The physicochemical parameters of dinucleotides suggested that supercoiling of DNA is relevant to its thermostability. Protein stability in thermophiles is realized primarily through global changes that increase charged residues (i.e., Glu, Arg, and Lys) on the molecular surface of all proteins. This kind of global change is attainable through a change in the amino acid composition coupled with alterations in the DNA base composition. The general strategies of thermophiles and psychrophiles for adaptation to higher and lower temperatures, respectively, that are suggested by the present study are discussed.

Adaptation, Physiological↗

A systematic investigation identifies a significant number of probable pseudogenes in the Escherichia coli genome.

Pseudogenes are open reading frames (ORFs) encoding dysfunctional proteins with high homology to known protein-coding genes. Although pseudogenes were reported to exist in the genomes of many eukaryotes and bacteria, no systematic search for pseudogenes in the Escherichia coli genome has been carried out. Genome comparisons of E. coli strains K-12 and O157 revealed that many protein-coding sequences have prematurely terminated orthologs encoding unstable proteins. To systematically screen for pseudogenes, we selected ORFs generated by premature termination of the orthologous protein-coding genes and subsequently excluded those possibly arising from sequence errors. Lastly we eliminated those with close homologs in this and other species, as these shortened ORFs may actually have functions. The process produced 95 and 101 pseudogene candidates in K-12 and O157, respectively. The assigned three-dimensional structures suggest that most of the encoded proteins cannot fold properly and thus are dysfunctional, indicating that they are probably pseudogenes. Therefore, the existence of a significant number of probable pseudogenes in E. coli is predicted, awaiting experimental verification. Most of them were found to be genes with paralogs or horizontally transferred genes or both. We suggest that pseudogenes constitute a small fraction of the genomes of free-living bacteria in general, reflecting the faster elimination than production of pseudogenes.

Amino Acid Sequence↗

Sequence-based approach for identification of cell wall proteins in Saccharomyces cerevisiae.

Open reading frames (ORFs) in the genome of Saccharomyces cerevisiae were screened for cell wall proteins and extracellular proteins, using an in silico sequence analysis combined with biochemical examination. The selection criteria used in the sequence analysis were the presence of a signal sequence for secretion and the absence of any targeting and retention signal to/in intracellular components. By using the PSORT II program, 163 ORFs/proteins were selected as potential extracellular proteins, including cell wall proteins. Of these, 51 ORFs/proteins of unknown localization and more than 120 amino acids in size were further studied on their cellular localization. A hemagglutinin (HA) epitope was inserted in the most C-terminus of each protein and the resulting HA-tagged protein was expressed under the authentic promoter in yeast cells. Out of the 51 constructs, 35 gave protein bands on Western blots. Examination of proteins in fractionated samples identified 11 extracellular proteins; six proteins that were weakly associated with the cell wall and five proteins that were relatively tightly associated with the cell wall.

Cell Wall↗

GTOP: a database of protein structures predicted from genome sequences.

Large-scale genome projects generate an unprecedented number of protein sequences, most of them are experimentally uncharacterized. Predicting the 3D structures of sequences provides important clues as to their functions. We constructed the Genomes TO Protein structures and functions (GTOP) database, containing protein fold predictions of a huge number of sequences. Predictions are mainly carried out with the homology search program PSI-BLAST, currently the most popular among high-sensitivity profile search methods. GTOP also includes the results of other analyses, e.g. homology and motif search, detection of transmembrane helices and repetitive sequences. We have completed analyzing the sequences of 41 organisms, with the number of proteins exceeding 120 000 in total. GTOP uses a graphical viewer to present the analytical results of each ORF in one page in a 'color-bar' format. The assigned 3D structures are presented by Chime plug-in or RasMol. The binding sites of ligands are also included, providing functional information. The GTOP server is available at http://spock.genes.nig.ac.jp/~genome/gtop.html.

Amino Acid Motifs↗