PubMed HealthSearch

Biomedical subjects

M Hirosawa

Publications and source records attributed to M Hirosawa.

17 recordsLinked to original sources

Analysis of sequence patterns surrounding the translation initiation sites on Cyanobacterium genome using the hidden Markov model.

Sequence patterns surrounding the translation initiation sites of Cyanobacterium were precisely analyzed by the hidden Markov model (HMM) based on the actual translation initiation sites. In a previous study, 72 actual protein coding regions and their translation initiation sites on the genome of Synechocystis sp. strain PCC6803 were determined by Sazuka et al. using protein two-dimensional electrophoresis and microsequening. In this work, we extracted the sequence patterns surrounding translation initiation sites as HMM using the computer program YEBIS. The constructed HMM could recognize all but one translation initiation site. The HMM contains an AG-rich region (5.7 bp on average), as the Shine-Dalgarno sequence exclusively contains purines, upstream of the translation initiation site (-9.7 position on average) and a CT rich region (4.2 bp on average) just upstream from the translation initiation site. In addition, we found that the second amino acid (-4.5,6) could be classified into two types, one of which had C as their second codon while another of which has a nucleotide distribution relatively similar to the distribution among amino acids in the 72 proteins. This fact corresponds well to our earlier finding that when the second nucleotide of the second amino acid of a translated protein was C, an initial methionine was processed and that otherwise the methionine was intact with high frequency.

Base Sequence

Detection of novel molecules recognized by anti-placental lactogen antibody in rat amniotic fluid.

Amniotic fluid contains various bioactive substances including the placental PRL family. In the present study, it was elucidated that rat amniotic fluid contained immunoreactive proteins which had different molecular sizes and pI values from the authentic placental lactogens (PLs) in the rat, recognized by antipeptide antibody to the N-terminal peptide of rat PL-I. Immunoreactive PLs residing in the amniotic fluid were characterized further by two-dimensional sodium dodecyl sulfate gel electrophoresis (2DE), immunoblotting and anion-exchange chromatography. Amniotic fluid collected from rats on day 12 of pregnancy contained two PL-like molecules, tentatively called A1 (MW 75 kDa, pI 4.6) and A2 (MW 99-102 kDa, pI 5.3-5.4). A1 and A2 are specific to the amniotic fluid, because no such molecules were found in the serum or placental extracts. Immunoblot analysis of amniotic fluid revealed that A1 levels increased, whereas those of A2 decreased to an undetectable level up to day 16 of pregnancy. When the A1 concentrations from days 12 to 20 were monitored intensively, they increased from day 12 to 14, were maintained until day 18, and then decreased dramatically by day 20. The expression pattern for A1 was therefore completely different from those of authentic placental PRL family members found in serum and placental tissue, indicating that the A1 is distinct from them. Partial purification by anion-exchange chromatography and 2DE revealed that A1 consisted of 5 isoforms.

Amniotic Fluid

Detection of short protein coding regions within the cyanobacterium genome: application of the hidden Markov model.

The gene-finding programs developed so far have not paid much attention to the detection of short protein coding regions (CDSs). However, the detection of short CDSs is important for the study of photosynthesis. We utilized GeneHacker, a gene-finding program based on the hidden Markov model (HMM), to detect short CDSs (from 90 to 300 bases) in a 1.0 mega contiguous sequence of cyanobacterium Synechocystis sp. strain PCC6803 which carries a complete set of genes for oxygenic photosynthesis. GeneHacker differs from other gene-finding programs based on the HMM in that it utilizes di-codon statistics as well. GeneHacker successfully detected seven out of the eight short CDSs annotated in this sequence and was clearly superior to GeneMark in this range of length. GeneHacker detected 94 potentially new CDSs, 9 of which have counterparts in the genetic databases. Four of the nine CDSs were less than 150 bases and were photosynthesis-related genes. The results show the effectiveness of GeneHacker in detecting very short CDSs corresponding to genes.

Algorithms

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. II. Sequence determination of the entire genome and assignment of potential protein-coding regions.

The sequence determination of the entire genome of the Synechocystis sp. strain PCC6803 was completed. The total length of the genome finally confirmed was 3,573,470 bp, including the previously reported sequence of 1,003,450 bp from map position 64% to 92% of the genome. The entire sequence was assembled from the sequences of the physical map-based contigs of cosmid clones and of lambda clones and long PCR products which were used for gap-filling. The accuracy of the sequence was guaranteed by analysis of both strands of DNA through the entire genome. The authenticity of the assembled sequence was supported by restriction analysis of long PCR products, which were directly amplified from the genomic DNA using the assembled sequence data. To predict the potential protein-coding regions, analysis of open reading frames (ORFs), analysis by the GeneMark program and similarity search to databases were performed. As a result, a total of 3,168 potential protein genes were assigned on the genome, in which 145 (4.6%) were identical to reported genes and 1,257 (39.6%) and 340 (10.8%) showed similarity to reported and hypothetical genes, respectively. The remaining 1,426 (45.0%) had no apparent similarity to any genes in databases. Among the potential protein genes assigned, 128 were related to the genes participating in photosynthetic reactions. The sum of the sequences coding for potential protein genes occupies 87% of the genome length. By adding rRNA and tRNA genes, therefore, the genome has a very compact arrangement of protein- and RNA-coding regions. A notable feature on the gene organization of the genome was that 99 ORFs, which showed similarity to transposase genes and could be classified into 6 groups, were found spread all over the genome, and at least 26 of them appeared to remain intact. The result implies that rearrangement of the genome occurred frequently during and after establishment of this species.

Bacterial Proteins

Gene recognition in cyanobacterium genomic sequence data using the hidden Markov model.

We have developed a hidden Markov model (HMM) to detect the protein coding regions within one megabase contiguous sequence data, registered in a database called GenBank in eight entries, of the genome of cyanobacterium, Synechocystis sp. strain PCC6803. Detection of the coding regions in the database entry was performed by using HMM whose parameters were determined by taking the statistics from the rests of the entries. This HMM has states modeling the di-codons and their frequencies within coding regions and those modeling its base contents in the intergenic regions. Results of the cross-validation showed that the HMM recognized 92.1% of coding regions assigned in sequence annotation. In addition, it suggested 94 potential new coding regions whose length are longer than 90 bases. The recognition accuracy calculated at the level of individual bases was 90.7% for the coding regions and 88.1% for the intergenic regions. This corresponds to a correlation coefficient for coding region recognition of 0.784. Comparison with its prediction accuracy with that by GeneMark showed that the HMM has the same level of prediction accuracy as GeneMark on average. Since we can extend the HMM to utilize information such as SD sequences, the prediction accuracy of the HMM will be enhanced. It was observed that correlation was positive between the prediction rate of the coding regions and the G + C content at the third position of the codon. This suggests the possibility that the prediction rate of coding regions in the cyanobacteria sequence can be enhanced by improving the present HMM into that reflects the classification of coding regions based on the G + C content.

Cyanobacteria

Computer survey for likely genes in the one megabase contiguous genomic sequence data of Synechocystis sp. strain PCC6803.

Using the computer program GeneMark, the open reading frames (ORFs) previously assigned within the one megabase sequence data of the genome of the cyanobacterium, Synechocystis sp. strain PCC6803 (Kaneko et al., DNA Res. 2: 153-166, 1995), were re-examined. Matrices required by GeneMark for its statistical calculation were generated and modified by running a script termed GeneMark-Genesis that performed recursive application of GeneMark against the Synechocystis data and evaluated the probability scores for optimization. Based on the matrices thus generated, 752 of the 818 previously assigned ORFs (92%) were supported by GeneMark as likely coding sequences, of which 26 were predicted to start at more internal positions than previously assigned. In addition, 50 ORFs were newly identified as likely coding sequences, most of them being shorter than 300 bp. Thus, the procedure was proven to be very powerful to locate likely coding regions within the genomic sequence data of Synechocystis without having prior information concerning their similarity to the genes of other organisms. However, GeneMark did not predict 66 previously assigned ORFs as likely genes: 14 of them showed significant degrees of similarity to known genes and 10 others were found within IS-like elements. It seems that these genes, many of which appear to be exogenous origin, escaped detection by GeneMark as in the case of "class 3 (horizontally transferred) genes" of E. coli, which in turn suggests that genes of different phylogenetic origins might also be detected as such by modifying the matrices.

Base Sequence

Comprehensive study on iterative algorithms of multiple sequence alignment.

Multiple sequence alignment is an important problem in the biosciences. To date, most multiple alignment systems have employed a tree-based algorithm, which combines the results of two-way dynamic programming in a tree-like order of sequence similarity. The alignment quality is not, however, high enough when the sequence similarity is low. Once an error occurs in the alignment process, that error can never be corrected. Recently, an effective new class of algorithms has been developed. These algorithms iteratively apply dynamic programming to partially aligned sequences to improve their alignment quality. The iteration corrects any errors that may have occurred in the alignment process. Such an iterative strategy requires heuristic search methods to solve practical alignment problems. Incorporating such methods yields various iterative algorithms. This paper reports our comprehensive comparison of iterative algorithms. We proved that performance improves remarkably when using a tree-based iterative method, which iteratively refines an alignment whenever two subalignments are merged in a tree-based way. We propose a tree-dependent, restricted partitioning technique to efficiently reduce the execution time of iterative algorithms.

Algorithms

DMI-1, a new DNA methyltransferase inhibitor produced by Streptomyces sp. strain No. 560.

A new inhibitor of DNA methyltransferase named DMI-1 has been discovered in the culture filtrate of Streptomyces sp. strain No. 560. DMI-1 was purified by extraction with ethyl acetate followed by Diaion HP-20SS and silica gel column chromatography. The structure of DMI-1 was determined to be 8-methylpentadecanoic acid (C16H32O2). DMI-1 is a novel inhibitor of methyltransferase isolated from microorganisms and is structurally different from sinefungin and A9145C which are structural analogs of S-adenosylmethionine (methyl donor). DMI-1 was a strong inhibitor of N6-methyladenine-DNA methyltransferase (M. Eco RI, EC 2.1.1.72) in a noncompetitive manner and its inhibition depended on the pH and temperature in the assay media.

DNA Modification Methylases

A cDNA encoding a new member of the rat placental lactogen family, PL-I mosaic (PL-Im).

We have isolated a cDNA encoding a novel placental lactogen (PL), PL-I mosaic (PL-Im), by screening the cDNA library in lambda ZAP of the mid-pregnant (day 12) rat placenta. The cDNA comprised an open reading frame of 687 bp encoding 229 amino acids, in which there were two putative N-glycosylation sites. Northern blot analysis showed that PL-Im mRNA was expressed specifically during mid-pregnancy (days 10 and 12) in the rat placenta. The cDNA was highly homologous with those of other rat PL family members; in particular, the homology among PL-Im, PL-I (mid-pregnancy-specific) and PL-Iv (late-pregnancy-specific) was over 90%. Interestingly, the nucleotide sequence of PL-Im cDNA was a mosaic of PL-I and PL-Iv and it did not possess its own particular sequence. As the genomic pattern determined using Southern blot analysis of PL-Im was distinct from that of PL-I and PL-Iv, the encoded area of PL-Im appears to be independent of those of PL-I and PL-Iv on the gene. In the dendrogram of the rat PL family constructed on the basis of the nucleotide sequence homologies, PL-Im was located between PL-Iv and PL-I in the process of molecular evolution. Therefore, PL-Im has a unique cDNA structure and may be a principal factor in the molecular evolution of PLs.

Amino Acid Sequence

Effect of protein nutrition on the mRNA content of insulin-like growth factor-binding protein-1 in liver and kidney of rats.

Effect of quantity and nutritional quality of dietary proteins on the content of mRNA of insulin-like growth factor-binding protein-1 (IGFBP-1) was studied in rat liver and kidney. IGFBP-1 mRNA content per unit RNA increased in liver and kidney of rats fed on a protein-free diet and in those of fasted rats compared with that in the rats fed on a casein diet. When rats were given a gluten diet for 7 d, IGFBP-1 mRNA content in liver did not change significantly but that in kidney increased considerably compared with that in those organs of the rats fed on the casein diet. Because IGFBP-1 mRNA has been demonstrated both in liver parenchymal and non-parenchymal cells (Takenaka et al. 1991), the effect of the protein-free diet on these two types of cells has been studied. An increase in IGFBP-1 mRNA content under protein deprivation was observed in both liver parenchymal and non-parenchymal cells, suggesting that these two types of cells are regulated in a similar mode as far as IGFBP-1 mRNA content is concerned. The physiological and nutritional significance of the previously stated results on protein anabolism are discussed when considered together with our previous observations on the plasma concentrations of IGF-1 (Takahashi et al. 1990) and IGFBP (Umezawa et al. 1991) and insulin-like growth factor-1 mRNA content in liver (Miura et al. 1991).

Animals

MASCOT: multiple alignment system for protein sequences based on three-way dynamic programming.

A multiple alignment methodology that can produce high-quality alignment is extremely important for predicting the structure of unknown proteins. Nearly all the methodologies developed so far have employed two-way alignment only. Although these methods are fast, the alignments they produce lose reliability as the similarity of sequences reduces. We developed the MASCOT multiple alignment system. MASCOT can sustain the reliability of alignment even when the similarity of sequences is low. MASCOT achieves high-quality alignment by employing three-way alignment in addition to two-way alignment. The resultant alignments are refined by simulated annealing to higher quality. We also use a cluster analysis of sequences to produce highly reliable alignments.

Algorithms

Characterization of rat placental lactogen-alpha (PL-alpha) with an antipeptide antibody directed against rat PLs.

We have already reported that the serum of the mid-pregnant rat contains placental lactogen-alpha (PL-alpha), which has a higher molecular weight than PL-I when analysed by gel-filtration chromatography. In order to characterize PL-alpha, we prepared antipeptide antisera directed against the hydrophilic regions of authentic PLs and developed RIAs with them. With these RIAs, antiserum AI-86 reacted only with PL-I, whereas antiserum AII-86 reacted with both PL-alpha and PL-I, which indicates that the antibody recognition sites of these molecules are not identical and their core proteins differ distinctly. As the use of the RIAs with AI-86 and AII-86 in combination enabled PL-alpha and PL-I to be discriminated, the glycoresidues and hydrophobicity of PL-alpha were analyzed further. Analysis with hydrophobic chromatography revealed that PL-alpha was less hydrophobic than PL-I and, furthermore, PL-alpha, but not PL-I, bound specifically to a wheat germ agglutinin (WGA) lectin column. Therefore, in view of their different amino acid sequences and glycoresidues, PL-alpha and PL-I are distinct molecular entities.

Amino Acid Sequence