PubMed Health⌕ Search

PubMed · 8877525

Gene recognition in cyanobacterium genomic sequence data using the hidden Markov model.

Abstract

We have developed a hidden Markov model (HMM) to detect the protein coding regions within one megabase contiguous sequence data, registered in a database called GenBank in eight entries, of the genome of cyanobacterium, Synechocystis sp. strain PCC6803. Detection of the coding regions in the database entry was performed by using HMM whose parameters were determined by taking the statistics from the rests of the entries. This HMM has states modeling the di-codons and their frequencies within coding regions and those modeling its base contents in the intergenic regions. Results of the cross-validation showed that the HMM recognized 92.1% of coding regions assigned in sequence annotation. In addition, it suggested 94 potential new coding regions whose length are longer than 90 bases. The recognition accuracy calculated at the level of individual bases was 90.7% for the coding regions and 88.1% for the intergenic regions. This corresponds to a correlation coefficient for coding region recognition of 0.784. Comparison with its prediction accuracy with that by GeneMark showed that the HMM has the same level of prediction accuracy as GeneMark on average. Since we can extend the HMM to utilize information such as SD sequences, the prediction accuracy of the HMM will be enhanced. It was observed that correlation was positive between the prediction rate of the coding regions and the G + C content at the third position of the codon. This suggests the possibility that the prediction rate of coding regions in the cyanobacteria sequence can be enhanced by improving the present HMM into that reflects the classification of coding regions based on the G + C content.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

T Yada, M Hirosawa. 1996. Gene recognition in cyanobacterium genomic sequence data using the hidden Markov model.. https://pubmed.ncbi.nlm.nih.gov/8877525/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Dual solvent cavities and hydrogen-bond networks define the chromophore environment in a far-red/orange-sensing cyanobacteriochrome.

Cyanobacteriochromes (CBCRs) are bilin-binding photoreceptors that exhibit remarkable spectral diversity and mediate light-dependent signaling in cyanobacteria. Far-red/orange-sensing CBCRs (froCBCRs) have attracted interest because of their unusually red-shifted absorption properties, yet structural information for their illuminated states has been lacking. Here, we report the first high-resolution (1.8 Å) crystal structure of the orange-absorbing (Po) state of the froCBCR ToFrO from Tolypothrix sp. PCC 7910. The structure reveals a compact, cyclic bilin configuration and water-mediated hydrogen-bonding networks within two solvent-accessible cavities. Within the GAF domain, the D-ring remains nearly perpendicular to the planar A-to-C ring system through interactions involving a flexible loop region. Comparative analyses of cryogenic synchrotron and room-temperature X-ray free-electron laser (XFEL) structures, together with molecular dynamics (MD) simulations, revealed alternative Met636 conformations associated with dynamic water exchange through the solvent-accessible cavity. Site-directed mutagenesis of cavity-lining and water-interacting residues resulted in modest spectral shifts. By contrast, mutations of two Trp residues, participating in π-π stacking with the D-ring and likely imposing a steric constraint near the A-ring, resulted in substantial blue and red shifts in the dark and illuminated states, respectively. Together with the observed chromophore geometry, these findings indicate that the spectral properties of ToFrO are governed by chromophore conformation and its direct interaction with surrounding residues through hydrogen-bonding, electrostatic, and π-π interactions. These results further suggest that cavity-mediated solvent organization contributes to stabilizing the local structural environment surrounding the chromophore and adjacent protein backbone. Collectively, these findings elucidate the structural basis for photoconversion and spectral tuning in froCBCRs.

Cyanobacteria↗

Characterization of Dapalides D and E and Genomic Comparison of the Two Co-Occurring Dapalide-Producing Dapis spp.

Marine cyanobacteria are a rich source of diverse bioactive natural products, targeting proteins involved in many diseases. Here, we combined metagenomic analysis to enhance the structure elucidation process of two new cyclodepsipeptides named dapalides D (1) and E (2) from a collection of a cyanobacterial mat containing multiple Dapis species from Guam. Dapalides D/E are composed of 11 amino acids, including multiple identical units with different configurations. Enantioselective amino acid identification of the acid hydrolyzate established the identity of amino acids, including the configuration of α/β-stereogenic centers. Identification and analysis of the dapalides D/E biosynthetic gene cluster from a metagenome-assembled genome aided the elucidation of α-configuration and establishment of the order of individual building blocks, collectively revealing the total structure. Phylogenomic analysis indicates that the dapalides D/E producer belongs to Dapis sp. (Dapis sp. VPG23-80 MAG-2), which shares a 95.2% average nucleotide identity with Dapis sp. VPG23-80 MAG-1, the producer of dapalides A-C that cooccurs in the same assemblage. Dapalide D (1) showed moderate growth inhibitory activity against various cancer cell lines. This work expands the dapalide structure class and further highlights the use of combined chemical and metagenomic analyses for natural product structure elucidation.

Cyanobacteria↗

Gloeotrichia echinulata genomes from the United States are nontoxigenic and likely geosmin producers.

Six Gloeotrichia echinulata genomes derived from planktonic harmful algal blooms (HABs) with similar colonial morphology have been sequenced from lakes in the west and northeast regions of USA, four of them to completion. The c. 7 Mbp genomes exhibit a high level of conservation, with 98-99% pairwise genome-wide average nucleotide identity and high levels of synteny, representing a single species cluster. We observed strong conservation of gene clusters responsible for the synthesis of the secondary metabolites and bioactive peptides that are characteristic of HAB-forming cyanobacteria. All six G. echinulata genomes lack genes for the synthesis of classic cyanotoxins, including microcystin, but possess genes responsible for the synthesis of the taste and odor compound geosmin. Interestingly, the geoA geosmin synthase gene in three genomes is homologous to other cyanobacterial geoA genes, while the other three geoA genes are related to actinomyces geoA. Phylogenomic analysis places the G. echinulata genomes within a clade of benthic Nostocales, reflecting an ecological niche featuring extensive growth on the sediment surface before colonies disperse into the epilimnion for planktonic growth. We identify genes conserved in all six genomes that could represent physiological adaptations supporting active growth on sediments and pelagic recruitment independent of wind-driven mixing: phycoerythrin light harvesting complexes for optimal photosynthesis at depth; gliding motility to access patchy nutrient distributions; and gas vesicles with relatively small GvpC proteins that predict resistance to higher hydrostatic pressure. The strong genomic similarity across geographically distant populations suggests that G. echinulata in the United States is a tightly related non-toxigenic species group with predictable properties relevant to public health and drinking water management.

Cyanobacteria↗