PubMed Health⌕ Search

Biomedical subjects

Hideki Noguchi

Publications and source records attributed to Hideki Noguchi.

12 recordsLinked to original sources

MetaGene: prokaryotic gene finding from environmental genome shotgun sequences.

Exhaustive gene identification is a fundamental goal in all metagenomics projects. However, most metagenomic sequences are unassembled anonymous fragments, and conventional gene-finding methods cannot be applied. We have developed a prokaryotic gene-finding program, MetaGene, which utilizes di-codon frequencies estimated by the GC content of a given sequence with other various measures. MetaGene can predict a whole range of prokaryotic genes based on the anonymous genomic sequences of a few hundred bases, with a sensitivity of 95% and a specificity of 90% for artificial shotgun sequences (700 bp fragments from 12 species). MetaGene has two sets of codon frequency interpolations, one for bacteria and one for archaea, and automatically selects the proper set for a given sequence using the domain classification method we propose. The domain classification works properly, correctly assigning domain information to more than 90% of the artificial shotgun sequences. Applied to the Sargasso Sea dataset, MetaGene predicted almost all of the annotated genes and a notable number of novel genes. MetaGene can be applied to wide variety of metagenomic projects and expands the utility of metagenomics.

Computational Biology↗

A wide-range phylogenetic analysis of Zic proteins: implications for correlations between protein structure conservation and body plan complexity.

We compared Zic homologues from a wide range of animals. Striking conservation was found in the zinc finger domains, in which an exon-intron boundary has been kept in all bilateralians but not cnidarians, suggesting that all of the bilateralian Zic genes are derived from a single gene in a bilateralian ancestor. There were additional conserved amino acid sequences, ZOC and ZF-NC. Combined analysis of the zinc finger, ZOC, and ZF-NC revealed the presence of two classes of Zic, based on the degree of protein structure conservation. The "conserved" class includes Zic proteins from the Arthropoda, Mollusca, Annelida, Echinodermata, and Chordata (vertebrates and cephalochordates), whereas the "diverged" class contains those from the Platyhelminthes, Cnidaria, Nematoda, and Chordata (urochordates). The result indicates that the ancestral bilateralian Zic protein had already acquired an entire set of conserved domains, but that this was lost and diverged in the platyhelminthes, nematodes, and urochordates.

Amino Acid Sequence↗

Human chromosome 11 DNA sequence and analysis including novel gene identification.

Chromosome 11, although average in size, is one of the most gene- and disease-rich chromosomes in the human genome. Initial gene annotation indicates an average gene density of 11.6 genes per megabase, including 1,524 protein-coding genes, some of which were identified using novel methods, and 765 pseudogenes. One-quarter of the protein-coding genes shows overlap with other genes. Of the 856 olfactory receptor genes in the human genome, more than 40% are located in 28 single- and multi-gene clusters along this chromosome. Out of the 171 disorders currently attributed to the chromosome, 86 remain for which the underlying molecular basis is not yet known, including several mendelian traits, cancer and susceptibility loci. The high-quality data presented here--nearly 134.5 million base pairs representing 99.8% coverage of the euchromatic sequence--provide scientists with a solid foundation for understanding the genetic basis of these disorders and other biological phenomena.

Chromosomes, Human, Pair 11↗

Comparative analysis of chimpanzee and human Y chromosomes unveils complex evolutionary pathway.

The mammalian Y chromosome has unique characteristics compared with the autosomes or X chromosomes. Here we report the finished sequence of the chimpanzee Y chromosome (PTRY), including 271 kb of the Y-specific pseudoautosomal region 1 and 12.7 Mb of the male-specific region of the Y chromosome. Greater sequence divergence between the human Y chromosome (HSAY) and PTRY (1.78%) than between their respective whole genomes (1.23%) confirmed the accelerated evolutionary rate of the Y chromosome. Each of the 19 PTRY protein-coding genes analyzed had at least one nonsynonymous substitution, and 11 genes had higher nonsynonymous substitution rates than synonymous ones, suggesting relaxation of selective constraint, positive selection or both. We also identified lineage-specific changes, including deletion of a 200-kb fragment from the pericentromeric region of HSAY, expansion of young Alu families in HSAY and accumulation of young L1 elements and long terminal repeat retrotransposons in PTRY. Reconstruction of the common ancestral Y chromosome reflects the dynamic changes in our genomes in the 5-6 million years since speciation.

Animals↗

Molecular characterization of ENU mouse mutagenesis and archives.

The large-scale mouse mutagenesis with ENU has provided forward-genetic resources for functional genomics. The frozen sperm archive of ENU-mutagenized generation-1 (G1) mice could also provide a "mutant mouse library" that allows us to conduct reverse genetics in any particular target genes. We have archived frozen sperm as well as genomic DNA from 9224 G1 mice. By genome-wide screening of 63 target loci covering a sum of 197 Mbp of the mouse genome, a total of 148 ENU-induced mutations have been directly identified. The sites of mutations were primarily identified by temperature gradient capillary electrophoresis method followed by direct sequencing. The molecular characterization revealed that all the identified mutations were point mutations and mostly independent events except a few cases of redundant mutations. The base-substitution spectra in this study were different from those of the phenotype-based mutagenesis. The ENU-based gene-driven mutagenesis in the mouse now becomes feasible and practical.

Animals↗

DNA sequence and analysis of human chromosome 18.

Chromosome 18 appears to have the lowest gene density of any human chromosome and is one of only three chromosomes for which trisomic individuals survive to term. There are also a number of genetic disorders stemming from chromosome 18 trisomy and aneuploidy. Here we report the finished sequence and gene annotation of human chromosome 18, which will allow a better understanding of the normal and disease biology of this chromosome. Despite the low density of protein-coding genes on chromosome 18, we find that the proportion of non-protein-coding sequences evolutionarily conserved among mammals is close to the genome-wide average. Extending this analysis to the entire human genome, we find that the density of conserved non-protein-coding sequences is largely uncorrelated with gene density. This has important implications for the nature and roles of non-protein-coding sequence elements.

Aneuploidy↗

Contribution of Asian mouse subspecies Mus musculus molossinus to genomic constitution of strain C57BL/6J, as defined by BAC-end sequence-SNP analysis.

MSM/Ms is an inbred strain derived from the Japanese wild mouse, Mus musculus molossinus. It is believed that subspecies molossinus has contributed substantially to the genome constitution of common laboratory strains of mice, although the majority of their genome is derived from the west European M. m. domesticus. Information on the molossinus genome is thus essential not only for genetic studies involving molossinus but also for characterization of common laboratory strains. Here, we report the construction of an arrayed bacterial artificial chromosome (BAC) library from male MSM/Ms genomic DNA, covering approximately 1x genome equivalent. Both ends of 176,256 BAC clone inserts were sequenced, and 62,988 BAC-end sequence (BES) pairs were mapped onto the C57BL/6J genome (NCBI mouse Build 30), covering 2,228,164 kbp or 89% of the total genome. Taking advantage of the BES map data, we established a computer-based clone screening system. Comparison of the MSM/Ms and C57BL/6J sequences revealed 489,200 candidate single nucleotide polymorphisms (SNPs) in 51,137,941 bp sequenced. The overall nucleotide substitution rate was as high as 0.0096. The distribution of SNPs along the C57BL/6J genome was not uniform: The majority of the genome showed a high SNP rate, and only 5.2% of the genome showed an extremely low SNP rate (percentage identity = 0.9997); these sequences are likely derived from the molossinus genome.

Animals↗

S-phase checkpoint proteins Tof1 and Mrc1 form a stable replication-pausing complex.

The checkpoint regulatory mechanism has an important role in maintaining the integrity of the genome. This is particularly important in S phase of the cell cycle, when genomic DNA is most susceptible to various environmental hazards. When chemical agents damage DNA, activation of checkpoint signalling pathways results in a temporary cessation of DNA replication. A replication-pausing complex is believed to be created at the arrested forks to activate further checkpoint cascades, leading to repair of the damaged DNA. Thus, checkpoint factors are thought to act not only to arrest replication but also to maintain a stable replication complex at replication forks. However, the molecular mechanism coupling checkpoint regulation and replication arrest is unknown. Here we demonstrate that the checkpoint regulatory proteins Tof1 and Mrc1 interact directly with the DNA replication machinery in Saccharomyces cerevisiae. When hydroxyurea blocks chromosomal replication, this assembly forms a stable pausing structure that serves to anchor subsequent DNA repair events.

Bromodeoxyuridine↗

Detection of herpes simplex virus DNA by in situ hybridization technique and polymerase chain reaction in Papanicolaou-stained cervicovaginal smears.

Papanicolaou-stained cervicovaginal smears from six patients with herpes simplex virus (HSV) infection were destained and reprocessed by means of in situ hybridization (ISH) technique to demonstrate the presence of HSV DNA utilizing biotinylated probe. Positive results were obtained in all six cases with intense staining signal for the HSV DNA in the nuclei of cells having a ground-glass nuclear appearance as well as in multinucleated giant cells. Furthermore, a hybridization signal was also noted in smears that had been prepared as much as 3 yr previously. HSV type 2-specific antigen was confirmed in six destained smears by means of immunoperoxidase staining. Moreover, polymerase chain reaction (PCR) was also performed for four patients from Pap-destained cervicovaginal smears. Amplified HSV DNA was detected in all four cases as 106 basepair PCR products by polyacrylamide gel electrophoresis. The combined use of cytology and the ISH technique and PCR amplification was of great value for the rapid cytodiagnosis of genital infection of HSV.

Adult↗

Comparative genomic sequence analysis of the human chromosome 21 Down syndrome critical region.

Comprehensive knowledge of the gene content of human chromosome 21 (HSA21) is essential for understanding the etiology of Down syndrome (DS). Here we report the largest comparison of finished mouse and human sequence to date for a 1.35-Mb region of mouse chromosome 16 (MMU16) that corresponds to human chromosome 21q22.2. This includes a portion of the commonly described "DS critical region," thought to contain a gene or genes whose dosage imbalance contributes to a number of phenotypes associated with DS. We used comparative sequence analysis to construct a DNA feature map of this region that includes all known genes, plus 144 conserved sequences > or =100 bp long that show > or =80% identity between mouse and human but do not match known exons. Twenty of these have matches to expressed sequence tag and cDNA databases, indicating that they may be transcribed sequences from chromosome 21. Eight putative CpG islands are found at conserved positions. Models for two human genes, DSCR4 and DSCR8, are not supported by conserved sequence, and close examination indicates that low-level transcripts from these loci are unlikely to encode proteins. Gene prediction programs give different results when used to analyze the well-conserved regions between mouse and human sequences. Our findings have implications for evolution and for modeling the genetic basis of DS in mice.

Animals↗

Hidden Markov model-based prediction of antigenic peptides that interact with MHC class II molecules.

Elucidating the interaction between major histocompatibility complex (MHC) molecules and antigenic peptides is fundamental to better understanding of the processes involved in immune responses and for the development of innovative immunotherapies. In the present study, hidden Markov models (HMM) were combined with the successive state splitting (SSS) algorithm for optimization of the HMM structure, to predict peptide binders to the human MHC class II molecule HLA-DRB1*0101. The predictive performance of our model (S-HMM) was compared with fully connected HMM and artificial neural network (ANN) methods using the relative operating characteristic (ROC) analysis. The S-HMM predictions had values of ROC > or = 0.85 which was at least as good, or better than the comparison methods. In addition, S-HMM is trained on positive data only and does not require exhaustive data preprocessing, such as peptide alignment. Our results demonstrated that S-HMM combines the high accuracy of predictions with the simplicity of implementation and is therefore useful for analyzing MHC class II binding peptides. In particular the S-HMM may be trained using only positive data and, the preprocessing of training data, such as peptide alignment and the selection of binding cores, is not required in this method.

Journal Article↗

A novel index which precisely derives protein coding regions from cross-species genome alignments.

We introduce here a novel index which precisely derives protein coding regions from cross-species genome alignments. The index is deeply related to frame recovery observed in coding sequence alignments, that is, if insertions or deletions of nucleotides causes frame shifts in coding regions, other in-dels which recover the reading frames will be often observed in the vicinity. In contrast, such frame recoveries are not observed in other conserved regions. We prepared two gene models: a model which finds gene by using sequence similarity and intrinsic gene measures (basic model), and the other model which finds gene by using frame recovery index in addition to sequence similarity and intrinsic gene measures (frame recovery model). We evaluated the prediction accuracies of the two models, and our benchmark test revealed that frame recovery model significantly improved the prediction accuracy in comparison with basic model.

Abstracting and Indexing↗