PubMed Health⌕ Search

Biomedical subjects

Shinsei Minoshima

Publications and source records attributed to Shinsei Minoshima.

At least 19 recordsLinked to original sources

DNA sequence and analysis of human chromosome 8.

The International Human Genome Sequencing Consortium (IHGSC) recently completed a sequence of the human genome. As part of this project, we have focused on chromosome 8. Although some chromosomes exhibit extreme characteristics in terms of length, gene content, repeat content and fraction segmentally duplicated, chromosome 8 is distinctly typical in character, being very close to the genome median in each of these aspects. This work describes a finished sequence and gene catalogue for the chromosome, which represents just over 5% of the euchromatic human genome. A unique feature of the chromosome is a vast region of approximately 15 megabases on distal 8p that appears to have a strikingly high mutation rate, which has accelerated in the hominids relative to other sequenced mammals. This fast-evolving region contains a number of genes related to innate immunity and the nervous system, including loci that appear to be under positive selection--these include the major defensin (DEF) gene cluster and MCPH1, a gene that may have contributed to the evolution of expanded brain size in the great apes. The data from chromosome 8 should allow a better understanding of both normal and disease biology and genome evolution.

Animals↗

Isolation and characterization of simple repeat sequences from the yellow fin sea bream Acanthopagrus latus (Sparidae).

We isolated DNA fragments containing various repetitive elements from the genome of a sea bream Acanthopagrus latus. Sequence analysis indicated that two fragments have particularly interesting features. Fragment AL87 contained a tetranucleotide repeat and a quasipalindromic sequence. Sequence comparison suggested that AL87 may be a part of a gene encoding a serine/threonine protein kinase, and that the quasipalindrome is situated at the junction of an intron and an exon. Moreover, the quasipalindrome is conserved in several other fishes, even though it has the potential to form a stem-loop structure at the splicing site. Fragment AL79 contained a minisatellite sequence made up of six 30-bp units in tandem. DNase I sensitivity assays and statistical analyses showed the repeat region to be flexible when subjected to bending stress. In addition, atomic force microscopic imaging of AL79 showed the presence of highly curved (kinked) segments flanking the repeat region. The structural features of these repetitive elements may be key factors facilitating the amplification of the repeats.

Animals↗

Identification and characterization of a novel gene family YPEL in a wide spectrum of eukaryotic species.

During comprehensive sequence analysis of human chromosome 22, we identified a novel gene family consisting of five members (YPEL1 through YPEL5) which has high homology with Drosophila yippee gene. We cloned and sequenced cDNAs for all five genes and determined their exon/intron organization. These YPEL genes showed high homology (43.8-96.6%) at amino acid sequence level among them. Mouse counterparts (Ypel1 through Ypel5) were also identified in the syntenic region of mouse chromosomes and their cDNAs were cloned and sequenced. Each of five pairs of human/mouse orthologs revealed extremely high homology. Thus, we named these genes as members of YPEL gene family. We searched YPEL family genes from the public databases, and found 100 genes from 68 species including animals, plants and fungi. Amino acid sequences of these 100 YPEL proteins were extremely similar and a consensus sequence of C-X(2)-C-X(19)-G-X(3)-L-X(5)-N-X(13)-G-X(8)-C-X(2)-C-X(4)-GWXY-X(10)-K-X(6)-E was established for all the YPEL family proteins without exception. Interestingly, the indirect immunofluorescent staining indicated that YPEL1-4 proteins are localized to the centrosome and nucleolus during interphase and at several dot-like structures around the mitotic apparatus during mitotic phase of COS-7 cells. YPEL5 protein is localized to the centrosome and nucleus during interphase and at the mitotic spindle during mitosis of the same cell line. Thus, the YPEL family proteins were found in essentially all the eukaryotes and hence they must play important roles in the maintenance of life. The subcellular localization of YPEL proteins in association with centrosome or mitotic spindle suggests a novel function involved in the cell division.

Amino Acid Sequence↗

Initial characterization of an uromodulin-like 1 gene on human chromosome 21q22.3.

We have isolated a novel gene, designated UMODL1, similar to uromodulin (UMOD)/Tamm-Horsfall glycoprotein, on human chromosome 21q22.3. Uromodulin like-1 (UMODL1) consists of 22 exons and spans approximately 80 kb in a direction from centromere to telomere. Two major transcripts produced by alternative splicing have been identified. These transcripts contain open reading frames of 4125 and 3741 bp encoding proteins of 1374 and 1246 amino acids, respectively. Expression of UMODL1 mRNA was detected only in 14 human tissues, e.g., kidney, testis, and fetal thymus at low level. Interestingly, two gene products (UMODL1L and UMODL1S) contain multiple domains including whey acidic protein, sea urchin sperm protein, enterokinase, and agrin, zona pellucida domain, and so on. Both proteins seemed to localize in cytoplasm, but UMODL1 is likely to be ubiquitinated and rapidly degraded in HEK293 cells. This gene may be a potent candidate for Down syndrome or bipolar affective disorder.

Alternative Splicing↗

Multiple gene organization of pufferfish Fugu rubripes tropomyosin isoforms and tissue distribution of their transcripts.

The Japanese pufferfish, torafugu (Fugu rubripes), has a haploid genome of about 400 Mb in size, which has been sequenced to approximately 90% coverage. Here we identified six Fugu tropomyosin (TPM) gene sequences by using the BLASTN program and the sequence of the white croaker TPM1 gene in our collection against the draft assembly of the Fugu genomic sequence database. TPM2, TPM3 and TPM4 genes were identified together with a set of two potentially duplicated genes of TPM1 (TPM1-1 and TPM1-2) as described in our previous report and TPM4 (TPM4-1 and TPM4-2) newly found in this study. The expression patterns of these Fugu TPM genes were determined by reverse transcription polymerase chain reaction (RT-PCR). A phylogenetic tree was constructed using the deduced amino acid sequences, which were encoded by the exons common to all vertebrate TPM genes. This indicated that the Fugu TPM1 and TPM4 genes had resulted from a gene duplication in the fish evolutionary lineage.

Alternative Splicing↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Comparative genomics of the keratin-associated protein (KAP) gene clusters in human, chimpanzee, and baboon.

We have previously identified a cluster of 16 genes that encode hair-specific proteins, called keratin-associated proteins (KAPs), located on human Chromosome (Chr) 21q22.3. Here, we have identified similar KAP gene clusters in two primates, chimpanzee and baboon. DNA sequence comparison revealed the common cluster structure consisting of 16 KAP genes for these three primates, but a significant difference was found in the baboon. Baboon possesses a new KAP gene not found in human and chimpanzee, whereas one KAP gene ( KRTAP18.12) that exists in human and chimpanzee was lost in baboon, making no change in the total number of KAP genes. Interestingly, the sequence for coding regions are highly variable among species owing to insertions and deletions, resulting in variation of gene size. On the contrary, the sequences for the 5' upstream region are highly conserved among species. These findings suggest that the ancestral KAP gene cluster was composed of 17 genes before the divergence of Old World monkeys (baboon) to the anthropoid (human and chimpanzee).

Amino Acid Sequence↗

A cluster of 21 keratin-associated protein genes within introns of another gene on human chromosome 21q22.3.

Recently, we identified multiple unique sequences in the 21q22.3 region and predicted them to be a cluster of genes encoding hair-specific keratin-associated proteins (KAPs). Detailed computer-aided analysis of these clustered genes revealed that the cluster spans over 165 kb and consists of 21 KAP-related sequences including 16 putative genes and 5 pseudogenes. These were further divided into two subfamilies, KRTAP12 (KRTAP12.1-12.4 and KRTAP12.5P) and KRTAP18 (KRTAP18.1-18.12 and KRTAP18.13P-18.16P). All 16 putative genes possess several intragenic repeat sequences and apparently belong to the high-sulfur KAP gene family (16-30% cysteine content) known for nonhuman mammalian species. Transcripts were detected by RT-PCR analysis for all 16 putative KAP genes and their expression was restricted to hair root cells (radix pili cells) and not found in 28 other tissues, including skin. All 16 KAP genes produced unspliced transcripts, indicating their nature to be that of active intronless genes. Interestingly, all these KAP-related genes are located within introns of the recently identified gene TSPEAR (approved gene symbol C21orf29), 214 kb in size. Surprisingly, the transcriptional direction of 8 of the 16 active genes is the same as that of C21orf29/TSPEAR. This finding suggests a novel transcription mechanism in which C21orf29/TSPEAR gene transcription passes over the multiple transcriptional termination sites of the KAP genes.

Amino Acid Sequence↗

Role of TBX1 in human del22q11.2 syndrome.

BACKGROUND: Del22q11.2 syndrome is the most frequent known chromosomal microdeletion syndrome, with an incidence of 1 in 4000-5000 livebirths. It is characterised by a 3-Mb deletion on chromosome 22q11.2, cardiac abnormalities, T-cell deficits, cleft palate facial anomalies, and hypocalcaemia. At least 30 genes have been mapped to the deleted region. However, the association of these genes with the cause of this syndrome is not clearly understood. METHODS: To test for the chromosomal deletion at 22q11.2, we did fluorescence in-situ hybridisation analysis with ten probes on 22q11.2 in 235 unrelated patients with clinically diagnosed del22q11.2 syndrome. To investigate mutations in the coding sequence of TBX1, we also did genetic analysis in 13 patients from ten families who have the 22q11.2 syndrome phenotype but no detectable deletion of 22q11.2. FINDINGS: 96% (225 of 235) of patients had a defined 1.5-3-Mb deletion at 22q11.2. We identified three mutations of TBX1 in two unrelated patients without the 22q11.2 deletion-one with sporadic conotruncal anomaly face syndrome/velocardiofacial syndrome and one with sporadic DiGeorge's syndrome-and in three patients from a family with conotruncal anomaly face syndrome/velocardiofacial syndrome. We did not record these three mutations in 555 healthy controls (1110 chromosomes; p<0.0001). INTERPRETATION: Our results suggest that the TBX1 mutation is responsible for five major phenotypes in del22q11.2 syndrome. Therefore, we conclude that TBX1 is a major genetic determinant of the del22q11.2 syndrome.

Abnormalities, Multiple↗

A novel giant gene CSMD3 encoding a protein with CUB and sushi multiple domains: a candidate gene for benign adult familial myoclonic epilepsy on human chromosome 8q23.3-q24.1.

We identified a novel giant gene encoding a transmembrane protein with CUB and sushi multiple domains on the human chromosome 8q23.3-q24.1 in which benign adult familial myoclonic epilepsy type 1 (BAFME1/FAME, OMIM:601068) has been mapped. This giant gene consists of 73 exons and spans over 1.2Mb on the genomic DNA region. It showed significant homology to two genes, CSMD1 gene on 8p23 and CSMD2 gene on 1p34, at reduced amino acid sequence level and hence we designated as CSMD3. The CSMD3 gene was expressed mainly in adult and fetal brains. We performed mutation analysis on the CSMD3 gene for seven patients with BAFME1/FAME, but no mutation was found in the coding sequence of the CSMD3 gene. Comparative genomic analysis revealed a conserved family of CSMD genes in the mouse and fugu genomes. Possible functions of the CSMD gene family are discussed.

Alleles↗

Frequent translocations occur between low copy repeats on chromosome 22q11.2 (LCR22s) and telomeric bands of partner chromosomes.

The chromosome 22q11.2 region is susceptible to rearrangements, mediated by low copy repeats (LCR22s). Deletions and duplications are mediated by homologous recombination events between LCR22s. The recurrent balanced constitutional translocation t(11;22)(q23;q11) breakpoint occurs in an LCR22 and is mediated by double strand breaks in AT-rich palindromes on both chromosomes 11 and 22. Recently, two cases of a t(17;22)(q11;q11) were reported, mediated by a similar mechanism (21). Except for these constitutional translocations, the molecular basis for non-recurrent, reciprocal 22q11.2 translocations is not known. To determine whether there are specific mechanisms that could mediate translocations, we analyzed cell lines derived from 14 different individuals by genotyping and FISH mapping. Somatic cell hybrid analysis was carried out for four cell lines. In five cell lines, the translocation breakpoints occurred in the same LCR22 as for the t(11;22) translocation, suggesting that similar molecular mechanisms are responsible. An additional three occurred in other LCR22s, and six were in non-LCR22 regions, mostly in the proximal half of the 22q11.2 region. The translocation breakpoints on the partner chromosomes were all located in the telomeric bands, proximal to the most telomeric unique sequence probe, in eight cell lines and distal to those loci in six. Therefore, several of the breakpoints were found to occur in the vicinity of highly dynamic regions of the genome, 22q11.2 and telomeric bands. We hypothesize that these regions are more susceptible to breakage and repair, resulting in translocations.

Base Sequence↗

Molecular cloning and expression analysis of a novel gene DGCR8 located in the DiGeorge syndrome chromosomal region.

We have identified and cloned a novel gene (DGCR8) from the human chromosome 22q11.2. This gene is located in the DiGeorge syndrome chromosomal region (DGCR). It consists of 14 exons spanning over 35kb and produces transcripts with ORF of 2322bp, encoding a protein of 773 amino acids. We also isolated a mouse ortholog Dgcr8 and found it has 95.3% identity with human DGCR8 at the amino acid sequence level. Northern blot analysis of human and mouse tissues from adult and fetus showed rather ubiquitous expression. However, the in situ hybridization of mouse embryos revealed that mouse Dgcr8 transcripts are localized in neuroepithelium of primary brain, limb bud, vessels, thymus, and around the palate during the developmental stages of embryos. The expression profile of Dgcr8 in developing mouse embryos is consistent with the clinical phenotypes including congenital heart defects and palate clefts associated with DiGeorge syndrome (DGS)/conotruncal anomaly face syndrome (CAFS)/velocardiofacial syndrome (VCFS), which are caused by monoallelic microdeletion of chromosome 22q11.2.

Amino Acid Motifs↗

Interarm interaction of DNA cruciform forming at a short inverted repeat sequence.

A novel interarm interaction of DNA cruciform forming at inverted repeat sequence was characterized using an S1 nuclease digestion, permanganate oxidation, and microscopic imaging. An inverted repeat consisting of 17 bp complementary sequences was isolated from the bluegill sunfish Lepomis macrochirus (Perciformes) and subcloned into the pUC19 plasmid, after which the supercoiled recombinant plasmid was subjected to enzymatic and chemical modification. In high salt conditions (200 mM NaCl, or 100-200 mM KCl), S1 nuclease cut supercoiled DNA at the center of palindromic symmetry, suggesting the formation of DNA cruciform. On the other hand, S1 nuclease in the presence of 150 mM NaCl or less cleaved mainly the 3'-half of the repeat, thereby forming an unusual structure in which the 3'-half of the inverted repeat, but not the 5'-half, was retained as an unpaired strand. Permanganate oxidation profiles also supported the presence of single-stranded part in the 3'-half of the inverted repeat in addition to the center of the symmetry. Both electron microscopy and atomic force microscopy have detected a thick protrusion on the supercoiled DNA harboring the inverted repeat. We hypothesize that the cruciform hairpins at conditions favoring triplex formation adopt a parallel side-by-side orientation of the arms allowing the interaction between them supposedly stabilized by hydrogen bonding of base triads.

Animals↗

Identification of eight members of the Argonaute family in the human genome.

A number of genes have been identified as members of the Argonaute family in various nonhuman organisms and these genes are considered to play important roles in the development and maintenance of germ-line stem cells. In this study, we identified the human Argonaute family, consisting of eight members. Proteins to be produced from these family members retain a common architecture with the PAZ motif in the middle and Piwi motif in the C-terminal region. Based on the sequence comparison, eight members of the Argonaute family were classified into two subfamilies: the PIWI subfamily (PIWIL1/HIWI, PIWIL2/HILI, PIWIL3, and PIWIL4/HIWI2) and the eIF2C/AGO subfamily (EIF2C1/hAGO1, EIF2C2/hAGO2, EIF2C3/hAGO3, and EIF2C4/hAGO4). PCR analysis using human multitissue cDNA panels indicated that all four members of the PIWI subfamily are expressed mainly in the testis, whereas all four members of the eIF2C/AGO subfamily are expressed in a variety of adult tissues. Immunoprecipitation and affinity binding experiments using human HEK293 cells cotransfected with cDNAs for FLAG-tagged DICER, a member of the ribonuclease III family, and the His-tagged members of the Argonaute family suggested that the proteins from members of both subfamilies are associated with DICER. We postulate that at least some members of the human Argonaute family may be involved in the development and maintenance of stem cells through the RNA-mediated gene-quelling mechanisms associated with DICER.

Amino Acid Motifs↗

Identification of novel tropomyosin 1 genes of pufferfish (Fugu rubripes) on genomic sequences and tissue distribution of their transcripts.

Fugu genome database enabled us to identify two novel tropomyosin 1 (TPM1) genes through in silico data mining and isolation of their corresponding cDNAs in vivo. The duplicate TPM1 genes in Japanese pufferfish Fugu rubripes suggest that additional an ancient segmental duplication or whole genome duplication occurred in fish lineage, which, like many other reported Fugu genes, showed reduction in genomic size in comparison with their human homologue. Computer analysis predicted that the coiled-coil probabilities, that were thought to be the most major function of TPM, were the same between the two TPM1 isoforms. We confirmed that the tissue expression profiles of the two TPM1 genes differed from each other, which implied that changes in expression pattern could fix duplicated TPM1 genes although the two TPM1 isoforms appear to have similar function.

Animals↗