PubMed HealthSearch

Biomedical subjects

G Lennon

Publications and source records attributed to G Lennon.

At least 19 recordsLinked to original sources

An encyclopedia of mouse genes.

The laboratory mouse is the premier model system for studies of mammalian development due to the powerful classical genetic analysis possible (see also the Jackson Laboratory web site, http://www.jax.org/) and the ever-expanding collection of molecular tools. To enhance the utility of the mouse system, we initiated a program to generate a large database of expressed sequence tags (ESTs) that can provide rapid access to genes. Of particular significance was the possibility that cDNA libraries could be prepared from very early stages of development, a situation unrealized in human EST projects. We report here the development of a comprehensive database of ESTs for the mouse. The project, initiated in March 1996, has focused on 5' end sequences from directionally cloned, oligo-dT primed cDNA libraries. As of 23 October 1998, 352,040 sequences had been generated, annotated and deposited in dbEST, where they comprised 93% of the total ESTs available for mouse. EST data are versatile and have been applied to gene identification, comparative sequence analysis, comparative gene mapping and candidate disease gene identification, genome sequence annotation, microarray development and the development of gene-based map resources.

Animals

Human endopeptidase 24.15 (THOP1) is localized on chromosome 19p13.3 and is excluded from the linkage region for late-onset Alzheimer disease.

The mapping position of human endopeptidase 24.15 (THOP1) has previously been reported to be within the linkage region for the late-onset Alzheimer disease AD2 locus on chromosome 19q13.3. After localizing THOP1 to the high-resolution cosmid contig map of human chromosome 19, we found that the previous report was incorrect. Results of the hybridization and FISH mapping of positive clones indicated localization of THOP1 to chromosome 19p13.3 and not 19q13. 3. This localization is a correction of wrong chromosomal delegation and excludes THOP1 from the region that shows evidence of linkage to late-onset familial Alzheimer disease.

Alzheimer Disease

High-resolution mapping of ribosomal protein genes to human chromosome 19.

In a systematic effort for mapping of all the human ribosomal protein (rp) genes, we have found that an unusually large number (12) of rp genes are present on chromosome 19 and subsequently determined their locations on the chromosome by a radiation-hybrid procedure. For this, we isolated cosmid clones corresponding to each gene and placed nine of them on a metric physical map of chromosome 19. Although most genes are scattered over the chromosome, we found three genes are clustered in a 0.6-Mb region at 19q13.3 and two of them, RPL13A and RPS11, within a single cosmid only 4.3 kb apart. To explore a possible relationship between rp gene defects and human disease, we compared map positions of the rp genes and disease loci on chromosome 19, which led us to find RPS9 gene in the same interval as the gene for retinitis pigmentosa 11. The disease locus has previously been mapped to the 6-cM interval at 19q13.4 between markers D19S572 and D19S926, which corresponds to less than 2-Mb region on the metric physical map. We mapped RPS9 about 800 kb distal to D19S572.

Base Sequence

Localization and genomic structure of human deoxyhypusine synthase gene on chromosome 19p13.2-distal 19p13.1.

The amino acid hypusine is formed post-translationally in a single cellular protein, the eukaryotic translation initiation factor 5A, by two enzymes, namely deoxyhypusine synthase and deoxyhypusine hydroxylase. Hypusine is found in all eukaryotes and in some archaebacteria, but not in eubacteria. The deoxyhypusine synthase cDNA was cloned and mapped by fluorescence in situ hybridization on chromosome 19p13.11-p13.12. Rare cDNAs containing internal deletions were also found. We localized the deoxyhypusine synthase gene on a high resolution cosmid/BAC contig map of chromosome 19 to a region in 19p13.2-distal 19p13.1 between MANB and JUNB. Analysis of the genomic exon/intron structure of the gene coding region showed that it consists of nine exons and spans a length of 6.6kb. From observation of the genomic structure, it seems likely that the internally deleted forms of mature RNA are the result of alternative splicing, rather than of artifacts.

Alternative Splicing

CAG/CTG and CGG/GCC repeats in human brain reference cDNAs: outcome in searching for new dynamic mutations.

CAG and CGG expansion is associated with 10 inherited neurological diseases and is thought to be involved in other human genetic diseases. To identify new candidate genes, we have undertaken a large-scale screening project for CAG/CTG ([CAG]n) and CGG/GCC ([CGG]n) repeats in human brain reference cDNAs. Here, we present the final classification for 597 cDNAs selected by CAG and CGG hybridization from two libraries (100,128 clones) and the updated characterization of [CAG]n- and [CGG]n-positive cDNAs (repeat polymorphism and cDNA localization). We have selected 124 CAG and 83 CGG hybridization-positive clones representing new genes, from which 49 CAG and 7 CGG repeats could be identified. New [CAG]n and [CGG]n with more than seven to nine units were rare (1/2000), and perfect [CAG]n 9 were more likely polymorphic. Overall, highly polymorphic to monomorphic new [CAG]n > 9 and [CGG]n > 7 were characterized. The comparison of our data with other [CAG]n and [CGG]n resources suggests that the screening of reference cDNAs leads to unique sources of new [CAG]n and [CGG]n and will enhance the study of enlarged triplet repeats in human genetic diseases.

Brain

Cloning, sequencing, gene organization, and localization of the human ribosomal protein RPL23A gene.

The intron-containing gene for human ribosomal protein RPL23A has been cloned, sequenced, and localized. The gene is approximately 4.0 kb in length and contains five exons and four introns. All splice sites exactly match the AG/GT consensus rule. The transcript is about 0.6 kb and is detected in all tissues examined. In adult tissues, the RPL23A transcript is dramatically more abundant in pancreas, skeletal muscle, and heart, while much less abundant in kidney, brain, placenta, lung, and liver. A full-length cDNA clone of 576 nt was identified, and the nucleotide sequence was found to match the exon sequence precisely. The open reading frame encodes a polypeptide of 156 amino acids, which is absolutely conserved with the rat RPL23A protein. In the 5' flanking region of the gene, a canonical TATA sequence and a defined CAAT box were found for the first time in a mammalian ribosomal protein gene. The intron-containing RPL23A gene was mapped to cytogenetic band 17q11 by fluorescence in situ hybridization.

Amino Acid Sequence

Isolation of chromosome 18-specific brain transcripts as positional candidates for bipolar disorder.

Several studies have proposed the existence of susceptibility loci for bipolar disorder on chromosome 18. To identify possible candidate genes for this disease, we isolated brain-expressed transcripts by direct cDNA selection on chromosome 18-specific biotinylated cosmid clones. Longer cognate cDNA clones of the selected cDNAs were isolated from a normalized infant brain cDNA library. Physical mapping by PCR on a panel of somatic cell hybrids was conducted by the use of primers derived from partial sequences on either the 5' or 3' ends of the clones. In our initial analysis, 48 cDNA clones were found to be chromosome 18-specific, mapping to different subchromosomal regions. Sequence redundancy among these clones yielded 30 unique transcripts, five of which were represented in previously known genes. Further sequencing of the remaining 25 unique cDNA clones confirmed the absence of significant homology to known genes, indicating that these transcripts represented novel genes. Mapping with the use of a radiation hybrid panel positioned the brain cDNAs to within = 100 to 1100 kb from reference sequence tag sites (STSs) and assembled them into six high resolution linkage groups. The majority of the transcripts were found to cluster to discrete locations on 18p and 18q, previously hypothesized as susceptibility regions for bipolar disorder, identifying them as positional candidate genes.

Bipolar Disorder

Mutations in the Cacnl1a4 calcium channel gene are associated with seizures, cerebellar degeneration, and ataxia in tottering and leaner mutant mice.

Tottering and leaner, two mutations of the mouse tottering locus, have been studied extensively as models for human epilepsy. Here we describe the isolation, mapping, and expression analysis of Cacnl1a4, a gene encoding the alpha subunit of a proposed P-type calcium channel, and also report the physical mapping and expression patterns of the orthologous human gene. DNA sequencing and gene expression data demonstrate that Cacnl1a4 mutations are the primary cause of seizures and ataxia in tottering and leaner mutant mice, and suggest that tottering locus mutations and human diseases, episodic ataxia 2 and familial hemiplegic migraine, represent mutations in mouse and human versions of the same channel-encoding gene.

Amino Acid Sequence

Crx, a novel Otx-like paired-homeodomain protein, binds to and transactivates photoreceptor cell-specific genes.

The otd/Otx gene family encodes paired-like homeodomain proteins that are involved in the regulation of anterior head structure and sensory organ development. Using the yeast one-hybrid screen with a bait containing the Ret 4 site from the bovine rhodopsin promoter, we have cloned a new member of the family, Crx (Cone rod homeobox). Crx encodes a 299 amino acid residue protein with a paired-like homeodomain near its N terminus. In the adult, it is expressed predominantly in photoreceptors and pinealocytes. In the developing mouse retina, it is expressed by embryonic day 12.5 (E12.5). Recombinant Crx binds in vitro not only to the Ret 4 site but also to the Ret 1 and BAT-1 sites. In transient transfection studies, Crx transactivates rhodopsin promoter-reporter constructs. Its activity is synergistic with that of Nrl. Crx also binds to and transactivates the genes for several other photoreceptor cell-specific proteins (interphotoreceptor retinoid-binding protein, beta-phosphodiesterase, and arrestin). Human Crx maps to chromosome 19q13.3, the site of a cone rod dystrophy (CORDII). These studies implicate Crx as a potentially important regulator of photoreceptor cell development and gene expression and also identify it as a candidate gene for CORDII and other retinal diseases.

Amino Acid Sequence

Large-scale concatenation cDNA sequencing.

A total of 100 kb of DNA derived from 69 individual human brain cDNA clones of 0.7-2.0 kb were sequenced by concatenated cDNA sequencing (CCS), whereby multiple individual DNA fragments are sequenced simultaneously in a single shotgun library. The method yielded accurate sequences and a similar efficiency compared with other shotgun libraries constructed from single DNA fragments (> 20 kb). Computer analyses were carried out on 65 cDNA clone sequences and their corresponding end sequences to examine both nucleic acid and amino acid sequence similarities in the databases. Thirty-seven clones revealed no DNA database matches, 12 clones generated exact matches (> or = 98% identity), and 16 clones generated nonexact matches (57%-97% identity) to either known human or other species genes. Of those 28 matched clones, 8 had corresponding end sequences that failed to identify similarities. In a protein similarity search, 27 clone sequences displayed significant matches, whereas only 20 of the end sequences had matches to known protein sequences. Our data indicate that full-length cDNA insert sequences provide significantly more nucleic acid and protein sequence similarity matches than expressed sequence tags (ESTs) for database searching.

DNA Transposable Elements

Isolation, sequencing, and mapping of the human homologue of the yeast transcription factor, SPT5.

We isolated the human homologue, SUPT5H, of the yeast transcription factor, SPT5. The human homologue is 1088 aa long compared to 1063 aa for the yeast gene. SUPT5H maps to 19q13, near the ryanodine receptor. Like its family member, SUPT6H, and like yeast SPT5, SUPT5H has a very acidic 5' domain. Like its family member, SUPT6H, but unlike yeast SPT5 or SPT6, SUPT5H has seven MAP kinase sites at its 5' end. In addition, SUPT5H lacks the novel 6-amino-acid repeat (consensus is S-T/A-W-G-G-A/Q) at the C-terminus of yeast SPT5. This argues that while there is functional similarity between SPT5 and SUPT5H, the molecules differ in the signals to which they respond.

Amino Acid Sequence

Survey of CAG/CTG repeats in human cDNAs representing new genes: candidates for inherited neurological disorders.

Expansion of polymorphic CAG and CTG repeats in transcripts is the cause of six inherited neurodegenerative or neuromuscular diseases and may be involved in several other genetic disorders of the central nervous system. To identify new candidate genes, we have undertaken a large-scale screening project for CAG and CTG repeats in human reference cDNAs. We screened 100 128 brain cDNAs by hybridization. We also scanned GenBank expressed sequence tags for the presence of long CAG/CTG repeats in the extremities of cDNAs from several human tissues. Of the selected clones, 286 were found to represent new genes, and 72 have thus far been shown to contain CAG/CTG repeats. Our data indicate that CAG/CTG repeated 10 or more times are more likely to be polymorphic, and that new 3'-directed cDNAs with such repeats are very rare (1/2862). Nine new cDNAs containing polymorphic (observed heterozygote frequency: 0.05-0.90) CAG/CTG repeats have been currently identified in cDNAs. All of the cDNAs have been assigned to chromosomes, and six of them could be mapped with YACs to 1q32-q41, 3p14, 4q28, 3p21 and 12q13.3, 13q13.1-q13.2, and 19q13.43. Three of these clones are highly polymorphic and represent the most likely candidate genes for inherited neurodegenerative diseases and, perhaps, neuropsychiatric disorders of multifactorial origin.

Brain

Normalization and subtraction: two approaches to facilitate gene discovery.

Large-scale sequencing of cDNAs randomly picked from libraries has proven to be a very powerful approach to discover (putatively) expressed sequences that, in turn, once mapped, may greatly expedite the process involved in the identification and cloning of human disease genes. However, the integrity of the data and the pace at which novel sequences can be identified depends to a great extent on the cDNA libraries that are used. Because altogether, in a typical cell, the mRNAs of the prevalent and intermediate frequency classes comprise as much as 50-65% of the total mRNA mass, but represent no more than 1000-2000 different mRNAs, redundant identification of mRNAs of these two frequency classes is destined to become overwhelming relatively early in any such random gene discovery programs, thus seriously compromising their cost-effectiveness. With the goal of facilitating such efforts, previously we developed a method to construct directionally cloned normalized cDNA libraries and applied it to generate infant brain (INIB) and fetal liver/spleen (INFLS) libraries, from which a total of 45,192 and 86,088 expressed sequence tags, respectively, have been derived. While improving the representation of the longest cDNAs in our libraries, we developed three additional methods to normalize cDNA libraries and generated over 35 libraries, most of which have been contributed to our integrated Molecular Analysis of Genomes and Their Expression (IMAGE) Consortium and thus distributed widely and used for sequencing and mapping. In an attempt to facilitate the process of gene discovery further, we have also developed a subtractive hybridization approach designed specifically to eliminate (or reduce significantly the representation of) large pools of arrayed and (mostly) sequenced clones from normalized libraries yet to be (or just partly) surveyed. Here we present a detailed description and a comparative analysis of four methods that we developed and used to generate normalize cDNA libraries from human (15), mouse (3), rat (2), as well as the parasite Schistosoma mansoni (1). In addition, we describe the construction and preliminary characterization of a subtracted liver/spleen library (INFLS-SI) that resulted from the elimination (or reduction of representation) of -5000 INFLS-IMAGE clones from the INFLS library.

Adult

Generation and analysis of 280,000 human expressed sequence tags.

We report the generation of 319,311 single-pass sequencing reactions (known as expressed sequence tags, or ESTs) obtained from the 5' and 3' ends of 194,031 human cDNA clones. Our goal has been to obtain tag sequences from many different genes and to deposit these in the publicly accessible Data Base for Expressed Sequence Tags. Highly efficient automatic screening of the data allows deposition of the annotated sequences without delay. Sequences have been generated from 26 oligo(dT) primed directionally cloned libraries, of which 18 were normalized. The libraries were constructed using mRNA isolated from 17 different tissues representing three developmental states. Comparisons of a subset of our data with nonredundant human mRNA and protein data bases show that the ESTs represent many known sequences and contain many that are novel. Analysis of protein families using Hidden Markov Models confirms this observation and supports the contention that although normalization reduces significantly the relative abundance of redundant cDNA clones, it does not result in the complete removal of members of gene families.

Adult

A continuous high-resolution physical map spanning 17 megabases of the q12, q13.1, and q13.2 cytogenetic bands of human chromosome 19.

We report the construction of a high-resolution physical map of a 17-Mb region that encompasses the entire q12, q13.1, and q13.2 bands of human chromosome 19. The continuous map extends from a region approximately 400 kb centromeric of the D19S7 marker to the excision repair cross-complementing rodent repair deficiency complementation group 1 (ERCC1) locus. The ordered clone map has been obtained starting from a foundation of cosmid contigs assembled by automated fingerprinting and localized to the cytogenetic map by fluorescence in situ hybridization (FISH). Clonal continuity of the map has been achieved by binning and linking the premapped cosmid contigs by means of yeast artificial chromosomes (YACs). The map consists of a single contig composed of 169 YAC members (minimal spanning path of 18 YACs) linking 165 cosmid contigs. Eighty percent, or about 13.2 Mb of the entire region spanned by the map, has been resolved to the EcoRI restriction map level. Twenty-nine sequence-tagged sites associated with genetic markers or derived from FISH-mapped cosmids have been placed on the map. In addition to the ERCC1 gene area, the map includes the location of the creatine kinase muscle locus (CKM), imidazoledipetidase (PEPD), glucophosphate isomerase (GPI), myelin-associated glycoprotein (MAG), the apolipoprotein E and C (APOE and APOC) genes, and the ryanodine receptor (RYR1) gene. This type of map provides a source of continuously overlapping DNA segments at a level of resolution two orders of magnitude higher than that obtained using YACs alone. In addition, it provides ready-to-use reagents for detailed analyses at the gene level, FISH studies of chromosomal aberrations, and DNA sequencing.

Chromosome Mapping

An integrated metric physical map of human chromosome 19.

A metric physical map of human chromosome 19 has been generated. The foundation of the map is sets of overlapping cosmids (contigs) generated by automated fingerprinting spanning over 95% of the euchromatin, about 50 megabases (Mb). Distances between selected cosmid clones were estimated using fluorescence in situ hybridization in sperm pronuclei, providing both order and distance between contigs. An average inter-marker separation of 230 kb has been obtained across the non-centromeric portion of the chromosome. Various types of larger insert clones were used to span gaps between contigs. Currently, the map consists of 51 'islands' containing multiple clone types, whose size, order and relative distance are known. Over 450 genes, genetic markers, sequence tagged sites (STSs), anonymous cDNAs, and other markers have been localized. In addition, EcoRI restriction maps have been generated for > 41 Mb (approximately 83%) of the chromosome.

Base Sequence