PubMed Health⌕ Search

Biomedical subjects

Gustavo Glusman

Publications and source records attributed to Gustavo Glusman.

16 recordsLinked to original sources

Genetic mapping at 3-kilobase resolution reveals inositol 1,4,5-triphosphate receptor 3 as a risk factor for type 1 diabetes in Sweden.

We mapped the genetic influences for type 1 diabetes (T1D), using 2,360 single-nucleotide polymorphism (SNP) markers in the 4.4-Mb human major histocompatibility complex (MHC) locus and the adjacent 493 kb centromeric to the MHC, initially in a survey of 363 Swedish T1D cases and controls. We confirmed prior studies showing association with T1D in the MHC, most significantly near HLA-DR/DQ. In the region centromeric to the MHC, we identified a peak of association within the inositol 1,4,5-triphosphate receptor 3 gene (ITPR3; formerly IP3R3). The most significant single SNP in this region was at the center of the ITPR3 peak of association (P=1.7 x 10(-4) for the survey study). For validation, we typed an additional 761 Swedish individuals. The P value for association computed from all 1,124 individuals was 1.30 x 10(-6) (recessive odds ratio 2.5; 95% confidence interval [CI] 1.7-3.9). The estimated population-attributable risk of 21.6% (95% CI 10.0%-31.0%) suggests that variation within ITPR3 reflects an important contribution to T1D in Sweden. Two-locus regression analysis supports an influence of ITPR3 variation on T1D that is distinct from that of any MHC class II gene.

Adolescent↗

KLK31P is a novel androgen regulated and transcribed pseudogene of kallikreins that is expressed at lower levels in prostate cancer cells than in normal prostate cells.

BACKGROUND: Fifteen human tissue kallikrein (KLK) genes have been identified as a cluster on chromosome 19. KLK expression is associated with various human diseases including cancers. Noncoding RNAs such as PCA3/DD3 and PCGEM1 have been identified in prostate cancer cells. METHODS: Using massively parallel signature sequencing (MPSS) technology, RT-PCR, and 5' rapid amplification of cDNA ends (RACE), we identified and cloned a novel gene that maps to the KLK locus. RESULTS: We have characterized this gene, named as KLK31P by the HUGO Gene Nomenclature Committee, as an unprocessed KLK pseudogene. It contains five exons, two of which are KLK-derived while the rest are "exonized" interspersed repeats. KLK31P is expressed abundantly in prostate tissues and is androgen regulated. KLK31P is expressed at lower levels in localized and metastatic prostate cancer cells than in normal prostate cells. CONCLUSIONS: KLK31P is a novel androgen regulated and transcribed pseudogene of kallikreins that may play a role in prostate carcinogenesis or maintenance.

Amino Acid Sequence↗

A third approach to gene prediction suggests thousands of additional human transcribed regions.

The identification and characterization of the complete ensemble of genes is a main goal of deciphering the digital information stored in the human genome. Many algorithms for computational gene prediction have been described, ultimately derived from two basic concepts: (1) modeling gene structure and (2) recognizing sequence similarity. Successful hybrid methods combining these two concepts have also been developed. We present a third orthogonal approach to gene prediction, based on detecting the genomic signatures of transcription, accumulated over evolutionary time. We discuss four algorithms based on this third concept: Greens and CHOWDER, which quantify mutational strand biases caused by transcription-coupled DNA repair, and ROAST and PASTA, which are based on strand-specific selection against polyadenylation signals. We combined these algorithms into an integrated method called FEAST, which we used to predict the location and orientation of thousands of putative transcription units not overlapping known genes. Many of the newly predicted transcriptional units do not appear to code for proteins. The new algorithms are particularly apt at detecting genes with long introns and lacking sequence conservation. They therefore complement existing gene prediction methods and will help identify functional transcripts within many apparent "genomic deserts."

Algorithms↗

The evolution of vertebrate Toll-like receptors.

The complete sequences of Takifugu Toll-like receptor (TLR) loci and gene predictions from many draft genomes enable comprehensive molecular phylogenetic analysis. Strong selective pressure for recognition of and response to pathogen-associated molecular patterns has maintained a largely unchanging TLR recognition in all vertebrates. There are six major families of vertebrate TLRs. This repertoire is distinct from that of invertebrates. TLRs within a family recognize a general class of pathogen-associated molecular patterns. Most vertebrates have exactly one gene ortholog for each TLR family. The family including TLR1 has more species-specific adaptations than other families. A major family including TLR11 is represented in humans only by a pseudogene. Coincidental evolution plays a minor role in TLR evolution. The sequencing phase of this study produced finished genomic sequences for the 12 Takifugu rubripes TLRs. In addition, we have produced >70 gene models, including sequences from the opossum, chicken, frog, dog, sea urchin, and sea squirt.

Animals↗

Interchromosomal segmental duplications explain the unusual structure of PRSS3, the gene for an inhibitor-resistant trypsinogen.

Homo sapiens possess several trypsinogen or trypsinogen-like genes of which three (PRSS1, PRSS2, and PRSS3) produce functional trypsins in the digestive tract. PRSS1 and PRSS2 are located on chromosome 7q35, while PRSS3 is found on chromosome 9p13. Here, we report a variation of the theme of new gene creation by duplication: the PRSS3 gene was formed by segmental duplications originating from chromosomes 7q35 and 11q24. As a result, PRSS3 transcripts display two variants of exon 1. The PRSS3 transcript whose gene organization most resembles PRSS1 and PRSS2 encodes a functional protein originally named mesotrypsinogen. The other variant is a fusion transcript, called trypsinogen IV. We show that the first exon of trypsinogen IV is derived from the noncoding first exon of LOC120224, a chromosome 11 gene. LOC120224 codes for a widely conserved transmembrane protein of unknown function. Comparative analyses suggest that these interchromosomal duplications occurred after the divergence of Old World monkeys and hominids. PRSS3 transcripts consist of a mixed population of mRNAs, some expressed in the pancreas and encoding an apparently functional trypsinogen and others of unknown function expressed in brain and a variety of other tissues. Analysis of the selection pressures acting on the trypsinogen gene family shows that, while the apparently functional genes are under mild to strong purifying selection overall, a few residues appear under positive selection. These residues could be involved in interactions with inhibitors.

Chromosomes, Human↗

T1DBase, a community web-based resource for type 1 diabetes research.

T1DBase (http://T1DBase.org) is a public website and database that supports the type 1 diabetes (T1D) research community. The site is currently focused on the molecular genetics and biology of T1D susceptibility and pathogenesis. It includes the following datasets: annotated genome sequence for human, rat and mouse; information on genetically identified T1D susceptibility regions in human, rat and mouse, and genetic linkage and association studies pertaining to T1D; descriptions of NOD mouse congenic strains; the Beta Cell Gene Expression Bank, which reports expression levels of genes in beta cells under various conditions, and annotations of gene function in beta cells; data on gene expression in a variety of tissues and organs; and biological pathways from KEGG and BioCarta. Tools on the site include the GBrowse genome browser, site-wide context dependent search, Connect-the-Dots for connecting gene and other identifiers from multiple data sources, Cytoscape for visualizing and analyzing biological networks, and the GESTALT workbench for genome annotation. All data are open access and all software is open source.

Animals↗

A comparison of the human and chimpanzee olfactory receptor gene repertoires.

Olfactory receptor (OR) genes constitute the basis of the sense of smell and are encoded by the largest mammalian gene superfamily, with >1000 members. In humans, but not in mice or dogs, the majority of OR genes have become pseudogenes, suggesting that OR genes in humans evolve under different selection pressures than in other mammals. To explore this further, we compare the OR gene repertoire of human with its closest living evolutionary relative, by taking advantage of the recently sequenced genome of the chimpanzee. In agreement with previous reports based on a small number of ORs, we find that humans have a significantly higher proportion of OR pseudogenes than chimpanzees. Moreover, we can reject the possibility that humans have been accumulating OR pseudogenes at a constant neutral rate since the divergence of human and chimpanzee. The comparison of the two repertoires reveals two chimpanzee-specific OR subfamily expansions and three expansions specific to humans. It also suggests that a subset of OR genes are under positive selection in either the human or the chimpanzee lineage. Thus, although overall there is relaxed constraint on human olfaction relative to chimpanzee, species-specific sensory requirements appear to have shaped the evolution of the functional OR gene repertoires in both species.

Animals↗

Structural and genetic diversity of group B streptococcus capsular polysaccharides.

Group B Streptococcus (GBS) is an important pathogen of neonates, pregnant women, and immunocompromised individuals. GBS isolates associated with human infection produce one of nine antigenically distinct capsular polysaccharides which are thought to play a key role in virulence. A comparison of GBS polysaccharide structures of all nine known GBS serotypes together with the predicted amino acid sequences of the proteins that direct their synthesis suggests that the evolution of serotype-specific capsular polysaccharides has proceeded through en bloc replacement of individual glycosyltransferase genes with DNA sequences that encode enzymes with new linkage specificities. We found striking heterogeneity in amino acid sequences of synthetic enzymes with very similar functions, an observation that supports horizontal gene transfer rather than stepwise mutagenesis as a mechanism for capsule variation. Eight of the nine serotypes appear to be closely related both structurally and genetically, whereas serotype VIII is more distantly related. This similarity in polysaccharide structure strongly suggests that the evolutionary pressure toward antigenic variation exerted by acquired immunity is counterbalanced by a survival advantage conferred by conserved structural motifs of the GBS polysaccharides.

Bacterial Capsules↗

An enigmatic fourth runt domain gene in the fugu genome: ancestral gene loss versus accelerated evolution.

BACKGROUND: The runt domain transcription factors are key regulators of developmental processes in bilaterians, involved both in cell proliferation and differentiation, and their disruption usually leads to disease. Three runt domain genes have been described in each vertebrate genome (the RUNX gene family), but only one in other chordates. Therefore, the common ancestor of vertebrates has been thought to have had a single runt domain gene. RESULTS: Analysis of the genome draft of the fugu pufferfish (Takifugu rubripes) reveals the existence of a fourth runt domain gene, FrRUNT, in addition to the orthologs of human RUNX1, RUNX2 and RUNX3. The tiny FrRUNT packs six exons and two putative promoters in just 3 kb of genomic sequence. The first exon is located within an intron of FrSUPT3H, the ortholog of human SUPT3H, and the first exon of FrSUPT3H resides within the first intron of FrRUNT. The two gene structures are therefore "interlocked". In the human genome, SUPT3H is instead interlocked with RUNX2. FrRUNT has no detectable ortholog in the genomes of mammals, birds or amphibians. We consider alternative explanations for an apparent contradiction between the phylogenetic data and the comparison of the genomic neighborhoods of human and fugu runt domain genes. We hypothesize that an ancient RUNT locus was lost in the tetrapod lineage, together with FrFSTL6, a member of a novel family of follistatin-like genes. CONCLUSIONS: Our results suggest that the runt domain family may have started expanding in chordates much earlier than previously thought, and exemplify the importance of detailed analysis of whole-genome draft sequence to provide new insights into gene evolution.

Amino Acid Sequence↗

Low-pass sequencing for microbial comparative genomics.

BACKGROUND: We studied four extremely halophilic archaea by low-pass shotgun sequencing: (1) the metabolically versatile Haloarcula marismortui; (2) the non-pigmented Natrialba asiatica; (3) the psychrophile Halorubrum lacusprofundi and (4) the Dead Sea isolate Halobaculum gomorrense. Approximately one thousand single pass genomic sequences per genome were obtained. The data were analyzed by comparative genomic analyses using the completed Halobacterium sp. NRC-1 genome as a reference. Low-pass shotgun sequencing is a simple, inexpensive, and rapid approach that can readily be performed on any cultured microbe. RESULTS: As expected, the four archaeal halophiles analyzed exhibit both bacterial and eukaryotic characteristics as well as uniquely archaeal traits. All five halophiles exhibit greater than sixty percent GC content and low isoelectric points (pI) for their predicted proteins. Multiple insertion sequence (IS) elements, often involved in genome rearrangements, were identified in H. lacusprofundi and H. marismortui. The core biological functions that govern cellular and genetic mechanisms of H. sp. NRC-1 appear to be conserved in these four other halophiles. Multiple TATA box binding protein (TBP) and transcription factor IIB (TFB) homologs were identified from most of the four shotgunned halophiles. The reconstructed molecular tree of all five halophiles shows a large divergence between these species, but with the closest relationship being between H. sp. NRC-1 and H. lacusprofundi. CONCLUSION: Despite the diverse habitats of these species, all five halophiles share (1) high GC content and (2) low protein isoelectric points, which are characteristics associated with environmental exposure to UV radiation and hypersalinity, respectively. Identification of multiple IS elements in the genome of H. lacusprofundi and H. marismortui suggest that genome structure and dynamic genome reorganization might be similar to that previously observed in the IS-element rich genome of H. sp. NRC-1. Identification of multiple TBP and TFB homologs in these four halophiles are consistent with the hypothesis that different types of complex transcriptional regulation may occur through multiple TBP-TFB combinations in response to rapidly changing environmental conditions. Low-pass shotgun sequence analyses of genomes permit extensive and diverse analyses, and should be generally useful for comparative microbial genomics.

Archaea↗

Genetic divergence of the rhesus macaque major histocompatibility complex.

The major histocompatibility complex (MHC) is comprised of the class I, class II, and class III regions, including the MHC class I and class II genes that play a primary role in the immune response and serve as an important model in studies of primate evolution. Although nonhuman primates contribute significantly to comparative human studies, relatively little is known about the genetic diversity and genomics underlying nonhuman primate immunity. To address this issue, we sequenced a complete rhesus macaque MHC spanning over 5.3 Mb, and obtained an additional 2.3 Mb from a second haplotype, including class II and portions of class I and class III. A major expansion of from six class I genes in humans to as many as 22 active MHC class I genes in rhesus and levels of sequence divergence some 10-fold higher than a similar human comparison were found, averaging from 2% to 6% throughout extended portions of class I and class II. These data pose new interpretations of the evolutionary constraints operating between MHC diversity and T-cell selection by contrasting with models predicting an optimal number of antigen presenting genes. For the clinical model, these data and derivative genetic tools can be implemented in ongoing genetic and disease studies that involve the rhesus macaque.

Animals↗

Genome sequence of Haloarcula marismortui: a halophilic archaeon from the Dead Sea.

We report the complete sequence of the 4,274,642-bp genome of Haloarcula marismortui, a halophilic archaeal isolate from the Dead Sea. The genome is organized into nine circular replicons of varying G+C compositions ranging from 54% to 62%. Comparison of the genome architectures of Halobacterium sp. NRC-1 and H. marismortui suggests a common ancestor for the two organisms and a genome of significantly reduced size in the former. Both of these halophilic archaea use the same strategy of high surface negative charge of folded proteins as means to circumvent the salting-out phenomenon in a hypersaline cytoplasm. A multitiered annotation approach, including primary sequence similarities, protein family signatures, structure prediction, and a protein function association network, has assigned putative functions for at least 58% of the 4242 predicted proteins, a far larger number than is usually achieved in most newly sequenced microorganisms. Among these assigned functions were genes encoding six opsins, 19 MCP and/or HAMP domain signal transducers, and an unusually large number of environmental response regulators-nearly five times as many as those encoded in Halobacterium sp. NRC-1--suggesting H. marismortui is significantly more physiologically capable of exploiting diverse environments. In comparing the physiologies of the two halophilic archaea, in addition to the expected extensive similarity, we discovered several differences in their metabolic strategies and physiological responses such as distinct pathways for arginine breakdown in each halophile. Finally, as expected from the larger genome, H. marismortui encodes many more functions and seems to have fewer nutritional requirements for survival than does Halobacterium sp. NRC-1.

Archaeal Proteins↗

Human Gene-Centric Databases at the Weizmann Institute of Science: GeneCards, UDB, CroW 21 and HORDE.

Recent enhancements and current research in the GeneCards (GC) (http://bioinfo.weizmann.ac.il/cards/) project are described, including the addition of gene expression profiles and integrated gene locations. Also highlighted are the contributions of specialized associated human gene-centric databases developed at the Weizmann Institute. These include the Unified Database (UDB) (http://bioinfo.weizmann.ac.il/udb) for human genome mapping, the human Chromosome 21 database at the Weizmann Insti-tute (CroW 21) (http://bioinfo.weizmann.ac.il/crow21), and the Human Olfactory Receptor Data Explora-torium (HORDE) (http://bioinfo.weizmann.ac.il/HORDE). The synergistic relationships amongst these efforts have positively impacted the quality, quantity and usefulness of the GeneCards gene compendium.

Algorithms↗

Initial sequencing and comparative analysis of the mouse genome.

The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.

Animals↗

Whole-genome shotgun assembly and analysis of the genome of Fugu rubripes.

The compact genome of Fugu rubripes has been sequenced to over 95% coverage, and more than 80% of the assembly is in multigene-sized scaffolds. In this 365-megabase vertebrate genome, repetitive DNA accounts for less than one-sixth of the sequence, and gene loci occupy about one-third of the genome. As with the human genome, gene loci are not evenly distributed, but are clustered into sparse and dense regions. Some "giant" genes were observed that had average coding sequence sizes but were spread over genomic lengths significantly larger than those of their human orthologs. Although three-quarters of predicted human proteins have a strong match to Fugu, approximately a quarter of the human proteins had highly diverged from or had no pufferfish homologs, highlighting the extent of protein evolution in the 450 million years since teleosts and mammals diverged. Conserved linkages between Fugu and human genes indicate the preservation of chromosomal segments from the common vertebrate ancestor, but with considerable scrambling of gene order.

Animals↗

Phylogenesis and regulated expression of the RUNT domain transcription factors RUNX1 and RUNX3.

The RUNX transcription factors are key regulators of lineage specific gene expression in developmental pathways. The mammalian RUNX genes arose early in evolution and maintained extensive structural similarities. Sequence analysis suggested that RUNX3 is the most ancient of the three mammalian genes, consistent with its role in neurogenesis of the monosynaptic reflex arc, the simplest neuronal response circuit, found in Cnidarians, the most primitive animals. All RUNX proteins bind to the same DNA motif and act as activators or repressors of transcription through recruitment of common transcriptional modulators. Nevertheless, analysis of Runx1 and Runx3 expression during embryogenesis revealed that their function is not redundant. In adults both Runx1 and Runx3 are highly expressed in the hematopoietic system. At early embryonic stages we found strong Runx3 expression in dorsal root ganglia neurons, confined to TrkC sensory neurons. In the absence of Runx3, knockout mice develop severe ataxia due to the early death of the TrkC neurons. Other phenotypic defects of Runx3 KO mice including abnormalities in thymopoiesis are also being investigated.

Animals↗