PubMed Health⌕ Search

Biomedical subjects

NISC Comparative Sequencing Program

Publications and source records attributed to NISC Comparative Sequencing Program.

At least 19 recordsLinked to original sources

Early history of mammals is elucidated with the ENCODE multiple species sequencing data.

Understanding the early evolution of placental mammals is one of the most challenging issues in mammalian phylogeny. Here, we addressed this question by using the sequence data of the ENCODE consortium, which include 1% of mammalian genomes in 18 species belonging to all main mammalian lineages. Phylogenetic reconstructions based on an unprecedented amount of coding sequences taken from 218 genes resulted in a highly supported tree placing the root of Placentalia between Afrotheria and Exafroplacentalia (Afrotheria hypothesis). This topology was validated by the phylogenetic analysis of a new class of genomic phylogenetic markers, the conserved noncoding sequences. Applying the tests of alternative topologies on the coding sequence dataset resulted in the rejection of the Atlantogenata hypothesis (Xenarthra grouping with Afrotheria), while this test rejected the second alternative scenario, the Epitheria hypothesis (Xenarthra at the base), when using the noncoding sequence dataset. Thus, the two datasets support the Afrotheria hypothesis; however, none can reject both of the remaining topological alternatives.

Animals↗

Gene duplication and inactivation in the HPRT gene family.

Hypoxanthine phosphoribosyltransferase (HPRT1) is a key enzyme in the purine salvage pathway, and mutations in HPRT1 cause Lesch-Nyhan disease. The studies described here utilized targeted comparative mapping and sequencing, in conjunction with database searches, to assemble a collection of 53 HPRT1 homologs from 28 vertebrates. Phylogenetic analysis of these homologs revealed that the HPRT gene family expanded as the result of ancient vertebrate-specific duplications and is composed of three groups consisting of HPRT1, phosphoribosyl transferase domain containing protein 1 (PRTFDC1), and HPRT1L genes. All members of the vertebrate HPRT gene family share a common intron-exon structure; however, we have found that the three gene groups have distinct rates of evolution and potentially divergent functions. Finally, we report our finding that PRTFDC1 was recently inactivated in the mouse lineage and propose the loss of function of this gene as a candidate genetic basis for the phenotypic disparity between HPRT-deficient humans and mice.

Animals↗

A preliminary comparative analysis of primate segmental duplications shows elevated substitution rates and a great-ape expansion of intrachromosomal duplications.

Compared with other sequenced animal genomes, human segmental duplications appear larger, more interspersed, and disproportionately represented as high-sequence identity alignments. Global sequence divergence estimates of human duplications have suggested an expansion relatively recently during hominoid evolution. Based on primate comparative sequence analysis of 37 unique duplication-transition regions, we establish a molecular clock for their divergence that shows a significant increase in their effective substitution rate when compared with unique genomic sequence. Fluorescent in situ hybridization (FISH) analyses from 1053 random nonhuman primate BACs indicate that great-ape species have been enriched for interspersed segmental duplications compared with representative Old World and New World monkeys. These findings support computational analyses that show a 12-fold excess of recent (>98%) intrachromosomal duplications when compared with duplications between nonhomologous chromosomes. These architectural shifts in genomic structure and elevated substitution rates have important implications for the emergence of new genes, gene-expression differences, and structural variation among humans and great apes.

Animals↗

Variable molecular clocks in hominoids.

Generation time is an important determinant of a neutral molecular clock. There are several human-specific life history traits that led to a substantially longer generation time in humans than in other hominoids. Indeed, a long generation time is considered an important trait that distinguishes humans from their closest relatives. Therefore, humans may exhibit a significantly slower molecular clock as compared to other hominoids. To investigate this hypothesis, we performed a large-scale analysis of lineage-specific rates of single-nucleotide substitutions among hominoids. We found that humans indeed exhibit a significant slowdown of molecular evolution compared to chimpanzees and other hominoids. However, the amount of fixed differences between humans and chimpanzees appears extremely small, suggesting a very recent evolution of human-specific life history traits. Notably, chimpanzees also exhibit a slower rate of molecular evolution compared to gorillas and orangutans in the regions analyzed.

Animals↗

The gene of retroviral origin Syncytin 1 is specific to hominoids and is inactive in Old World monkeys.

Syncytin 1 is one of the best known examples of recent acquisition of a new gene from an endogenous retrovirus (HERV) in the human genome and has been implicated in placental physiology. Within primates, Syncytin 1 is conserved in all hominoids but has not been characterized in Old World monkeys (OWMs). In this study, we investigated the status of Syncytin 1 in 14 hominoid and OWM species. We show that although the HERV-W provirus responsible for the origin of this gene was present in the genome of the most recent common ancestor of hominoids and OWMs, Syncytin 1 is inactive in OWMs. In addition, we were able to determine that the evolution of Syncytin 1 in hominoids involved an accumulation of amino acid changes and showed signatures of both positive and purifying selection. Our results indicate that Syncytin 1 is indeed a hominoid-specific gene and illustrate the complex and dynamic process associated with the origin of new genes.

Animals↗

Deletion of long-range sequences at Sox10 compromises developmental expression in a mouse model of Waardenburg-Shah (WS4) syndrome.

The transcription factor SOX10 is mutated in the human neurocristopathy Waardenburg-Shah syndrome (WS4), which is characterized by enteric aganglionosis and pigmentation defects. SOX10 directly regulates genes expressed in neural crest lineages, including the enteric ganglia and melanocytes. Although some SOX10 target genes have been reported, the mechanisms by which SOX10 expression is regulated remain elusive. Here, we describe a transgene-insertion mutant mouse line (Hry) that displays partial enteric aganglionosis, a loss of melanocytes, and decreased Sox10 expression in homozygous embryos. Mutation analysis of Sox10 coding sequences was negative, suggesting that non-coding regulatory sequences are disrupted. To isolate the Hry molecular defect, Sox10 genomic sequences were collected from multiple species, comparative sequence analysis was performed and software was designed (ExactPlus) to identify identical sequences shared among species. Mutation analysis of conserved sequences revealed a 15.9 kb deletion located 47.3 kb upstream of Sox10 in Hry mice. ExactPlus revealed three clusters of highly conserved sequences within the deletion, one of which shows strong enhancer potential in cultured melanocytes. These studies: (i) present a novel hypomorphic Sox10 mutation that results in a WS4-like phenotype in mice; (ii) demonstrate that a 15.9 kb deletion underlies the observed phenotype and likely removes sequences essential for Sox10 expression; (iii) combine a novel in silico method for comparative sequence analysis with in vitro functional assays to identify candidate regulatory sequences deleted in this strain. These studies will direct further analyses of Sox10 regulation and provide candidate sequences for mutation detection in WS4 patients lacking a SOX10-coding mutation.

Algorithms↗

Progressive proximal expansion of the primate X chromosome centromere.

Previous studies of the pericentromeric region of the human X chromosome short arm (Xp) revealed an age gradient from ancient DNA that contains expressed genes to recent human-specific DNA at the functional centromere. We analyzed the finished sequence of this human genomic region to investigate its evolutionary history. Phylogenetic analysis of >1,500 alpha-satellite monomers from the region revealed the presence of five physical domains, each containing monomers from a distinct phylogenetic clade. The most distal domain contains long interspersed nucleotide element repeats that were active >35 million years ago, whereas the four proximal domains contain more recently active long interspersed nucleotide element repeats. An out-of-register, unequal recombination (i.e., crossover) detected at the edge of the X chromosome-specific alpha-satellite array (DXZ1) may reflect the most recent of a series of punctuating events during evolution that resulted in a proximal physical expansion of the X centromere. The first 18 kb of this array has 97-99% pairwise identity among all 2-kb repeat units. To perform more detailed evolutionary comparisons, we sequenced the junction between the ancient DNA of Xp and the primate-specific alpha satellite in chimpanzee, gorilla, orangutan, vervet, macaque, and baboon. The striking conservation found in all cases supports the ancestral nature of the alpha satellite at this location. These studies demonstrate that the primate X centromere appears to have evolved through repeated expansion events occurring within the central, active region of centromeric DNA, with the newly added sequences then conferring centromere function.

Animals↗

Distribution and intensity of constraint in mammalian genomic sequence.

Comparisons of orthologous genomic DNA sequences can be used to characterize regions that have been subject to purifying selection and are enriched for functional elements. We here present the results of such an analysis on an alignment of sequences from 29 mammalian species. The alignment captures approximately 3.9 neutral substitutions per site and spans approximately 1.9 Mbp of the human genome. We identify constrained elements from 3 bp to over 1 kbp in length, covering approximately 5.5% of the human locus. Our estimate for the total amount of nonexonic constraint experienced by this locus is roughly twice that for exonic constraint. Constrained elements tend to cluster, and we identify large constrained regions that correspond well with known functional elements. While constraint density inversely correlates with mobile element density, we also show the presence of unambiguously constrained elements overlapping mammalian ancestral repeats. In addition, we describe a number of elements in this region that have undergone intense purifying selection throughout mammalian evolution, and we show that these important elements are more numerous than previously thought. These results were obtained with Genomic Evolutionary Rate Profiling (GERP), a statistically rigorous and biologically transparent framework for constrained element identification. GERP identifies regions at high resolution that exhibit nucleotide substitution deficits, and measures these deficits as "rejected substitutions". Rejected substitutions reflect the intensity of past purifying selection and are used to rank and characterize constrained elements. We anticipate that GERP and the types of analyses it facilitates will provide further insights and improved annotation for the human genome as mammalian genome sequence data become richer.

Animals↗

An initial strategy for the systematic identification of functional elements in the human genome by low-redundancy comparative sequencing.

With the recent completion of a high-quality sequence of the human genome, the challenge is now to understand the functional elements that it encodes. Comparative genomic analysis offers a powerful approach for finding such elements by identifying sequences that have been highly conserved during evolution. Here, we propose an initial strategy for detecting such regions by generating low-redundancy sequence from a collection of 16 eutherian mammals, beyond the 7 for which genome sequence data are already available. We show that such sequence can be accurately aligned to the human genome and used to identify most of the highly conserved regions. Although not a long-term substitute for generating high-quality genomic sequences from many mammalian species, this strategy represents a practical initial approach for rapidly annotating the most evolutionarily conserved sequences in the human genome, providing a key resource for the systematic study of human genome function.

Animals↗

Comparative sequencing provides insights about the structure and conservation of marsupial and monotreme genomes.

Sequencing and comparative analyses of genomes from multiple vertebrates are providing insights about the genetic basis for biological diversity. To date, these efforts largely have focused on eutherian mammals, chicken, and fish. In this article, we describe the generation and study of genomic sequences from noneutherian mammals, a group of species occupying unusual phylogenetic positions. A large sequence data set (totaling >5 Mb) was generated for the same orthologous region in three marsupial (North American opossum, South American opossum, and Australian tammar wallaby) and one monotreme (platypus) genomes. These ancient mammalian genomes are characterized by unusual architectural features with respect to G + C and repeat content, as well as compression relative to human. Approximately 14% and 34% of the human sequence forms alignments with the orthologous sequence from platypus and the marsupials, respectively; these numbers are distinctly lower than that observed with nonprimate eutherian mammals (45-70%). The alignable sequences between human and each marsupial species are not completely overlapping (only 80% common to all three species) nor are the platypus-alignable sequences completely contained within the marsupial-alignable sequences. Phylogenetic analysis of synonymous coding positions reveals that platypus has a notably long branch length, with the human-platypus substitution rate being on average 55% greater than that seen with human-marsupial pairs. Finally, analyses of the major mammalian lineages reveal distinct patterns with respect to the common presence of evolutionarily conserved vertebrate sequences. Our results confirm that genomic sequence from noneutherian mammals can contribute uniquely to unraveling the functional and evolutionary histories of the mammalian genome.

Animals↗

Detection of potential GDF6 regulatory elements by multispecies sequence comparisons and identification of a skeletal joint enhancer.

The identification of noncoding functional elements within vertebrate genomes, such as those that regulate gene expression, is a major challenge. Comparisons of orthologous sequences from multiple species are effective at detecting highly conserved regions and can reveal potential regulatory sequences. The GDF6 gene controls developmental patterning of skeletal joints and is associated with numerous, distant cis-acting regulatory elements. Using sequence data from 14 vertebrate species, we performed novel multispecies comparative analyses to detect highly conserved sequences flanking GDF6. The complementary tools WebMCS and ExactPlus identified a series of multispecies conserved sequences (MCSs). Of particular interest are MCSs within noncoding regions previously shown to contain GDF6 regulatory elements. A previously reported conserved sequence at -64 kb was also detected by both WebMCS and ExactPlus. Analysis of LacZ-reporter transgenic mice revealed that a 440-bp segment from this region contains an enhancer for Gdf6 expression in developing proximal limb joints. Several other MCSs represent candidate GDF6 regulatory elements; many of these are not conserved in fish or frog, but are strongly conserved in mammals.

Animals↗

Uprobe: a genome-wide universal probe resource for comparative physical mapping in vertebrates.

Interspecies comparisons are important for deciphering the functional content and evolution of genomes. The expansive array of >70 public vertebrate genomic bacterial artificial chromosome (BAC) libraries can provide a means of comparative mapping, sequencing, and functional analysis of targeted chromosomal segments that is independent and complementary to whole-genome sequencing. However, at the present time, no complementary resource exists for the efficient targeted physical mapping of the majority of these BAC libraries. Universal overgo-hybridization probes, designed from regions of sequenced genomes that are highly conserved between species, have been demonstrated to be an effective resource for the isolation of orthologous regions from multiple BAC libraries in parallel. Here we report the application of the universal probe design principal across entire genomes, and the subsequent creation of a complementary probe resource, Uprobe, for screening vertebrate BAC libraries. Uprobe currently consists of whole-genome sets of universal overgo-hybridization probes designed for screening mammalian or avian/reptilian libraries. Retrospective analysis, experimental validation of the probe design process on a panel of representative BAC libraries, and estimates of probe coverage across the genome indicate that the majority of all eutherian and avian/reptilian genes or regions of interest can be isolated using Uprobe. Future implementation of the universal probe design strategy will be used to create an expanded number of whole-genome probe sets that will encompass all vertebrate genomes.

Alligators and Crocodiles↗

An intermediate grade of finished genomic sequence suitable for comparative analyses.

Although the cost of generating draft-quality genomic sequence continues to decline, refining that sequence by the process of "sequence finishing" remains expensive. Near-perfect finished sequence is an appropriate goal for the human genome and a small set of reference genomes; however, such a high-quality product cannot be cost-justified for large numbers of additional genomes, at least for the foreseeable future. Here we describe the generation and quality of an intermediate grade of finished genomic sequence (termed comparative-grade finished sequence), which is tailored for use in multispecies sequence comparisons. Our analyses indicate that this sequence is very high quality (with the residual gaps and errors mostly falling within repetitive elements) and reflects 99% of the total sequence. Importantly, comparative-grade sequence finishing requires approximately 40-fold less reagents and approximately 10-fold less personnel effort compared to the generation of near-perfect finished sequence, such as that produced for the human genome. Although applied here to finishing sequence derived from individual bacterial artificial chromosome (BAC) clones, one could envision establishing routines for refining sequences emanating from whole-genome shotgun sequencing projects to a similar quality level. Our experience to date demonstrates that comparative-grade sequence finishing represents a practical and affordable option for sequence refinement en route to comparative analyses.

Animals↗

Multi-species sequence comparison reveals dynamic evolution of the elastin gene that has involved purifying selection and lineage-specific insertions/deletions.

BACKGROUND: The elastin gene (ELN) is implicated as a factor in both supravalvular aortic stenosis (SVAS) and Williams Beuren Syndrome (WBS), two diseases involving pronounced complications in mental or physical development. Although the complete spectrum of functional roles of the processed gene product remains to be established, these roles are inferred to be analogous in human and mouse. This view is supported by genomic sequence comparison, in which there are no large-scale differences in the ~1.8 Mb sequence block encompassing the common region deleted in WBS, with the exception of an overall reversed physical orientation between human and mouse. RESULTS: Conserved synteny around ELN does not translate to a high level of conservation in the gene itself. In fact, ELN orthologs in mammals show more sequence divergence than expected for a gene with a critical role in development. The pattern of divergence is non-conventional due to an unusually high ratio of gaps to substitutions. Specifically, multi-sequence alignments of eight mammalian sequences reveal numerous non-aligning regions caused by species-specific insertions and deletions, in spite of the fact that the vast majority of aligning sites appear to be conserved and undergoing purifying selection. CONCLUSIONS: The pattern of lineage-specific, in-frame insertions/deletions in the coding exons of ELN orthologous genes is unusual and has led to unique features of the gene in each lineage. These differences may indicate that the gene has a slightly different functional mechanism in mammalian lineages, or that the corresponding regions are functionally inert. Identified regions that undergo purifying selection reflect a functional importance associated with evolutionary pressure to retain those features.

Animals↗

Comparative sequence analysis of the Gdf6 locus reveals a duplicon-mediated chromosomal rearrangement in rodents and rapidly diverging coding and regulatory sequences.

Duplicated segments of genomic DNA can catalyze both gene evolution and chromosome evolution. Here we describe a rodent-specific duplication involving the Uqcrb gene, a cis-regulatory element for the Gdf6 gene, and a chromosomal rearrangement. Comparisons of Gdf6 sequences from several placental mammals and platypus revealed many strongly conserved regions flanking Gdf6 and the adjacent Uqcrb gene. However, in rat and mouse a synteny break resides approximately 70 kb upstream of Gdf6, such that Gdf6 and Uqcrb are on separate chromosomes. In rodents, Gdf6 and Uqcrb are both associated with homologous duplicons that may have catalyzed a rearrangement separating the two genes. However, the duplicon spanned both Uqcrb and a cis-regulatory element that controls Gdf6 transcription in limb skeletal joints. In mouse and rat, one duplicon now contains a degrading Uqcrb pseudogene but retains strongly conserved sequences within a Gdf6 enhancer. In contrast, the other duplicon has retained the intact Uqcrb gene and (in mouse) a copy of the Gdf6 enhancer that has acquired novel mutations. The duplicons have separately maintained distinct functions of the ancestral sequence, consistent with a "subfunction partitioning" evolutionary model. These findings also provide an example of a duplication that mobilized a tissue-specific enhancer from its cognate gene, and new evidence that duplications can be associated with chromosomal rearrangements. Furthermore, these data suggest that segmental duplications could lead to evolution of novel gene expression patterns via diversification of regulatory elements.

Animals↗

MultiPipMaker and supporting tools: Alignments and analysis of multiple genomic DNA sequences.

Analysis of multiple sequence alignments can generate important, testable hypotheses about the phylogenetic history and cellular function of genomic sequences. We describe the MultiPipMaker server, which aligns multiple, long genomic DNA sequences quickly and with good sensitivity (available at http://bio.cse.psu.edu/ since May 2001). Alignments are computed between a contiguous reference sequence and one or more secondary sequences, which can be finished or draft sequence. The outputs include a stacked set of percent identity plots, called a MultiPip, comparing the reference sequence with subsequent sequences, and a nucleotide-level multiple alignment. New tools are provided to search MultiPipMaker output for conserved matches to a user-specified pattern and for conserved matches to position weight matrices that describe transcription factor binding sites (singly and in clusters). We illustrate the use of MultiPipMaker to identify candidate regulatory regions in WNT2 and then demonstrate by transfection assays that they are functional. Analysis of the alignments also confirms the phylogenetic inference that horses are more closely related to cats than to cows.

Algorithms↗

LAGAN and Multi-LAGAN: efficient tools for large-scale multiple alignment of genomic DNA.

To compare entire genomes from different species, biologists increasingly need alignment methods that are efficient enough to handle long sequences, and accurate enough to correctly align the conserved biological features between distant species. We present LAGAN, a system for rapid global alignment of two homologous genomic sequences, and Multi-LAGAN, a system for multiple global alignment of genomic sequences. We tested our systems on a data set consisting of greater than 12 Mb of high-quality sequence from 12 vertebrate species. All the sequence was derived from the genomic region orthologous to an approximately 1.5-Mb region on human chromosome 7q31.3. We found that both LAGAN and Multi-LAGAN compare favorably with other leading alignment methods in correctly aligning protein-coding exons, especially between distant homologs such as human and chicken, or human and fugu. Multi-LAGAN produced the most accurate alignments, while requiring just 75 minutes on a personal computer to obtain the multiple alignment of all 12 sequences. Multi-LAGAN is a practical method for generating multiple alignments of long genomic sequences at any evolutionary distance. Our systems are publicly available at http://lagan.stanford.edu.

Animals↗

Transcription-associated mutational asymmetry in mammalian evolution.

Although mutation is commonly thought of as a random process, evolutionary studies show that different types of nucleotide substitution occur with widely varying rates that presumably reflect biases intrinsic to mutation and repair mechanisms. A strand asymmetry, the occurrence of particular substitution types at higher rates than their complementary types, that is associated with DNA replication has been found in bacteria and mitochondria. A strand asymmetry that is associated with transcription and attributable to higher rates of cytosine deamination on the coding strand has been observed in enterobacteria. Here, we describe a qualitatively different transcription-associated strand asymmetry in mammals, which may be a byproduct of transcription-coupled repair in germline cells. This mutational asymmetry has acted over long periods of time to produce a compositional asymmetry, an excess of G+T over A+C on the coding strand, in most genes. The mutational and compositional asymmetries can be used to detect the orientations and approximate extents of transcribed regions.

Animals↗