PubMed Health⌕ Search

Biomedical subjects

James C Mullikin

Publications and source records attributed to James C Mullikin.

4 recordsLinked to original sources

Revisiting the mouse mitochondrial DNA sequence.

The existence of reliable mtDNA reference sequences for each species is of great relevance in a variety of fields, from phylogenetic and population genetics studies to pathogenetic determination of mtDNA variants in humans or in animal models of mtDNA-linked diseases. We present compelling evidence for the existence of sequencing errors on the current mouse mtDNA reference sequence. This includes the deletion of a full codon in two genes, the substitution of one amino acid on five occasions and also the involvement of tRNA and rRNA genes. The conclusions are supported by: (i) the re-sequencing of the original cell line used by Bibb and Clayton, the LA9 cell line, (ii) the sequencing of a second L-derivative clone (L929), and (iii) the comparison with 12 other mtDNA sequences from live mice, 10 of them maternally related with the mouse from which the L cells were generated. Two of the latest sequences are reported for the first time in this study (Balb/cJ and C57BL/6J). In addition, we found that both the LA9 and L929 mtDNAs also contain private clone polymorphic variants that, at least in the case of L929, promote functional impairment of the oxidative phosphorylation system. Consequently, the mtDNA of the strain used for the mouse genome project (C57BL/6J) is proposed as the new standard for the mouse mtDNA sequence.

Animals↗

The phusion assembler.

The Phusion assembler has assembled the mouse genome from the whole-genome shotgun (WGS) dataset collected by the Mouse Genome Sequencing Consortium, at ~7.5x sequence coverage, producing a high-quality draft assembly 2.6 gigabases in size, of which 90% of these bases are in 479 scaffolds. For the mouse genome, which is a large and repeat-rich genome, the input dataset was designed to include a high proportion of paired end sequences of various size selected inserts, from 2-200 kbp lengths, into various host vector templates. Phusion uses sequence data, called reads, and information about reads that share common templates, called read pairs, to drive the assembly of this large genome to highly accurate results. The preassembly stage, which clusters the reads into sensible groups, is a key element of the entire assembler, because it permits a simple approach to parallelization of the assembly stage, as each cluster can be treated independent of the others. In addition to the application of Phusion to the mouse genome, we will also present results from the WGS assembly of Caenorhabditis briggsae sequenced to about 11x coverage. The C. briggsae assembly was accessioned through EMBL, http://www.ebi.ac.uk/services/index.html, using the series CAAC01000001-CAAC01000578, however, the Phusion mouse assembly described here was not accessioned. The mouse data was generated by the Mouse Genome Sequencing Consortium. The C. briggsae sequence was generated at The Wellcome Trust Sanger Institute and the Genome Sequencing Center, Washington University School of Medicine.

Animals↗

The mosaic structure of variation in the laboratory mouse genome.

Most inbred laboratory mouse strains are known to have originated from a mixed but limited founder population in a few laboratories. However, the effect of this breeding history on patterns of genetic variation among these strains and the implications for their use are not well understood. Here we present an analysis of the fine structure of variation in the mouse genome, using single nucleotide polymorphisms (SNPs). When the recently assembled genome sequence from the C57BL/6J strain is aligned with sample sequence from other strains, we observe long segments of either extremely high (approximately 40 SNPs per 10 kb) or extremely low (approximately 0.5 SNPs per 10 kb) polymorphism rates. In all strain-to-strain comparisons examined, only one-third of the genome falls into long regions (averaging >1 Mb) of a high SNP rate, consistent with estimated divergence rates between Mus musculus domesticus and either M. m. musculus or M. m. castaneus. These data suggest that the genomes of these inbred strains are mosaics with the vast majority of segments derived from domesticus and musculus sources. These observations have important implications for the design and interpretation of positional cloning experiments.

Albinism↗

Human genome sequence variation and the influence of gene history, mutation and recombination.

Variation in the human genome sequence is key to understanding susceptibility to disease in modern populations and the history of ancestral populations. Unlocking this information requires knowledge of the patterns and underlying causes of human sequence diversity. By applying a new population-genetic framework to two genome-wide polymorphism surveys, we find that the human genome contains sizeable regions (stretching over tens of thousands of base pairs) that have intrinsically high and low rates of sequence variation. We show that the primary determinant of these patterns is shared genealogical history. Only a fraction of the variation (at most 25%) is due to the local mutation rate. By measuring the average distance over which genealogical histories are typically preserved, these data provide the first genome-wide estimate of the average extent of correlation among variants (linkage disequilibrium). The results are best explained by extreme variability in the recombination rate at a fine scale, and provide the first empirical evidence that such recombination 'hot spots' are a general feature of the human genome and have a principal role in shaping genetic variation in the human population.

Animals↗