PubMed Health⌕ Search

Biomedical subjects

Richard K Wilson

Publications and source records attributed to Richard K Wilson.

At least 19 recordsLinked to original sources

Molecular refinement of gibbon genome rearrangements.

The gibbon karyotype is known to be extensively rearranged when compared to the human and to the ancestral primate karyotype. By combining a bioinformatics (paired-end sequence analysis) approach and a molecular cytogenetics approach, we have refined the synteny block arrangement of the white-cheeked gibbon (Nomascus leucogenys, NLE) with respect to the human genome. We provide the first detailed clone framework map of the gibbon genome and refine the location of 86 evolutionary breakpoints to <1 Mb resolution. An additional 12 breakpoints, mapping primarily to centromeric and telomeric regions, were mapped to approximately 5 Mb resolution. Our combined FISH and BES analysis indicates that we have effectively subcloned 49 of these breakpoints within NLE gibbon BAC clones, mapped to a median resolution of 79.7 kb. Interestingly, many of the intervals associated with translocations were gene-rich, including some genes associated with normal skeletal development. Comparisons of NLE breakpoints with those of other gibbon species reveal variability in the position, suggesting that chromosomal rearrangement has been a longstanding property of this particular ape lineage. Our data emphasize the synergistic effect of combining computational genomics and cytogenetics and provide a framework for ultimate sequence and assembly of the gibbon genome.

Animals↗

Identification and analysis of genes expressed in the adult filarial parasitic nematode Dirofilaria immitis.

The heartworm Dirofilaria immitis is a filarial parasitic nematode infecting dogs and other mammals worldwide causing fatal complications. Here, we present the first large-scale survey of the adult heartworm transcriptome by generation and analysis of 4005 expressed sequence tags, identifying about 1800 genes and expanding the available sequence information for the parasite significantly. Brugia malayi genomic data offered the most valuable information to interpret heartworm genes, with about 70% of D. immitis genes showing significant similarities to the assembly. Comparative genomic analyses revealed both genes common to metazoans or nematodes and genes specific to filarial parasites that may relate to parasitism. Characterization of abundant transcripts suggested important roles for genes involved in energy generation and antioxidant defense in adults. In particular, we proposed that adult heartworm likely adopted an anaerobic electron transfer-based energy generation system distinct from the aerobic pathway utilized by its mammalian host, making it a promising target in developing next generation macrofilaricides and other treatments. Our survey provided novel insights into the D. immitis transcriptome and laid a foundation for further comparative studies on biology, parasitism and evolution within the phylum Nematoda.

Animals↗

Use of cigarette-smoking history to estimate the likelihood of mutations in epidermal growth factor receptor gene exons 19 and 21 in lung adenocarcinomas.

PURPOSE: Lung adenocarcinomas with mutations in exons 19 and 21 of the epidermal growth factor receptor gene (EGFR) demonstrate sensitivity to gefitinib or erlotinib. Investigators have reported an association between EGFR mutations and the amount and duration of cigarette smoking, with the highest incidence of mutations seen in never smokers. METHODS: EGFR exon 19 and 21 mutation status was determined in 265 tumor samples using direct sequencing, polymerase chain reaction (PCR), or PCR-based restriction fragment length polymorphism analysis. A detailed smoking history was obtained. Patients were categorized as never smokers (< 100 lifetime cigarettes), former smokers (quit > or = 1 year ago), or current smokers (quit < 1 year ago). RESULTS: We detected EGFR mutations in 34 (51%) of 67 never smokers (95% CI, 38% to 64%), 29 (19%) of 151 former smokers (95% CI, 13% to 27%), and two (4%) of 47 current smokers (95% CI, 1% to 16%). Significantly fewer EGFR mutations were found in people who smoked for more than 15 pack-years (P < .001) or stopped smoking less than 25 years ago (P < .02) compared with individuals who never smoked. The number of smoking pack-years and smoke-free years predicted the prevalence of EGFR mutations (areas under receiver operating characteristic curve = 0.78 and 0.77, respectively). CONCLUSION: The likelihood of EGFR mutations in exons 19 and 21 decreases as the number of pack-years increases. Mutations were less common in people who smoked for more than 15 pack-years or who stopped smoking cigarettes less than 25 years ago. These data can assist clinicians in assessing the likelihood of exon 19 and 21 EGFR mutations in patients with lung adenocarcinoma when mutational analysis is not feasible.

Adenocarcinoma↗

Application of a superword array in genome assembly.

We introduce a data structure called a superword array for finding quickly matches between DNA sequences. The superword array possesses some desirable features of the lookup table and suffix array. We describe simple algorithms for constructing and using a superword array to find pairs of sequences that share a unique superword. The algorithms are implemented in a genome assembly program called PCAP.REP for computation of overlaps between reads. Experimental results produced by PCAP.REP and PCAP on a whole-genome dataset show that PCAP.REP produced a more accurate and contiguous assembly than PCAP.

Algorithms↗

Physical map-assisted whole-genome shotgun sequence assemblies.

We describe a targeted approach to improve the contiguity of whole-genome shotgun sequence (WGS) assemblies at run-time, using information from Bacterial Artificial Chromosome (BAC)-based physical maps. Clone sizes and overlaps derived from clone fingerprints are used for the calculation of length constraints between any two BAC neighbors sharing 40% of their size. These constraints are used to promote the linkage and guide the arrangement of sequence contigs within a sequence scaffold at the layout phase of WGS assemblies. This process is facilitated by FASSI, a stand-alone application that calculates BAC end and BAC overlap length constraints from clone fingerprint map contigs created by the FPC package. FASSI is designed to work with the assembly tool PCAP, but its output can be formatted to work with other WGS assembly algorithms able to use length constraints for individual clones. The FASSI method is simple to implement, potentially cost-effective, and has resulted in the increase of scaffold contiguity for both the Drosophila melanogaster and Cryptococcus gattii genomes when compared to a control assembly without map-derived constraints. A 6.5-fold coverage draft DNA sequence of the Pan troglodytes (chimpanzee) genome was assembled using map-derived constraints and resulted in a 26.1% increase in scaffold contiguity.

Animals↗

After the duplication: gene loss and adaptation in Saccharomyces genomes.

The ancient duplication of the Saccharomyces cerevisiae genome and subsequent massive loss of duplicated genes is apparent when it is compared to the genomes of related species that diverged before the duplication event. To learn more about the evolutionary effects of the duplication event, we compared the S. cerevisiae genome to other Saccharomyces genomes. We demonstrate that the whole genome duplication occurred before S. castellii diverged from S. cerevisiae. In addition to more accurately dating the duplication event, this finding allowed us to study the effects of the duplication on two separate lineages. Analyses of the duplication regions of the genomes indicate that most of the duplicated genes (approximately 85%) were lost before the speciation. Only a small amount of paralogous gene loss (4-6%) occurred after speciation. On the other hand, S. castellii appears to have lost several hundred genes that were not retained as duplicated paralogs. These losses could be related to genomic rearrangements that reduced the number of chromosomes from 16 to 9. In addition to S. castellii, other Saccharomyces sensu lato species likely diverged from S. cerevisiae after the duplication. A thorough analysis of these species will likely reveal other important outcomes of the whole genome duplication.

Adaptation, Physiological↗

A genome-wide comparison of recent chimpanzee and human segmental duplications.

We present a global comparison of differences in content of segmental duplication between human and chimpanzee, and determine that 33% of human duplications (> 94% sequence identity) are not duplicated in chimpanzee, including some human disease-causing duplications. Combining experimental and computational approaches, we estimate a genomic duplication rate of 4-5 megabases per million years since divergence. These changes have resulted in gene expression differences between the species. In terms of numbers of base pairs affected, we determine that de novo duplication has contributed most significantly to differences between the species, followed by deletion of ancestral duplications. Post-speciation gene conversion accounts for less than 10% of recent segmental duplication. Chimpanzee-specific hyperexpansion (> 100 copies) of particular segments of DNA have resulted in marked quantitative differences and alterations in the genome landscape between chimpanzee and human. Almost all of the most extreme differences relate to changes in chromosome structure, including the emergence of African great ape subterminal heterochromatin. Nevertheless, base per base, large segmental duplication events have had a greater impact (2.7%) in altering the genomic landscape of these two species than single-base-pair substitution (1.2%).

Animals↗

Conservation of Y-linked genes during human evolution revealed by comparative sequencing in chimpanzee.

The human Y chromosome, transmitted clonally through males, contains far fewer genes than the sexually recombining autosome from which it evolved. The enormity of this evolutionary decline has led to predictions that the Y chromosome will be completely bereft of functional genes within ten million years. Although recent evidence of gene conversion within massive Y-linked palindromes runs counter to this hypothesis, most unique Y-linked genes are not situated in palindromes and have no gene conversion partners. The 'impending demise' hypothesis thus rests on understanding the degree of conservation of these genes. Here we find, by systematically comparing the DNA sequences of unique, Y-linked genes in chimpanzee and human, which diverged about six million years ago, evidence that in the human lineage, all such genes were conserved through purifying selection. In the chimpanzee lineage, by contrast, several genes have sustained inactivating mutations. Gene decay in the chimpanzee lineage might be a consequence of positive selection focused elsewhere on the Y chromosome and driven by sperm competition.

Animals↗

Reduced PU.1 expression causes myeloid progenitor expansion and increased leukemia penetrance in mice expressing PML-RARalpha.

PU.1 is a member of the ETS family of transcription factors that is known to be important for hematopoietic development. Recently, haploinsufficiency for PU.1 has been shown to cause a shift in myelomonocytic progenitor fate toward the myeloid lineage. We have previously shown that transgenic mice expressing PML-RARalpha (PR) and RARalpha-PML frequently develop acute promyelocytic leukemia (APL) in association with a large (>20 Mb) interstitial deletion of chromosome 2 that includes PU.1. To directly assess the relevance of levels of expression of PU.1 for leukemia progression, we bred hCG-PR mice with PU.1+/- mice and assessed their phenotype. Young, nonleukemic hCG-PR x PU.1+/- mice developed splenomegaly because of the abnormal expansion of myeloid cells in their spleens. hCG-PR x PU.1+/- mice developed a typical APL syndrome after a long latent period, but the penetrance of disease was 84%, compared with 7% in hCG-PR x PU.1+/+ mice (P < 0.0001). The residual PU.1 allele in hCG-PR x PU.1+/- APL cells was expressed, and complete exonic resequencing revealed no detectable mutations in nine of nine samples. However, PR expression in U937 myelomonocytic cells and primary murine myeloid bone marrow cells caused a reduction in PU.1 mRNA levels. Therefore, the loss of one copy of PU.1 through a deletional mechanism, plus down-regulation of the residual allele caused by PR expression, may synergize to expand the pool of myeloid progenitors that are susceptible to transformation, increasing the penetrance of APL.

Animals↗

Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.

We have conducted a comprehensive search for conserved elements in vertebrate genomes, using genome-wide multiple alignments of five vertebrate species (human, mouse, rat, chicken, and Fugu rubripes). Parallel searches have been performed with multiple alignments of four insect species (three species of Drosophila and Anopheles gambiae), two species of Caenorhabditis, and seven species of Saccharomyces. Conserved elements were identified with a computer program called phastCons, which is based on a two-state phylogenetic hidden Markov model (phylo-HMM). PhastCons works by fitting a phylo-HMM to the data by maximum likelihood, subject to constraints designed to calibrate the model across species groups, and then predicting conserved elements based on this model. The predicted elements cover roughly 3%-8% of the human genome (depending on the details of the calibration procedure) and substantially higher fractions of the more compact Drosophila melanogaster (37%-53%), Caenorhabditis elegans (18%-37%), and Saccharaomyces cerevisiae (47%-68%) genomes. From yeasts to vertebrates, in order of increasing genome size and general biological complexity, increasing fractions of conserved bases are found to lie outside of the exons of known protein-coding genes. In all groups, the most highly conserved elements (HCEs), by log-odds score, are hundreds or thousands of bases long. These elements share certain properties with ultraconserved elements, but they tend to be longer and less perfectly conserved, and they overlap genes of somewhat different functional categories. In vertebrates, HCEs are associated with the 3' UTRs of regulatory genes, stable gene deserts, and megabase-sized regions rich in moderately conserved noncoding sequences. Noncoding HCEs also show strong statistical evidence of an enrichment for RNA secondary structure.

3' Untranslated Regions↗

Punctuated duplication seeding events during the evolution of human chromosome 2p11.

Primate genomic sequence comparisons are becoming increasingly useful for elucidating the evolutionary history and organization of our own genome. Such studies are particularly informative within human pericentromeric regions--areas of particularly rapid change in genomic structure. Here, we present a systematic analysis of the evolutionary history of one approximately 700-kb region of 2p11, including the first autosomal transition from pericentromeric sequence to higher-order alpha-satellite DNA. We show that this region is composed of segmental duplications corresponding to 14 ancestral segments ranging in size from 4 kb to approximately 115 kb. These duplicons show 94%-98.5% sequence identity to their ancestral loci. Comparative FISH and phylogenetic analysis indicate that these duplicons are differentially distributed in human, chimpanzee, and gorilla genomes, whereas baboon has a single putative ancestral locus for all but one of the duplications. Our analysis supports a model where duplicative transposition events occurred during a narrow window of evolution after the separation of the human/ape lineage from the Old World monkeys (10-20 million years ago). Although dramatic secondary dispersal events occurred during the radiation of the human, chimpanzee, and gorilla lineages, duplicative transposition seeding events of new material to this particular pericentromeric region abruptly ceased after this time period. The multiplicity of initial duplicative transpositions prior to the separation of humans and great-apes suggests a punctuated model for the formation of highly duplicated pericentromeric regions within the human genome. The data further indicate that factors other than sequence are important determinants for such bursts of duplicative transposition from the euchromatin to pericentromeric regions.

Animals↗

Investigating hookworm genomes by comparative analysis of two Ancylostoma species.

BACKGROUND: Hookworms, infecting over one billion people, are the mostly closely related major human parasites to the model nematode Caenorhabditis elegans. Applying genomics techniques to these species, we analyzed 3,840 and 3,149 genes from Ancylostoma caninum and A. ceylanicum. RESULTS: Transcripts originated from libraries representing infective L3 larva, stimulated L3, arrested L3, and adults. Most genes are represented in single stages including abundant transcripts like hsp-20 in infective L3 and vit-3 in adults. Over 80% of the genes have homologs in C. elegans, and nearly 30% of these were with observable RNA interference phenotypes. Homologies were identified to nematode-specific and clade V specific gene families. To study the evolution of hookworm genes, 574 A. caninum/A. ceylanicum orthologs were identified, all of which were found to be under purifying selection with distribution ratios of nonsynonymous to synonymous amino acid substitutions similar to that reported for C. elegans/C. briggsae orthologs. The phylogenetic distance between A. caninum and A. ceylanicum is almost identical to that for C. elegans/C. briggsae. CONCLUSION: The genes discovered should substantially accelerate research toward better understanding of the parasites' basic biology as well as new therapies including vaccines and novel anthelmintics.

Ancylostoma↗

Comparing low coverage random shotgun sequence data from Brassica oleracea and Oryza sativa genome sequence for their ability to add to the annotation of Arabidopsis thaliana.

Since the completion of the Arabidopsis thaliana genome sequence, there is an ongoing effort to annotate the genome as accurately as possible. Comparing genome sequences of related species complements the current annotation strategies by identifying genes and improving gene structure. A total of 595,321 Brassica oleracea shotgun reads were sequenced by TIGR (The Institute for Genome Research) and the collaboration of Washington University and Cold Spring Harbor. Vicogenta (a genome viewer based on GMOD and GBrowse) was created to view the current annotation and sequence alignments for Arabidopsis. Brassica reads were compared with the Arabidopsis genome and proteome databases using BLAST. Hypothetical genes and conserved unannotated regions on the short arm of chromosome 4 from Arabidopsis were experimentally verified using RT-PCR. We were able to improve the Arabidopsis annotation by identifying 25 genes that were missed, and confirming expression of 43 hypothetical genes in Arabidopsis. We were also able to detect conservation in genes whose transcription is normally suppressed due to methylation. We also examined how useful the O. sativa genome and ESTs from other species are, compared with Brassica, in improving the Arabidopsis annotation.

Amino Acid Sequence↗

Genome science: a video tour of the Washington University Genome Sequencing Center for high school and undergraduate students.

Sequencing of the human genome has ushered in a new era of biology. The technologies developed to facilitate the sequencing of the human genome are now being applied to the sequencing of other genomes. In 2004, a partnership was formed between Washington University School of Medicine Genome Sequencing Center's Outreach Program and Washington University Department of Biology Science Outreach to create a video tour depicting the processes involved in large-scale sequencing. "Sequencing a Genome: Inside the Washington University Genome Sequencing Center" is a tour of the laboratory that follows the steps in the sequencing pipeline, interspersed with animated explanations of the scientific procedures used at the facility. Accompanying interviews with the staff illustrate different entry levels for a career in genome science. This video project serves as an example of how research and academic institutions can provide teachers and students with access and exposure to innovative technologies at the forefront of biomedical research. Initial feedback on the video from undergraduate students, high school teachers, and high school students provides suggestions for use of this video in a classroom setting to supplement present curricula.

Feedback↗

A physical map of the chicken genome.

Strategies for assembling large, complex genomes have evolved to include a combination of whole-genome shotgun sequencing and hierarchal map-assisted sequencing. Whole-genome maps of all types can aid genome assemblies, generally starting with low-resolution cytogenetic maps and ending with the highest resolution of sequence. Fingerprint clone maps are based upon complete restriction enzyme digests of clones representative of the target genome, and ultimately comprise a near-contiguous path of clones across the genome. Such clone-based maps are used to validate sequence assembly order, supply long-range linking information for assembled sequences, anchor sequences to the genetic map and provide templates for closing gaps. Fingerprint maps are also a critical resource for subsequent functional genomic studies, because they provide a redundant and ordered sampling of the genome with clones. In an accompanying paper we describe the draft genome sequence of the chicken, Gallus gallus, the first species sequenced that is both a model organism and a global food source. Here we present a clone-based physical map of the chicken genome at 20-fold coverage, containing 260 contigs of overlapping clones. This map represents approximately 91% of the chicken genome and enables identification of chicken clones aligned to positions in other sequenced genomes.

Animals↗

A genetic variation map for chicken with 2.8 million single-nucleotide polymorphisms.

We describe a genetic variation map for the chicken genome containing 2.8 million single-nucleotide polymorphisms (SNPs). This map is based on a comparison of the sequences of three domestic chicken breeds (a broiler, a layer and a Chinese silkie) with that of their wild ancestor, red jungle fowl. Subsequent experiments indicate that at least 90% of the variant sites are true SNPs, and at least 70% are common SNPs that segregate in many domestic breeds. Mean nucleotide diversity is about five SNPs per kilobase for almost every possible comparison between red jungle fowl and domestic lines, between two different domestic lines, and within domestic lines--in contrast to the notion that domestic animals are highly inbred relative to their wild ancestors. In fact, most of the SNPs originated before domestication, and there is little evidence of selective sweeps for adaptive alleles on length scales greater than 100 kilobases.

Alleles↗

Comparison of genome degradation in Paratyphi A and Typhi, human-restricted serovars of Salmonella enterica that cause typhoid.

Salmonella enterica serovars often have a broad host range, and some cause both gastrointestinal and systemic disease. But the serovars Paratyphi A and Typhi are restricted to humans and cause only systemic disease. It has been estimated that Typhi arose in the last few thousand years. The sequence and microarray analysis of the Paratyphi A genome indicates that it is similar to the Typhi genome but suggests that it has a more recent evolutionary origin. Both genomes have independently accumulated many pseudogenes among their approximately 4,400 protein coding sequences: 173 in Paratyphi A and approximately 210 in Typhi. The recent convergence of these two similar genomes on a similar phenotype is subtly reflected in their genotypes: only 30 genes are degraded in both serovars. Nevertheless, these 30 genes include three known to be important in gastroenteritis, which does not occur in these serovars, and four for Salmonella-translocated effectors, which are normally secreted into host cells to subvert host functions. Loss of function also occurs by mutation in different genes in the same pathway (e.g., in chemotaxis and in the production of fimbriae).

Base Sequence↗