PubMed Health⌕ Search

Biomedical subjects

W R Gish

Publications and source records attributed to W R Gish.

6 recordsLinked to original sources

Initial sequencing and analysis of the human genome.

The human genome holds an extraordinary trove of information about human development, physiology, medicine and evolution. Here we report the results of an international collaboration to produce and make freely available a draft sequence of the human genome. We also present an initial analysis of the data, describing some of the insights that can be gleaned from the sequence.

Animals↗

Rapid gene mapping in Caenorhabditis elegans using a high density polymorphism map.

Single nucleotide polymorphisms (SNPs) are valuable genetic markers of human disease. They also comprise the highest potential density marker set available for mapping experimentally derived mutations in model organisms such as Caenorhabditis elegans. To facilitate the positional cloning of mutations we have identified polymorphisms in CB4856, an isolate from a Hawaiian island that shows a uniformly high density of polymorphisms compared with the reference Bristol N2 strain. Based on 5.4 Mbp of aligned sequences, we predicted 6,222 polymorphisms. Furthermore, 3,457 of these markers modify restriction enzyme recognition sites ('snip-SNPs') and are therefore easily detected as RFLPs. Of these, 493 were experimentally confirmed by restriction digest to produce a snip-SNP map of the worm genome. A mapping strategy using snip-SNPs and bulked segregant analysis (BSA) is outlined. CB4856 is crossed into a mutant strain, and exclusion of CB4856 alleles of a subset of snip-SNPs in mutant progeny is assessed with BSA. The proximity of a linked marker to the mutation is estimated by the relative proportion of each form of the biallelic marker in populations of wildtype and mutant genomes. The usefulness of this approach is illustrated by the rapid mapping of the dyf-5 gene.

Animals↗

Gene structure prediction and alternative splicing analysis using genomically aligned ESTs.

With the availability of a nearly complete sequence of the human genome, aligning expressed sequence tags (EST) to the genomic sequence has become a practical and powerful strategy for gene prediction. Elucidating gene structure is a complex problem requiring the identification of splice junctions, gene boundaries, and alternative splicing variants. We have developed a software tool, Transcript Assembly Program (TAP), to delineate gene structures using genomically aligned EST sequences. TAP assembles the joint gene structure of the entire genomic region from individual splice junction pairs, using a novel algorithm that uses the EST-encoded connectivity and redundancy information to sort out the complex alternative splicing patterns. A method called polyadenylation site scan (PASS) has been developed to detect poly-A sites in the genome. TAP uses these predictions to identify gene boundaries by segmenting the joint gene structure at polyadenylated terminal exons. Reconstructing 1007 known transcripts, TAP scored a sensitivity (Sn) of 60% and a specificity (Sp) of 92% at the exon level. The gene boundary identification process was found to be accurate 78% of the time. also reports alternative splicing patterns in EST alignments. An analysis of alternative splicing in 1124 genic regions suggested that more than half of human genes undergo alternative splicing. Surprisingly, we saw an absolute majority of the detected alternative splicing events affect the coding region. Furthermore, the evolutionary conservation of alternative splicing between human and mouse was analyzed using an EST-based approach. (See http://stl.wustl.edu/~zkan/TAP/)

Alternative Splicing↗

Surveying Saccharomyces genomes to identify functional elements by comparative DNA sequence analysis.

Comparative sequence analysis has facilitated the discovery of protein coding genes and important functional sequences within proteins, but has been less useful for identifying functional sequence elements in nonprotein-coding DNA because the relatively rapid rate of change of nonprotein-coding sequences and the relative simplicity of non-coding regulatory sequence elements necessitates the comparison of sequences of relatively closely related species. We tested the use of comparative DNA sequence analysis to aid identification of promoter regulatory elements, nonprotein-coding RNA genes, and small protein-coding genes by surveying random DNA sequences of several Saccharomyces yeast species, with the goal of learning which species are best suited for comparisons with S. cerevisiae. We also determined the DNA sequence of a few specific promoters and RNA genes of several Saccharomyces species to determine the degree of conservation of known functional elements within the genome. Our results lead us to conclude that comparative DNA sequence analysis will enable identification of functionally conserved elements within the yeast genome, and suggest a path for obtaining this information.

Base Sequence↗

A general approach to single-nucleotide polymorphism discovery.

Single-nucleotide polymorphisms (SNPs) are the most abundant form of human genetic variation and a resource for mapping complex genetic traits. The large volume of data produced by high-throughput sequencing projects is a rich and largely untapped source of SNPs (refs 2, 3, 4, 5). We present here a unified approach to the discovery of variations in genetic sequence data of arbitrary DNA sources. We propose to use the rapidly emerging genomic sequence as a template on which to layer often unmapped, fragmentary sequence data and to use base quality values to discern true allelic variations from sequencing errors. By taking advantage of the genomic sequence we are able to use simpler yet more accurate methods for sequence organization: fragment clustering, paralogue identification and multiple alignment. We analyse these sequences with a novel, Bayesian inference engine, POLYBAYES, to calculate the probability that a given site is polymorphic. Rigorous treatment of base quality permits completely automated evaluation of the full length of all sequences, without limitations on alignment depth. We demonstrate this approach by accurate SNP predictions in human ESTs aligned to finished and working-draft quality genomic sequences, a data set representative of the typical challenges of sequence-based SNP discovery.

Algorithms↗

Simian virus 40-transformed human cells that express large T antigens defective for viral DNA replication.

Many types of human cells cultured in vitro are generally semipermissive for simian virus 40 (SV40) replication. Consequently, subpopulations of stably transformed human cells often carry free viral DNA, which is presumed to arise via spontaneous excision from an integrated DNA template. Stably transformed human cell lines that do not have detectable free DNA are therefore likely to harbor harbor mutant viral genomes incapable of excision and replication, or these cells may synthesize variant cellular proteins necessary for viral replication. We examined four such cell lines and conclude that for the three lines SV80, GM638, and GM639, the cells did indeed harbor spontaneous T-antigen mutants. For the SV80 line, marker rescue (determined by a plaque assay) and DNA sequence analysis of cloned DNA showed that a single point mutation converting serine 147 to asparagine was the cause of the mutation. Similarly, a point mutation converting leucine 457 to methionine for the GM638 mutant T allele was found. Moreover, the SV80 line maintained its permissivity for SV40 DNA replication but did not complement the SV40 tsA209 mutant at its nonpermissive temperature. The cloned SV80 T-antigen allele, though replication incompetent, maintained its ability to transform rodent cells at wild-type efficiencies. A compilation of spontaneously occurring SV40 mutations which cannot replicate but can transform shows that these mutations tend to cluster in two regions of the T-antigen gene, one ascribed to the site-specific DNA-binding ability of the protein, and the other to the ATPase activity which is linked to its helicase activity.

Animals↗