PubMed Health⌕ Search

Biomedical subjects

Webb Miller

Publications and source records attributed to Webb Miller.

17 recordsLinked to original sources

EnteriX 2003: Visualization tools for genome alignments of Enterobacteriaceae.

We describe EnteriX, a suite of three web-based visualization tools for graphically portraying alignment information from comparisons among several fixed and user-supplied sequences from related enterobacterial species, anchored on a reference genome (http://bio.cse.psu.edu/). The first visualization, Enteric, displays stacked pairwise alignments between a reference genome and each of the related bacteria, represented schematically as PIPs (Percent Identity Plots). Encoded in the views are large-scale genomic rearrangement events and functional landmarks. The second visualization, Menteric, computes and displays 1 Kb views of nucleotide-level multiple alignments of the sequences, together with annotations of genes, regulatory sites and conserved regions. The third, a Java-based tool named Maj, displays alignment information in two formats, corresponding roughly to the Enteric and Menteric views, and adds zoom-in capabilities. The uses of such tools are diverse, from examining the multiple sequence alignment to infer conserved sites with potential regulatory roles, to scrutinizing the commonalities and differences between the genomes for pathogenicity or phylogenetic studies. The EnteriX suite currently includes >15 enterobacterial genomes, generates views centered on four different anchor genomes and provides support for including user sequences in the alignments.

Computer Graphics↗

MultiPipMaker and supporting tools: Alignments and analysis of multiple genomic DNA sequences.

Analysis of multiple sequence alignments can generate important, testable hypotheses about the phylogenetic history and cellular function of genomic sequences. We describe the MultiPipMaker server, which aligns multiple, long genomic DNA sequences quickly and with good sensitivity (available at http://bio.cse.psu.edu/ since May 2001). Alignments are computed between a contiguous reference sequence and one or more secondary sequences, which can be finished or draft sequence. The outputs include a stacked set of percent identity plots, called a MultiPip, comparing the reference sequence with subsequent sequences, and a nucleotide-level multiple alignment. New tools are provided to search MultiPipMaker output for conserved matches to a user-specified pattern and for conserved matches to position weight matrices that describe transcription factor binding sites (singly and in clusters). We illustrate the use of MultiPipMaker to identify candidate regulatory regions in WNT2 and then demonstrate by transfection assays that they are functional. Analysis of the alignments also confirms the phylogenetic inference that horses are more closely related to cats than to cows.

Algorithms↗

Transcription-associated mutational asymmetry in mammalian evolution.

Although mutation is commonly thought of as a random process, evolutionary studies show that different types of nucleotide substitution occur with widely varying rates that presumably reflect biases intrinsic to mutation and repair mechanisms. A strand asymmetry, the occurrence of particular substitution types at higher rates than their complementary types, that is associated with DNA replication has been found in bacteria and mitochondria. A strand asymmetry that is associated with transcription and attributable to higher rates of cytosine deamination on the coding strand has been observed in enterobacteria. Here, we describe a qualitatively different transcription-associated strand asymmetry in mammals, which may be a byproduct of transcription-coupled repair in germline cells. This mutational asymmetry has acted over long periods of time to produce a compositional asymmetry, an excess of G+T over A+C on the coding strand, in most genes. The mutational and compositional asymmetries can be used to detect the orientations and approximate extents of transcribed regions.

Animals↗

Multispecies comparative analysis of a mammalian-specific genomic domain encoding secretory proteins.

The mammalian-specific casein gene cluster comprises 3 or 4 evolutionarily related genes and 1 physically linked gene with a functional association. To gain a better understanding of the mechanisms regulating the entire casein cluster at the genomic level we initiated a multispecies comparative sequence analysis. Despite the high level of divergence at the coding level, these studies have identified uncharacterized family members within two species and the presence at orthologous positions of previously uncharacterized genes. Also the previous suggestion that the histatin/statherin gene family, located in this region, was primate specific was ruled out. All 11 genes identified in this region appear to encode secretory proteins. Conservation of a number of noncoding regions was observed; one coincides with an element previously suggested to be important for beta-casein gene expression in human and cow. The conserved regions might have biological importance for the regulation of genes in this genomic "neighborhood."

Amino Acid Sequence↗

Significance of interspecies matches when evolutionary rate varies.

We develop techniques to estimate the statistical significance of gap-free alignments between two genomic DNA sequences, using human-mouse alignments as an example. The sequences are assumed to be sufficiently similar that some but not all of the neutrally evolving regions (i.e., those under no evolutionary constraint) can be reliably aligned. Our goal is to model the situation in which the neutral rate of evolution, and hence the extent of the aligning intervals, varies across the genome. In some cases, this permits the weaker of two matches to be judged as less likely to have arisen by chance, provided it lies in a genomic interval with a high level of background divergence. We employ a hidden Markov model to capture variations in divergence rates and assign probability values to gap-free alignments using techniques of Dembo and Karlin, which are related to those used for the same purpose by BLAST. Our methods are illustrated in detail using a 1.49 Mb genomic region. Results obtained from the analysis of human chromosome 22 using these techniques are also provided.

Animals↗

GALA, a database for genomic sequence alignments and annotations.

We have developed a relational database to contain whole genome sequence alignments between human and mouse with extensive annotations of the human sequence. Complex queries are supported on recorded features, both directly and on proximity among them. Searches can reveal a wide variety of relationships, such as finding all genes expressed in a designated tissue that have a highly conserved noncoding sequence 5' to the start site. Other examples are finding single nucleotide polymorphisms that occur in conserved noncoding regions upstream of genes and identifying CpG islands that overlap the 5' ends of divergently transcribed genes. The database is available online at http://globin.cse.psu.edu/ and http://bio.cse.psu.edu/.

5' Untranslated Regions↗

Human-mouse alignments with BLASTZ.

The Mouse Genome Analysis Consortium aligned the human and mouse genome sequences for a variety of purposes, using alignment programs that suited the various needs. For investigating issues regarding genome evolution, a particularly sensitive method was needed to permit alignment of a large proportion of the neutrally evolving regions. We selected a program called BLASTZ, an independent implementation of the Gapped BLAST algorithm specifically designed for aligning two long genomic sequences. BLASTZ was subsequently modified, both to attain efficiency adequate for aligning entire mammalian genomes and to increase its sensitivity. This work describes BLASTZ, its modifications, the hardware environment on which we run it, and several empirical studies to validate its results.

Animals↗

Distinguishing regulatory DNA from neutral sites.

We explore several computational approaches to analyzing interspecies genomic sequence alignments, aiming to distinguish regulatory regions from neutrally evolving DNA. Human-mouse genomic alignments were collected for three sets of human regions: (1) experimentally defined gene regulatory regions, (2) well-characterized exons (coding sequences, as a positive control), and (3) interspersed repeats thought to have inserted before the human-mouse split (a good model for neutrally evolving DNA). Models that potentially could distinguish functional noncoding sequences from neutral DNA were evaluated on these three data sets, as well as bulk genome alignments. Our analyses show that discrimination based on frequencies of individual nucleotide pairs or gaps (i.e., of possible alignment columns) is only partially successful. In contrast, scoring procedures that include the alignment context, based on frequencies of short runs of alignment columns, dramatically improve separation between regulatory and neutral features. Such scoring functions should aid in the identification of putative regulatory regions throughout the human genome.

Animals↗

Covariation in frequencies of substitution, deletion, transposition, and recombination during eutherian evolution.

Six measures of evolutionary change in the human genome were studied, three derived from the aligned human and mouse genomes in conjunction with the Mouse Genome Sequencing Consortium, consisting of (1) nucleotide substitution per fourfold degenerate site in coding regions, (2) nucleotide substitution per site in relics of transposable elements active only before the human-mouse speciation, and (3) the nonaligning fraction of human DNA that is nonrepetitive or in ancestral repeats; and three derived from human genome data alone, consisting of (4) SNP density, (5) frequency of insertion of transposable elements, and (6) rate of recombination. Features 1 and 2 are measures of nucleotide substitutions at two classes of "neutral" sites, whereas 4 is a measure of recent mutations. Feature 3 is a measure dominated by deletions in mouse, whereas 5 represents insertions in human. It was found that all six vary significantly in megabase-sized regions genome-wide, and many vary together. This indicates that some regions of a genome change slowly by all processes that alter DNA, and others change faster. Regional variation in all processes is correlated with, but not completely accounted for, by GC content in human and the difference between GC content in human and mouse.

Animals↗

Use of subtractive hybridization for comprehensive surveys of prokaryotic genome differences.

Comparative bacterial genomics shows that even different isolates of the same bacterial species can vary significantly in gene content. An effective means to survey differences across whole genomes would be highly advantageous for understanding this variation. Here we show that suppression subtractive hybridization (SSH) provides high, representative coverage of regions that differ between similar genomes. Using Helicobacter pylori strains 26695 and J99 as a model, SSH identified approximately 95% of the unique open reading frames in each strain, showing that the approach is effective. Furthermore, combining data from parallel SSH experiments using different restriction enzymes significantly increased coverage compared to using a single enzyme. These results suggest a powerful approach for assessing genome differences among closely related strains when one member of the group has been completely sequenced.

DNA Restriction Enzymes↗

Functional and binding studies of HS3.2 of the beta-globin locus control region.

The distal locus control region (LCR) is required for high-level expression of the complex of genes (HBBC) encoding the beta-like globins of mammals in erythroid cells. Several major DNase hypersensitive sites (HSs 1-5) mark the LCR. Sequence conservation and direct experimental evidence have implicated sequences within and between the HS cores in function of the LCR. In this report we confirm the mapping of a minor HS between HS3 and HS4, called HS3.2, and show that sequences including it increase the number of random integration sites at which a drug resistance gene is expressed. We also show that nuclear proteins including GATA1 and Oct1 bind specifically to sequences within HS3.2. However, the protein Pbx1, whose binding site is the best match to one highly conserved sequence, does not bind strongly. GATA1 and Oct1 also bind in the HS cores of the LCR and to promoters in HBBC. Their binding to this minor HS suggests that they may be used in assembly of a large complex containing multiple regulatory sequences.

Base Sequence↗

Genomic structure and functional control of the Dlx3-7 bigene cluster.

The Dlx genes are involved in early vertebrate morphogenesis, notably of the head. The six Dlx genes of mammals are arranged in three convergently transcribed bigene clusters. In this study, we examine the regulation of the Dlx3-7 cluster of the mouse. We obtained and sequenced human and mouse P1 clones covering the entire Dlx3-7 cluster. Comparative analysis of the human and mouse sequences revealed several highly conserved noncoding regions within 30 kb of the Dlx3-7-coding regions. These conserved elements were located both 5' of the coding exons of each gene and in the intergenic region 3' of the exons, suggesting that some enhancers might be shared between genes. We also found that the protein sequence of Dlx7 is evolving more rapidly than that of Dlx3. We conducted a functional study of the 79-kb mouse genomic clone to locate cis-element activity able to reproduce the endogenous expression pattern by using transgenic mice. We inserted a lacZ reporter gene into the first exon of the Dlx3 gene by using homologous recombination in yeast. Strong lacZ expression in embryonic (E) stage E9.5 and E10.5 mouse embryos was found in the limb buds and first and second visceral arches, consistent with the endogenous Dlx3 expression pattern. This result shows that the 79-kb region contains the major cis-elements required to direct the endogenous expression of Dlx3 at stage E10.5. To test for enhancer location, we divided the construct in the mid-intergenic region and injected the Dlx3 gene portion. This shortened fragment lacking Dlx7-flanking sequences is able to drive expression in the limb buds but not in the visceral arches. This observation is consistent with a cis-regulatory enhancer-sharing model within the Dlx bigene cluster.

Animals↗

HbVar: A relational database of human hemoglobin variants and thalassemia mutations at the globin gene server.

We have constructed a relational database of hemoglobin variants and thalassemia mutations, called HbVar, which can be accessed on the web at http://globin.cse.psu.edu. Extensive information is recorded for each variant and mutation, including a description of the variant and associated pathology, hematology, electrophoretic mobility, methods of isolation, stability information, ethnic occurrence, structure studies, functional studies, and references. The initial information was derived from books by Dr. Titus Huisman and colleagues [Huisman et al., 1996, 1997, 1998]. The current database is updated regularly with the addition of new data and corrections to previous data. Queries can be formulated based on fields in the database. Tables of common categories of variants, such as all those involving the alpha1-globin gene (HBA1) or all those that result in high oxygen affinity, are maintained by automated queries on the database. Users can formulate more precise queries, such as identifying "all beta-globin variants associated with instability and found in Scottish populations." This new database should be useful for clinical diagnosis as well as in fundamental studies of hemoglobin biochemistry, globin gene regulation, and human sequence variation at these loci.

Databases, Genetic↗

Candidate genes required for embryonic development: a comparative analysis of distal mouse chromosome 14 and human chromosome 13q22.

Mice homozygous for the Ednrb(s-1Acrg) deletion arrest at embryonic day 8.5 from defects associated with mesoderm development. To determine the molecular basis of this phenotype, we initiated a positional cloning of the Acrg minimal region. This region was predicted to be gene-poor by several criteria. From comparative analysis with the syntenic human locus at 13q22 and gene prediction program analysis, we found a single cluster of four genes within the 1.4-to 2-Mb contig over the Acrg minimal region that is flanked by a gene desert. We also found 130 highly conserved nonexonic sequences that were distributed over the gene cluster and desert. The four genes encode the TBC (Tre-2, BUB2, CDC16) domain-containing protein KIAA0603, the ubiquitin carboxy-terminal hydrolase L3 (UCHL3), the F-box/PDZ/LIM domain protein LMO7,and a novel gene. On the basis of their expression profile during development, all four genes are candidates for the Ednrb(s-1Acrg) embryonic lethality. Because we determined that a mutant of Uchl3 was viable, three candidate genes remain within the region.

Animals↗

PipTools: a computational toolkit to annotate and analyze pairwise comparisons of genomic sequences.

Sequence conservation between species is useful both for locating coding regions of genes and for identifying functional noncoding segments. Hence interspecies alignment of genomic sequences is an important computational technique. However, its utility is limited without extensive annotation. We describe a suite of software tools, PipTools, and related programs that facilitate the annotation of genes and putative regulatory elements in pairwise alignments. The alignment server PipMaker uses the output of these tools to display detailed information needed to interpret alignments. These programs are provided in a portable format for use on common desktop computers and both the toolkit and the PipMaker server can be found at our Web site (http://bio.cse.psu.edu/). We illustrate the utility of the toolkit using annotation of a pairwise comparison of the mouse MHC class II and class III regions with orthologous human sequences and subsequently identify conserved, noncoding sequences that are DNase I hypersensitive sites in chromatin of mouse cells.

Animals↗

Generation and comparative analysis of approximately 3.3 Mb of mouse genomic sequence orthologous to the region of human chromosome 7q11.23 implicated in Williams syndrome.

Williams syndrome is a complex developmental disorder that results from the heterozygous deletion of a approximately 1.6-Mb segment of human chromosome 7q11.23. These deletions are mediated by large (approximately 300 kb) duplicated blocks of DNA of near-identical sequence. Previously, we showed that the orthologous region of the mouse genome is devoid of such duplicated segments. Here, we extend our studies to include the generation of approximately 3.3 Mb of genomic sequence from the mouse Williams syndrome region, of which just over 1.4 Mb is finished to high accuracy. Comparative analyses of the mouse and human sequences within and immediately flanking the interval commonly deleted in Williams syndrome have facilitated the identification of nine previously unreported genes, provided detailed sequence-based information regarding 30 genes residing in the region, and revealed a number of potentially interesting conserved noncoding sequences. Finally, to facilitate comparative sequence analysis, we implemented several enhancements to the program, including the addition of links from annotated features within a generated percent-identity plot to specific records in public databases. Taken together, the results reported here provide an important comparative sequence resource that should catalyze additional studies of Williams syndrome, including those that aim to characterize genes within the commonly deleted interval and to develop mouse models of the disorder.

Animals↗

The Chlamydomonas reinhardtii plastid chromosome: islands of genes in a sea of repeats.

Chlamydomonas reinhardtii is a unicellular eukaryotic alga possessing a single chloroplast that is widely used as a model system for the study of photosynthetic processes. This report analyzes the surprising structural and evolutionary features of the completely sequenced 203,395-bp plastid chromosome. The genome is divided by 21.2-kb inverted repeats into two single-copy regions of approximately 80 kb and contains only 99 genes, including a full complement of tRNAs and atypical genes encoding the RNA polymerase. A remarkable feature is that >20% of the genome is repetitive DNA: the majority of intergenic regions consist of numerous classes of short dispersed repeats (SDRs), which may have structural or evolutionary significance. Among other sequenced chlorophyte plastid genomes, only that of the green alga Chlorella vulgaris appears to share this feature. The program MultiPipMaker was used to compare the genic complement of Chlamydomonas with those of other chloroplast genomes and to scan the genomes for sequence similarities and repetitive DNAs. Among the results was evidence that the SDRs were not derived from extant coding sequences, although some SDRs may have arisen from other genomic fragments. Phylogenetic reconstruction of changes in plastid genome content revealed that an accelerated rate of gene loss also characterized the Chlamydomonas/Chlorella lineage, a phenomenon that might be independent of the proliferation of SDRs. Together, our results reveal a dynamic and unusual plastid genome whose existence in a model organism will allow its features to be tested functionally.

Amino Acid Sequence↗