PubMed HealthSearch

Biomedical subjects

R C Hardison

Publications and source records attributed to R C Hardison.

18 recordsLinked to original sources

The 5' ends of LINE1 repeats in rabbit DNA define subfamilies and reveal a short sequence conserved between rabbits and humans.

The 5' ends of five full-length LINE1 (L1) repeats from the rabbit genome (L1Oc) were mapped and their nucleotide sequences determined. Computer-generated alignments showed that these five L1Oc repeats can be divided into subfamilies, each of which has a characteristic sequence upstream of the first open reading frame (ORF1). These five L1Ocs range in size from 6.5 to 7.3 kb, with 5' ends located 76 to 1125 bp upstream of ORF1. Two of these subfamilies appear to have diverged from a common ancestor at least 66 million years ago. Comparisons of the 5' ends of L1s from rabbit, human, mouse, and rat show no common sequence 5' to ORF1, except for a 22-bp sequence that is found near the beginning of all characterized full-length L1s from rabbit and human. A statistical analysis indicates that this 22-bp aligned block is highly significant. Part of this 22-bp sequence matches the microE1 binding site in immunoglobulin gene enhancers. This strong conservation suggests that the microE1 binding site may be part of a transcriptional regulatory element at the 5' ends of rabbit and human L1 repeats.

Animals

Parallelization of a local similarity algorithm.

The local similarity problem is to determine the similar regions within two given sequences. We recently developed a dynamic programming algorithm for the local similarity problem that requires only space proportional to the sum of the two sequence lengths, whereas earlier methods use space proportional to the product of the lengths. In this paper, we describe how to parallelize the new algorithm and present results of experimental studies on an Intel hypercube. The parallel method provides rapid, high-resolution alignments for users of our software toolkit for pairwise sequence comparison, as illustrated here by a comparison of the chloroplast genomes of tobacco and liverwort.

Algorithms

Analysis of conserved domains and sequence motifs in cellular regulatory proteins and locus control regions using new software tools for multiple alignment and visualization.

With the tremendous expansion of molecular sequence data in recent years, multiple alignment is arguably one of the two most important analytic techniques (the other being fast database searching). A number of useful approaches to this problem have previously been developed, but often they are limited to only a subset of multiple-alignment applications and cannot easily deal with the complex structural organization seen in an increasing number of sequences. For example, a single sequence may contain several domains of different evolutionary origins, and the multiplicities and relative ordering of these domains may be quite different among related sequences. Here we describe an integrated set of interactive Unix tools that combines several multiple-alignment techniques with traditional "dot-plot" visualization to provide a flexible environment for approaching complex sequence analysis problems. We apply these tools to the identification and characterization of "catalytic" domains in ras and rho/rac GTPase-activating proteins, to "Src homology" (SH2, SH3) domains in cytoplasmic signaling proteins, to repetitive sequence motifs in the alpha and beta subunits of protein prenyltransferases, and to regulatory DNA sequences in the locus control region of the beta-globin gene cluster.

Alkyl and Aryl Transferases

Software tools for analyzing pairwise alignments of long sequences.

Pairwise comparison of long stretches of genomic DNA sequence can identify regions conserved across species, which often indicate functional significance. However, the novel insights frequently must be windowed from a flood of information; for instance, running an alignment program on two 50-kilobase sequences might yield over a hundred pages of alignments. Direct inspection of such a volume of printed output is infeasible, or at best highly undesirable, and computer tools are needed to summarize the information, to assist in its analysis, and to report the findings. This paper describes two such software tools. One tool prepares publication-quality pictorial representations of alignments, while another facilitates interactive browsing of pairwise alignment data. Their effectiveness is illustrated by comparing the beta-like globin gene clusters between humans and rabbits. A second example compares the chloroplast genomes of tobacco and liverwort.

Animals

An apparent pause site in the transcription unit of the rabbit alpha-globin gene.

Transcription of the rabbit alpha-globin gene begins primarily at the cap site, although some upstream start sites are also observed. Analysis by RNA polymerase run-on assays in nuclei shows that transcription continues at a high level past the polyadenylation site, after which the polymerase density actually increases in a region of about 400 nucleotides, followed by a gradual decline over the 700 nucleotides. These features are also observed in the transcription unit of the rabbit beta-globin gene. The region with the unexpectedly high nascent RNA hybridization signal in the 3' flank contains a conserved sequence, KGCAGCWGGR (K = G or T, W = A or T, R = A or G), followed by an inverted repeat. The inverted repeat (perhaps with the conserved sequence) may be a pause site for RNA polymerase II, thus accounting for the increase in polymerase density. This sequence and inverted repeat are found in the 3' flank of several globin genes and the simian virus 40 (SV40) early genes, as well as in the regions implicated in pausing or termination of transcription of eight different genes. Deletion of the conserved sequence and inverted repeat from the 3' flank of the SV40 early region causes a small increase in the levels of transcription downstream from this site. Replacement with the conserved sequence and inverted repeat from the rabbit alpha-globin gene causes an accumulation of polymerases, supporting the hypothesis that polymerases pause at this site. This proposed pause site may affect the efficiency of termination at some sites further downstream, perhaps by loss of a processivity factor.

Amino Acid Sequence

Survey of plastid RNA abundance during tomato fruit ripening: the amounts of RNA from the ORF 2280 region increase in chromoplasts.

A comprehensive survey of the levels of plastid RNAs at progressive stages of tomato fruit ripening was conducted by hybridizing total RNA with labeled Pst I fragments that cover almost the entire tomato plastid genome and with gene-specific probes. Two different cultivars of tomato (Lycopersicon esculentum Mill.) were examined, Traveler 76 and Count II. One of the tomato probes, P7, revealed a pronounced increase in the amount of an 8.3 kb RNA in ripe fruit. The homologous region of the tobacco plastid genome contains several genes for ribosomal proteins and a large unidentified open reading frame (2280 codons). Little change was observed in the levels of many transcripts during ripening. However, in some cases (e.g. psbA and psbC/D) the amount of RNA decreased during ripening of Count II but showed little or no change in Traveler 76. The contrast between Traveler 76 and Count II tomatoes shows that the level of plastid transcripts can vary substantially during fruit ripening with no obvious effect on the chloroplast to chromoplast transition. The large RNA from the P7 region may encode a protein that functions predominantly in chromoplasts.

Amino Acid Sequence

Localization of the alpha-like globin gene cluster to region q12 of rabbit chromosome 6 by in situ hybridization.

The rabbit (Oryctolagus cuniculus) alpha-like globin gene cluster (HBAC) contains several block duplications of zeta-, alpha and theta-globin genes. Using in situ hybridizations to metaphase chromosome spreads, the gene cluster has been mapped to region q12 of chromosome 6. Given that human HBAC maps to the short arm of chromosome 16, the mapping of rabbit HBAC to 6q12 confirms the assignment of homology between OCU6q and HSA16p based on similarities of chromosomal banding patterns. In both species, HBAC is in a very G + C-rich region within the most distal band of the chromosome.

Animals

Subfamily relationships and clustering of rabbit C repeats.

C repeats constitute the predominant family of short interspersed repeats (SINEs) in the rabbit genome. Determination of the nucleotide sequence 5' to rabbit zeta-globin genes reveals clusters of C repeats, and analysis of these and other sequenced regions of rabbit chromosomes shows that the C repeats have a strong tendency to insert within or in close proximity to other C repeats. An alignment of 44 members of the C repeat family shows that they are composites of different sequences, including a tRNA-like sequence, a conserved central core, a stretch of repeating CT dinucleotides, and an A-rich tract. Cladograms generated by both parsimony and cluster analysis subdivide the C repeats into at least three distinct subfamilies. Nucleotides at sites diagnostic for subfamilies appear to have changed in a punctuated and progressive manner during evolution, indicating that a limited number of progenitors have given rise to new repeats in waves of dispersion. C repeats that insert into preexisting C repeats belong to subfamilies that are proposed to have been propagated more recently; hence, these data support the model of dispersion in successive waves. The divergence among the oldest group of C repeats is greater than that observed for the analogous Alu repeats in humans, indicating that rabbit C repeats have been propagating longer than human Alu repeats. The improved consensus sequence for these repeats is similar to that of the predominant artiodactyl SINE in both the tRNA-like region and a central region. Because members of different subfamilies cross-hybridize very poorly, hybridization data with representatives of each subfamily provide a new minimal estimate, 234,000, for the copy number of C repeats in the rabbit haploid genome, although it is likely that the actual value is closer to 1 million.

Animals

Restriction site and genetic map of Cucurbita pepo chloroplast DNA.

A detailed restriction map of squash chloroplast DNA (cpDNA) was constructed with five restriction endonucleases, SalI, PvuII, BglI, SacII, and PstI. The cleavage sites were mapped by sequential digestion of cpDNA using low-gelling temperature agarose. The restriction map shows that squash cpDNA is an approximately 153 kilobase (kb) circle with a large inverted repeat sequence of 23.3 kb, separated by a large (83.7 kb) and a small (22.7 kb) single copy region. Genes for a number of chloroplast polypeptides were localized on the map by hybridizing the cpDNA restriction fragments to heterologous gene-specific probes from tobacco, pea, tomato, maize, and spinach chloroplasts. The gene locations and organization of squash cpDNA are highly conserved and similar to chloroplast genomes of tomato, pepper, and Ginkgo.

Chloroplasts

A space-efficient algorithm for local similarities.

Existing dynamic-programming algorithms for identifying similar regions of two sequences require time and space proportional to the product of the sequence lengths. Often this space requirement is more limiting than the time requirement. We describe a dynamic-programming local-similarity algorithm that needs only space proportional to the sum of the sequence lengths. The method can also find repeats within a single long sequence. To illustrate the algorithm's potential, we discuss comparison of a 73,360 nucleotide sequence containing the human beta-like globin gene cluster and a corresponding 44,594 nucleotide sequence for rabbit, a problem well beyond the capabilities of other dynamic-programming software.

Algorithms

Short interspersed repeats in rabbit DNA can provide functional polyadenylation signals.

Analysis of 37 short repetitive elements (SINEs) in rabbit DNA that are known as C repeats has revealed three that contribute functional polyadenylation signals to genes into which they have been inserted. Similar roles have been attributed to particular individual SINEs in rodents and primates before, suggesting that these roles may be common to SINEs in all mammalian orders. Although most SINEs appear to have little influence on the genome individually, the observation that three of 36 rabbit C repeats provide functional sequences suggests a mechanism for the maintenance of SINEs within mammalian genomes.

Animals

Complete nucleotide sequence of the rabbit beta-like globin gene cluster. Analysis of intergenic sequences and comparison with the human beta-like globin gene cluster.

The nucleotide sequence of the entire beta-like globin gene cluster of rabbits has been determined. This sequence of a continuous stretch of 44.5 x 10(3) base-pairs (bp) starts about 6 x 10(3) bp upstream from epsilon (the 5'-most gene) and ends about 12 x 10(3) bp downstream from beta (the 3'-most gene). Analysis of the sequence reveals that: (1) the sequence is relatively A + T rich (about 60%); (2) regions with high G + C content are associated with OcC repeats, a short interspersed repeated DNA in rabbits; (3) the distribution of polypurines, polypyrimidines and alternating purine/pyrimidine tracts is not random within the cluster; (4) most open reading frames are associated with known globin coding regions, OcC repeats or long interspersed repeats (L1 repeats); (5) the most prominent open reading frames are found in the L1 repeats; (6) different strand asymmetries in base composition are associated with embyronic and adult genes as well as the tandem L1 repeats at the 3' end of the cluster; and (7) essentially all the repeats appear to have been inserted by a transposon mechanism. A comparison of the sequence with itself by a dot-plot analysis has revealed nine new members of the OcC family of repeats in addition to the six previously reported. The OcC repeats tend to be clustered, particularly in the epsilon-gamma and gamma-psi delta intergenic regions. Dot-plot comparisons between the rabbit and the human clusters have revealed extensive sequence matches. Homology starts about 6 x 10(3) bp 5' to epsilon or as far upstream as the rabbit sequence is available. It continues throughout the entire cluster and stops about 0.7 x 10(3) bp 3' to beta, at which point several repeats have inserted in both rabbits and humans. Throughout the gene cluster, the homology is interrupted mainly by insertions or deletions in either the rabbit or the human genome. Almost all of the insertions are of known short or long repeated DNAs. The positions of the insertions are different in the two gene clusters, which indicates that both short and long repeats have been transposing throughout the genome for the time since the mammalian radiation. An alignment of rabbit and human sequences allows the calculation of the substitution rate around epsilon. Sequences far removed from the gene are evolving at a rate equivalent to the pseudogene rate, although some short regions show an apparently higher rate.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals

The L1 family of long interspersed repetitive DNA in rabbits: sequence, copy number, conserved open reading frames, and similarity to keratin.

The L1 family of long interspersed repetitive DNA in the rabbit genome (L1Oc) has been studied by determining the sequence of the five L1 repeats in the rabbit beta-like globin gene cluster and by hybridization analysis of other L1 repeats in the genome. L1Oc repeats have a common 3' end that terminates in a poly A addition signal and an A-rich tract, but individual repeats have different 5' ends, indicating a polar truncation from the 5' end during their synthesis or propagation. As a result of the polar truncations, the 5' end of L1Oc is present in about 11,000 copies per haploid genome, whereas the 3' end is present in at least 66,000 copies per haploid genome. One type of L1Oc repeat has internal direct repeats of 78 bp in the 3' untranslated region, whereas other L1Oc repeats have only one copy of this sequence. The longest repeat sequenced, L1Oc5, is 6.5 kb long, and genomic blot-hybridization data using probes from the 5' end of L1Oc5 indicate that a full length L1Oc repeat is about 7.5 kb long, extending about 1 kb 5' to the sequenced region. The L1Oc5 sequence has long open reading frames (ORFs) that correspond to ORF-1 and ORF-2 described in the mouse L1 sequence. In contrast to the overlapping reading frames seen for mouse L1, ORF-1 and ORF-2 are in the same reading frame in rabbit and human L1s, resulting in a discistronic structure. The region between the likely stop codon for ORF-1 and the proposed start codon for ORF-2 is not conserved in interspecies comparisons, which is further evidence that this short region does not encode part of a protein. ORF-1 appears to be a hybrid of sequences, of which the 3' half is unique to and conserved in mammalian L1 repeats. The 5' half of ORF-1 is not conserved between mammalian L1 repeats, but this segment of L1Oc is related significantly to type II cytoskeletal keratin.

Animals

The linkage arrangement of four rabbit beta-like globin genes.

Four different regions of rabbit beta-like globin gene sequences designated beta 1, beta 2, beta 3 and beta 4 were identified in a set of clones isolated from a bacteriophage lambda library of chromosomal DNA fragments (Maniatis et al., 1978). Restriction mapping and blot hybridization (Southern, 1975) studies indicate that a subset of these clones containing beta 1 and beta 2 hybridizes to an adult beta-globin cDNA clone (Maniatis et al., 1976) more efficiently than to a human gamma-globin cDNA clone (Wilson et al., 1978), while another subset containing beta 3 and beta 4 displays the converse hybridization specificity. beta 1 was identified as the adult beta-globin gene, while beta 2, beta 3 and beta 4 have not been identified with any known rabbit globin polypeptides. Cross-hybridization and transcriptional orientation experiments indicate that the set of beta-like gene clones contains overlapping restriction fragments encompassing 44 kb of rabbit chromosomal DNA. In addition, all four genes have the same transcriptional orientation and are arranged in the order 5'-beta 4-beta 3-beta 2-beta 1-3'.

Animals

The structure and transcription of four linked rabbit beta-like globin genes.

Rabbit chromosomal DNA contains a cluster of four linked beta-like globin genes arranged in the orientation 5'-beta 4-(8kb)-beta 3-(5 kb)-beta 2-(7-kb)-beta 1-3'. Determination of the nucleotide sequence of gene beta 1 confirms that this gene corresponds to the second type of two common co-dominant alleles encoding the adult beta-globin chain. With the exception of two nucleotide substitutions in the large intervening sequence (intron), the intron and flanking sequences are identical with the nucleotide sequence of the first type determined by Weissmann et al. (1979). A 14S polyadenylated transcript containing large intron sequences (possibly a mRNA precursor) is detected in the bone marrow cells of anemic rabbits. Gene beta 2 has limited sequence homology to adult and embryonic beta-globin probes and lacks a detectable mRNA transcript in the erythropoietic tissues examined. It contains at least one intervening sequence analogous to the large intron in gene beta 1. Genes beta 3 and beta 4 both contain an intron of 0.8 kb. Partial DNA sequence analysis indicates that the large intron in beta 4 is located between codons for amino acids lysine and leucine in an analogous position to that of the large intron in beta 1. In addition, a second smaller intron interrupts the 5' coding sequences of gene beta 4. Both genes beta 3 and beta 4 are transcribed in embryonic globin-producing cells. Their DNA sequence homology is limited, however, to a segment of approximately 0.2 kb located on the 5' side of the large intron.

Animals

The isolation of structural genes from libraries of eucaryotic DNA.

We present a procedure for eucaryotic structural gene isolation which involves the construction and screening of cloned libraries of genomic DNA. Large random DNA fragments are joined to phage lambda vectors by using synthetic DNA linkers. The recombinant molecules are packaged into viable phage particles in vitro and amplified to establish a permanent library. We isolated structural genes together with their associated sequences from three libraries constructed from Drosophila, silkmoth and rabbit genomic DNA. In particular, we obtained a large number of phage recombinants bearing the chorion gene sequence from the silkmoth library and several independent clones of beta-globin genes from the rabbit library. Restriction mapping and hybridization studies reveal the presence of closely linked beta-globin genes.

Base Sequence

Histone neighbors in nuclei and extended chromatin.

Histone neighbors in compact and extended chromatin have been investigated by cross-linking histones in nuclei and in nucleohistone extended with 6 M urea, using the bifunctional reversible reagent methyl-4-mercaptobutyrimidate (MMB). Similar histone dimers are found in both conformational states of chromatin. The dimers most frequently found are H2b-H2a, H2b-H3 and H3-H2a; dimers found less frequently are H3-H4, H3-H3 and H2b-H4. More H3-H3 is found in nuclei than in extended chromatin. H1 is found predominantly as poly-H1, although it can be cross-linked to H2b or H3. After reaction with MMB, native compact chromatin is no longer extendable in 6 M urea, which shows that the reagent is capable of linking together histones holding the chromatin in a compact conformation. Thus the histone propinquity in extended chromatin mimics and intimate histone associations in compact chromatin.

Cell Nucleus

An approach to histone nearest neighbours in extended chromatin.

The primary sequence organization of histones upon the DNA molecule in chromatin has been analyzed by extension of the nucleoprotein at very low ionic strength and crosslinking with a reversible crosslinking reagent, methyl-4-mercaptobutyrimidate. Histones extracted after limited reaction were fractionated into different classes and the composition of the oligomers analyzed after reduction of the crosslinked material. We have found that the following dimers occur at a high frequency: (F3-F2b), (F3-F2a2), and (F2b-F2a2), whereas (F2b-F2al), (F3-F2al) and (F3-F3) occur with a lower frequency. F1 appears to polymerize rapidly to largely homogeneous polymers of high molecular weight. These results are analyzed in terms of several models proposed for chromatin structure.

Animals