PubMed HealthSearch

Biomedical subjects

P A Stockwell

Publications and source records attributed to P A Stockwell.

15 recordsLinked to original sources

Phylogenetic relationships among transposon-like elements in human and primate DNA.

A THE-1 sequence in intron 7 of the human dystrophin gene has been found to represent a new subfamily of THE-1 elements. The sequence is closely related to the MstII family of repetitive sequences and is more like single-copy sequences found in the galago genome than any other THE-1 sequence previously reported. This new THE-1 sequence has been compared with two other complete THE-1 sequences and three related long-terminal repeat elements that we have previously found in intron 7 of the dystrophin gene, and with members of the same family from elsewhere in the primate genome. Parsimony and deletion analysis show that the cluster of THE-1 sequences in intron 7 of the dystrophin gene has arisen from at least three individual insertion events, rather than from the insertion and duplication of a single progenitor sequence.

Animals

Sequence of a transposon identified as Tn1000 (gamma delta).

We report the complete sequence of a transposon found in a cosmid clone of a human DNA sequence. The transposon is identified as the Escherichia coli transposon Tn1000 (also known as gamma delta) on the basis of the identity of the restriction map of the new sequence with that previously recorded for Tn1000 and homology between parts of the new sequence and that of published fragments of Tn1000 sequence. The transposon, which comprises 5,981 nucleotides including two 35 bp inverted terminal repeat sequences (ITRs), contains three open reading frames. The sequence of the resolvase coding region (tnpR) is identical to that published by others. A second reading frame can be identified as the tnpA gene, coding for the transposase, on the grounds of its strong homology with the corresponding gene from transposon Tn3. The third reading frame has the potential to code for a protein of unknown function containing 698 amino acids.

Amino Acid Sequence

The translational termination signal database (TransTerm) now also includes initiation contexts.

The TransTerm database of termination codon contexts has been extended to include sense codon usage, and initiation codon contexts. The database was constructed from 23,721 coding sequences from 93 organisms. The database contains: a) the sequence around the termination codon (-10, +10); b) the sequence around the initiation codon (-20, +10); c) the length, 'G+C%' of the third position of codons (GC3), the 'codon adaptation index' (CAI) and the 'effective number of codons' statistic (Nc); d) summary tables for each organism including total codon usage, stop codon and tetranucleotide stop-signal usage, and matrices tallying base frequencies at each position around the initiation and termination codons. The data are arranged to facilitate investigation of the relationships between the three phases of protein synthesis. The database is available electronically from EMBL.

Animals

The putative single-stranded DNA-binding protein of the filamentous bacteriophage, Ifl. Amino acid sequence of the protein and structure of the gene.

The protein product corresponding to the gene located in the region of the coliphage Ifl genome shown to contain the code for the single-stranded DNA (ssDNA)-binding proteins of all filamentous phages so far studied has been isolated from infected bacterial cells and its amino acid sequence determined. The mature protein contains 95 amino acids (calculated molecular mass 10553 Da). Its sequence corresponds to that predicted from the DNA sequence but lacks the initiating methionine residue. Although there is little direct sequence homology between the phage Ifl protein and the ssDNA-binding proteins of the other filamentous phages that have been studied, computer-based comparisons of various physical and structural parameters showed that the phage Ifl protein contains a domain that is closely related to domains in the coliphage T4 gene 32 protein and the Pseudomonas phage Pfl ssDNA-binding protein and suggest that the Ifl protein does have a ssDNA-binding function although we were unable to show this directly.

Amino Acid Sequence

Sequence analysis suggests that tetra-nucleotides signal the termination of protein synthesis in eukaryotes.

An increasing number of cases where tri-nucleotide stop codons do not signal the termination of protein synthesis are being reported. In order to identify what constitutes an efficient stop signal, we analysed the region around natural stop codons in genes from a wide variety of eukaryotic species and gene families. Certain stop codons and nucleotides following stop codons are over-represented, and this pattern is accentuated in highly expressed genes. For example, the preferred signal for Saccharomyces cerevisiae and Drosophila melanogaster highly expressed genes is UAAG, and generally the signals UAA(A/G) and UGA(A/G) are preferred in eukaryotes. The GC% of the organism or DNA region can affect whether there is A or G in the second or fourth positions. We suggest therefore, that the stop codon and the nucleotide following it comprise a tetra-nucleotide stop signal. A model is proposed in which the polypeptide chain release factor, a protein, recognises this sequence, but will tolerate some substitution, particularly A to G in the second or third positions.

Animals

The signal for the termination of protein synthesis in procaryotes.

The sequences around the stop codons of 862 Escherichia coli genes have been analysed to identify any additional features which contribute to the signal for the termination of protein synthesis. Highly significant deviations from the expected nucleotide distribution were observed, both before and after the stop codon. Immediately prior to UAA stop codons in E. coli there is a preference for codons of the form NAR (any base, adenine, purine), and in particular those that code for glutamine or the basic amino acids. In contrast, codons for threonine or branched nonpolar amino acids were under-represented. Uridine was over-represented in the nucleotide position immediately following all three stop codons, whereas adenine and cytosine were under-represented. This pattern is accentuated in highly expressed genes, but is not as marked in either lowly expressed genes or those that terminate in UAG, the codon specifically recognised by polypeptide chain release factor-1. These observations suggest that for the efficient termination of protein synthesis in E. coli, the 'stop signal' may be a tetranucleotide, rather than simply a tri-nucleotide codon, and that polypeptide chain release factor-2 recognises this extended signal. The sequence following stop codons was analysed in genes from several other procaryotes and bacteriophages. Salmonella typhimurium, Bacillus subtilis, bacteriophages and the methanogenic archaebacteria showed a similar bias to E. coli.

Amino Acids

Laser excitation of fluorescent-labeled polypeptides in polyacrylamide gels.

A laser beam at 488 nm, converted into a fan of light by a surface-coated mirror oscillated in response to a triangular wave, was inserted into the base of a polyacrylamide gel. The laser light was trapped by internal reflection and gave uniform illumination throughout the entire gel slab. Photography with color film detected 50 fmol of fluorescein covalently coupled to ovalbumin, gave 80-fold greater sensitivity than transillumination in detection of fluorescein-labeled polypeptides, and was about 25-fold more sensitive than protein staining with silver. Laser illumination visualized end-labeled beta-galactosidase, afforded quality control of such preparations, and demonstrated that the end-labeled derivative contained about 25-fold less fluorescein than uniformly labeled beta-galactosidase. The latter result was confirmed by dot-blot analysis using a polyclonal antibody specific for fluorescein. The application of end-labeling to the location of features of protein primary structure is discussed.

Electrophoresis, Polyacrylamide Gel

A homologue of retroviral pseudoproteases in the parapoxvirus, orf virus.

The nucleotide sequence of a near-terminal region of orf virus DNA was determined. Examination of the sequence revealed an open reading frame encoding a peptide with significant amino acid homology to the pseudoprotease domains recently identified in a number of retroviruses including mouse mammary tumor virus, simian Mason-Pfizer virus, maedi-visna virus, and equine infectious anaemia virus. The orf virus pseudoprotease shares up to 28% amino acid homology with retroviral pseudoproteases and appears to be a discrete transcriptional unit rather than a subunit of a larger polypeptide as is the case in retroviruses. The sharing of amino acid composition across such wide taxonomic boundaries suggests that this polypeptide has a functional significance in both retroviruses and poxviruses.

Animals

Messenger RNA recognition in Escherichia coli: a possible second site of interaction with 16S ribosomal RNA.

Examination of the nucleotides following the ATG or GTG initiation codons of a file of 251 genes from Escherichia coli has shown that 247 (98.4%) of them contain a sequence of at least three and 168 (66.9%) of them a sequence of at least four consecutive nucleotides that is complementary to some part of the 16 nt at the 5' terminus of the bacterial 16S rRNA. It is proposed that this sequence, which falls within the first 24 nt coding for the genetic message, might be involved in mRNA recognition through a mechanism analogous to the well-established 'Shine--Dalgarno' interaction with the 3' terminus of the 16S rRNA. Comparison of these data with data derived from a file of 117 'false' gene starts that have a Shine--Dalgarno-like sequence followed by a suitably spaced ATG or GTG triplet but which are believed not to lie at the beginnings of genetic messages shows the association that we have found to be statistically significant at the 99.9% level.

Escherichia coli

HOMED: a homologous sequence editor.

The alignment of homologous sequences with each other and their display has proved a difficult task, despite a frequent requirement for this process. HOMED enables related sequences to be edited and listed in parallel with each other. The editor function uses a full screen editor which emulates the text editors KED and EDT (on PDP-11 and VAX-11 respectively) and which can be adapted to emulate other text editors. This emulation has been adopted to simplify user learning of editing functions. HOMED provides functions for listing the sequences in a variety of formats and for generating a consensus sequence as well as providing a series of tools for maintenance of the sequence database. HOMED has been implemented in Pascal in a modular fashion to enhance portability.

Algorithms

VTUTIN: a full screen gel management editor.

Large DNA sequences are now routinely sequenced by the cloning of randomly generated fragments into single-stranded DNA phage vectors (the 'shotgun' method). Various programs exist for computerized assembly of such fragments, including the phases of data entry, homology searching and gel-management/editing. Many gel-management editors are rudimentary in nature, using either line-editing techniques or using unnatural displays or command systems. Others are available only on restricted types of computer system. The program VTUTIN makes full screen editing along the lines of modern text editors available for the complex data type of sets of sequence gels and their consensus. Not only are the data displayed on the VDU screen in a natural manner, but VTUTIN has also been written to model the command system of a well-established text editor (PDP-11 KED or VAX/VMS EDT) to simplify editor use and learning. VTUTIN has been written in Pascal in a modular form so that wide-spread portability is facilitated. VTUTIN is currently implemented to work on VT-100 type terminals although the modularity of the code should allow straightforward conversion for other terminal types and should also permit simple alteration to model any other text editor.

Algorithms

A large database DNA sequence handling program with generalized searching specifications.

The program described allows for the creation and manipulation of files of DNA sequence data up to very great lengths. The program uses its own paging system to load segments of the sequence into a small internal buffer so that the program does not have excessive memory requirements. The program offers a menu of functions to the user, and has been written to be forgiving of user errors. A code for the generalised specification of bases as a series of groups (i.e. A or T, Purine, etc.) has been devised and can be used in search specifications or in sequence files. Versions of the program have been developed to run with special efficiency under DIGITAL's RT11 operating system or to run under systems with a suitable implementation of FORTRAN VI.

Base Sequence

Gene duplication in tetraploid fish: model for gene silencing at unlinked duplicated loci.

Several groups of fishes, including salmonids and catastomids, appear to have originated through genome duplication events. However, these two groups retain approximately 50% of the loci examined as functioning duplicates, despite the passage of 50 million years or more of mutation and selection. Although other effects are not excluded, this apparently slow rate of duplicate silencing can be explained in terms of the effects of selection against defective double homozygotes to unlinked duplicates. We have derived a computer simulation of genetic drift that affords direct evaluation of the effects of population size (N), mutation rate (micron), initial allele frequencies, back mutation, fitness, and time on the probability of fixation for null alleles at unlinked duplicate loci. The results show that this probability is approximately linearly related to population size for N greater than or equal to 10(3). Specifically, for naive populations, the time for 50% probability of gene silencing is approximately equal to 15N + micron-3/4 generations. The retention of 50% of the loci as functional duplicates may therefore result from the large effective size of salmonid and catastomid populations. The results also show that, under most conditions for populations of 2000--3000 or larger, unlinked duplicate loci will be sustained in the functional state longer than tandem (linked) duplicates and hence are available for evolution of new functions for a longer time.

Alleles