PubMed HealthSearch

Biomedical subjects

R Staden

Publications and source records attributed to R Staden.

18 recordsLinked to original sources

The C. elegans genome sequencing project: a beginning.

The long-term goal of this project is the elucidation of the complete sequence of the Caenorhabditis elegans genome. During the first year methods have been developed and a strategy implemented that is amenable to large-scale sequencing. The three cosmids sequenced in this initial phase are surprisingly rich in genes, many of which have mammalian homologues.

Animals

Indexing the sequence libraries: software providing a common indexing system for all the standard sequence libraries.

We describe a set of programs for creating and using indexes for the distributed forms of the major sequence libraries. The indexes conform to the specification of those distributed on cd-rom by the EMBL sequence library. The programs create entry name, accession number, author and freetext indexes and a brief directory index. If a suitable application program is given an entry name or accession number these indexes allow rapid retrieval of sequences or annotation. Similarly the author and freetext indexes provide the data for extremely fast searching on author names and "keywords". The indexing programs can create indexes for EMBL, Swiss-Prot, GenBANK, PIR and NRL3d libraries. We also describe the organisation and use of the different sequence libraries and their index files.

Abstracting and Indexing

A standard file format for data from DNA sequencing instruments.

There are now a number of machines for determining DNA sequences. These devices are currently of two types: those such as the Applied Biosystems 373A and the Pharmacia A.L.F. which interpret the sequences of samples as they run on gels within the machine, and those, such as the Bio-Rad and Amersham readers that scan and analyse conventional autoradiographs. Both types of machine can produce their data in the form of traces which represent the band intensity of each of the four base types at each position in the sequence. At present all the machines write files in different formats. We describe a machine independent format for storing data derived from automatic sequencing machines. Files in this format can store the derived sequence, the traces and a set of confidence measures for each base. We have adopted the format as the standard for our sequence handling software.

Databases, Factual

A sequence assembly and editing program for efficient management of large projects.

We describe a sequence assembly and editing program for managing large and small projects. It is being used to sequence complete cosmids and has substantially reduced the time taken to process the data. In addition to handling conventionally derived sequences it can use data obtained from Applied Biosystems,Inc. 373A and Pharmacia A.L.F. fluorescent sequencing machines. Readings are assembled automatically. All editing is performed using a mouse operated contig editor that displays aligned sequences and their traces together on the screen. The editor, which can be used on single contigs or for joining contigs, permits rapid movement along the aligned sequences. Insertions, deletions and replacements can be made in individual aligned readings and global changes can be made by editing the consensus. All changes are recorded. A click on a mouse button will display the traces covering the current cursor position, hence allowing quick resolution of problems. Another function automatically moves the cursor to the next unresolved character. The editor also provides facilities for annotating the sequences. Typical annotations include flagging the positions of primers used for walking, or for marking sites, such as compressions, that have caused problems during sequencing. Graphical displays aid the assessment of progress.

Base Sequence

Screening protein and nucleic acid sequences against libraries of patterns.

We describe programs that can screen nucleic acid and protein sequences against libraries of motifs and patterns. Such comparisons are likely to play an important role in interpreting the function of sequences determined during large scale sequencing projects. In addition we report programs for converting the Prosite protein motif library into a form that is compatible with our searching programs. The programs work on VAX and SUN computers.

Amino Acid Sequence

An improved sequence handling package that runs on the Apple Macintosh.

We report improvements to our sequence analysis package and adaptation to run on the Apple Macintosh range of machines. The 'standard' version of the programs, which run on a VAX, has been given a new user interface that makes the programs very much easier to work with and has facilitated the move to the Macintosh. The reorganization of the code should simplify moves to other systems that offer WIMP user interfaces. In addition to a large number of small but useful extra features, some important new analytical functions have been devised. These include sequence and contig editors; optimal alignment and comparison methods; and a new method for comparing the observed and expected frequencies of selected oligonucleotides.

Base Sequence

Methods for calculating the probabilities of finding patterns in sequences.

This paper describes the use of probability-generating functions for calculating the probabilities of finding motifs in nucleic acid and protein sequences. Equations and algorithms are given for calculating the probabilities associated with nine different ways of defining motifs. Comparisons are made with searches of random sequences. A higher level structure--the pattern--is defined as a list of motifs. A pattern also specifies the permitted ranges of spacing allowed between its constituent motifs. Equations for calculating the expected numbers of matches to patterns are given.

Algorithms

Methods for discovering novel motifs in nucleic acid sequences.

We describe a computer tool to aid the discovery of new motifs in nucleic acid sequences. A typical use would be to analyse a set of upstream regions from a family of related genes in order to find possible control sequences. The heart of the method is the creation of dictionaries of related subsequences. These dictionaries can then be analysed to look for the commonest or best-defined subsequences, those that occur in the highest number of different sequences, or for those in equivalent positions within the family. We show the application of the method to a set of E. coli promoter sequences.

Base Sequence

Deviations from expected frequencies of CpG dinucleotides in herpesvirus DNAs may be diagnostic of differences in the states of their latent genomes.

The DNA sequences of genomes from G + C-rich and A + T-rich lymphotropic herpesviruses [i.e. gammaherpesviruses; Epstein-Barr virus and herpesvirus saimiri (HVS)] are deficient in CpG dinucleotides and contain an excess of TpG and CpA dinucleotides relative to frequencies predicted from their mononucleotide compositions. In contrast, for sequences from genomes of G + C-rich and A + T-rich neurotropic herpesviruses (i.e. alphaherpesviruses; herpes simplex virus and varicellazoster virus) and human cytomegalovirus (HCMV; a betaherpesvirus) the mean observed frequencies of these dinucleotides are close to those expected from their mononucleotide compositions. Comparisons between DNA sequences that encode proteins conserved in all these viruses also show that sequences of these lymphotropic viruses are CpG-deficient whereas the homologous genes from the neurotropic viruses and the HCMV are not. Analyses of local variations in dinucleotide frequencies reveal some occurrences of clustered CpG dinucleotides in generally deficient genomes (e.g. upstream of the thymidylate synthase gene of HVS) and locally CpG-deficient regions within some generally non-deficient genomes (e.g. the major immediate early genes of human, simian and murine CMVs). A relative deficiency in CpG and an excess of TpG and CpA dinucleotides is a diagnostic feature of higher eukaryotic DNA sequences that have been subjected to methylation of cytosine residues in CpG doublets with the resulting increase in mutations to give TpG (and thereby its complement, CpA). The available evidence implicates the latent genome as the site of methylation of these herpesviruses. We conclude that in the neurotropic herpesviruses the normal latent precursors to infectious progeny are not methylated whereas there is local methylation of the immediate early locus in the latent genomes of CMVs, and the latent genomes of these lymphotropic herpesviruses are extensively methylated.

Base Sequence

A strategy of DNA sequencing employing computer programs.

With modern fast sequencing techniques and suitable computer programs it is now possible to sequence whole genomes without the need of restriction maps. This paper describes computer programs that can be used to order both sequence gel readings and clones. A method of coding for uncertainties in gel readings is described. These programs are available on request.

Base Sequence

Nucleotide sequence of bacteriophage G4 DNA.

The 5,577 nucleotide long sequence of bacteriophage G4 DNA has been determined using the 'plus and minus' and chain termination methods of DNA sequencing. This sequence has been compared with that of the closely related bacteriophage phiX174 (refs 1, 55). In the coding regions there is an average of 33.1% nucleotide sequence differences between the two genomes, but the distribution of these changes is not random and the sequence of some genes is more conserved than others. There is less sequence similarity between the untranslated intergenic regions of G4 and phiX174, but despite this the sequences of the J/F, F/G and H/A untranslated spaces in both genomes have similar sized hairpin loops, which may be related to their function.

Bacteriophages

Sequence data handling by computer.

The speed of the new DNA sequencing techniques has created a need for computer programs to handle the data produced. This paper describes simple programs designed specifically for use by people with little or no computer experience. The programs are for use on small computers and provide facilities for storage, editing and analysis of both DNA and amino acid sequences. A magnetic tape containing these programs is available on request.

Amino Acid Sequence