PubMed HealthSearch

Biomedical subjects

R Mott

Publications and source records attributed to R Mott.

At least 19 recordsLinked to original sources

Approximate statistics of gapped alignments.

A heuristic approximation to the score distribution of gapped alignments in the logarithmic domain is presented. The method applies to comparisons between random, unrelated protein sequences, using standard score matrices and arbitrary gap penalties. It is shown that gapped alignment behavior is essentially governed by a single parameter, alpha, depending on the penalty scheme and sequence composition. This treatment also predicts the position of the transition point between logarithmic and linear behavior. The approximation is tested by simulation and shown to be accurate over a range of commonly used substitution matrices and gap-penalties.

Amino Acid Sequence

Local sequence alignments with monotonic gap penalties.

MOTIVATION: Sequence alignments obtained using affine gap penalties are not always biologically correct, because the insertion of long gaps is over-penalised. There is a need for an efficient algorithm which can find local alignments using non-linear gap penalties. RESULTS: A dynamic programming algorithm is described which computes optimal local sequence alignments for arbitrary, monotonically increasing gap penalties, i.e. where the cost g(k) of inserting a gap of k symbols is such that g(k) >/= g(k-1). The running time of the algorithm is dependent on the scoring scheme; if the expected score of an alignment between random, unrelated sequences of lengths m, n is proportional to log mn, then with one exception, the algorithm has expected running time O(mn). Elsewhere, the running time is no greater than O(mn(m+n)). Optimisations are described which appear to reduce the worst-case run-time to O(mn) in many cases. We show how using a non-affine gap penalty can dramatically increase the probability of detecting a similarity containing a long gap. AVAILABILITY: The source code is available to academic collaborators under licence.

Algorithms

Comparative gene expression profiling by oligonucleotide fingerprinting.

The use of hybridisation of synthetic oligonucleotides to cDNAs under high stringency to characterise gene sequences has been demonstrated by a number of groups. We have used two cDNA libraries of 9 and 12 day mouse embryos (24 133 and 34 783 clones respectively) in a pilot study to characterise expressed genes by hybridisation with 110 hybridisation probes. We have identified 33 369 clusters of cDNA clones, that ranged in representation from 1 to 487 copies (0.7%). 737 were assigned to known rodent genes, and a further 13 845 showed significant homologies. A total of 404 clusters were identified as significantly differentially represented (P < 0.01) between the two cDNA libraries. This study demonstrates the utility of the fingerprinting approach for the generation of comparative gene expression profiles through the analysis of cDNAs derived from different biological materials.

Animals

Trace alignment and some of its applications.

MOTIVATION: Extra useful information can be extracted from a DNA chromatogram trace, over that contained in the base-called DNA sequence. Many sequencing applications can benefit from examination of these traces. RESULTS: An algorithm, based on dynamic programming, for aligning a DNA chromatogram to a DNA sequence is described and implemented. Its applications to vector clipping, EST alignment and mutation detection are discussed.

Algorithms

Sequence assembly with CAFTOOLS.

Large-scale genomic sequencing requires a software infrastructure to support and integrate applications that are not directly compatible. We describe a suite of software tools built around the Common Assembly Format (CAF), a comprehensive representation of a sequence assembly as a text file. These tools form the backbone of sequencing informatics at the Sanger Centre and the Genome Sequencing Center. The CAF format is intentionally flexible, and our Perl and C libraries, which parse and manipulate it, provide powerful tools for creating new applications as well as wrappers to incorporate other software. The tools are available free by anonymous FTP from ftp://ftp.sanger.ac.uk/pub/badger/.

Algorithms

Instability of highly expanded CAG repeats in mice transgenic for the Huntington's disease mutation.

Six inherited neurodegenerative diseases are caused by a CAG/polyglutamine expansion, including spinal and bulbar muscular atrophy (SBMA), Huntington's disease (HD), spinocerebellar ataxia type 1 (SCA1), dentatorubral pallidoluysian atrophy (DRPLA) Machado-Joseph disease (MJD or SCA3) and SCA2. Normal and expanded HD allele sizes of 6-39 and 35-121 repeats have been reported, and the allele distributions for the other diseases are comparable. Intergenerational instability has been described in all cases, and repeats tend to be more unstable on paternal transmission. This may present as larger increases on paternal inheritance as in HD, or as a tendency to increase on male and decrease on female transmission as in SCA1 (ref. 15). Somatic repeat instability is also apparent and appears most pronounced in the CNS. The major exception is the cerebellum, which in HD, DRPLA, SCA1 and MJD has a smaller repeat relative to the other brain regions tested. Of non-CNS tissues, instability was observed in blood, liver, kidney and colon. A mouse model of CAG repeat instability would be helpful in unravelling its molecular basis although an absence of CAG repeat instability in transgenic mice has so far been reported. These studies include (CAG) in the androgen receptor cDNA, (CAG) in the HD cDNA, (CAG) in the SCA1 cDNA, (CAG) in the SCA3 cDNA and as an isolated (CAG) tract.

Animals

FPC: a system for building contigs from restriction fingerprinted clones.

MOTIVATION: To meet the demands of large-scale sequencing, thousands of clones must be fingerprinted and assembled into contigs. To determine the order of clones, a typical experiment is to digest the clones with one or more restriction enzymes and measure the resulting fragments. The probability of two clones overlapping is based on the similarity of their fragments. A contig contains two or more overlapping clones and a minimal tiling path of clones is selected to be sequenced. Interactive software with algorithmic support is necessary to assemble the clones into contigs quickly. RESULTS: FPC (fingerprinted contigs) is an interactive program for building contigs from restriction fingerprinted clones. FPC uses an algorithm to cluster clones into contigs based on their probability of coincidence score. For each contig, it builds a consensus band (CB) map which is similar to a restriction map; but it does not try to resolve all the errors. The CB map is used to assign coordinates to the clones based on their alignment to the map and to provide a detailed visualization of the clone overlap. FPC has editing facilities for the user to refine the coordinates and to remove poorly fingerprinted clones. Functions are available for updating an FPC database with new clones. Contigs can easily be merged, split or deleted. Markers can be added to clones and are displayed with the appropriate contig. Sequence-ready clones can be selected and their sequencing status displayed. As such, FPC is an integrated program for the assembly of sequence-ready clones for large-scale sequencing projects.

Algorithms

Cloning and expression of cystolic phospholipase A2 (cPLA2) and a naturally occurring variant. Phosphorylation of Ser505 of recombinant cPLA2 by p42 mitogen-activated protein kinase results in an increase in specific activity.

Full-length cytosolic phospholipase A2 (cPLA2) was cloned from U937 cells and polymorphonuclear leukocytes (PMNLs) while a naturally occurring variant of cPLA2, which lacks residues Val473-Ala749 but has a C-terminal extension of ILMNLSEYMLWMSKVKRFM (DcPLA2) was cloned from PMNLs and mononuclear leukocytes. We were unable to clone DcPLA2 from U937 cells. When cPLA2 and DcPLA2 were expressed in insect cells, both proteins were detected in cell lysates by SDS/PAGE as single bands of apparent molecular masses 100 kDa and 57 kDa, respectively. Full-length cPLA2 was phosphorylated stoichiometrically by p42 mitogen-activated protein (MAP) kinase in vitro at a similar rate to other physiological substrates of this protein kinase and the major site of phosphorylation was identified by amino acid sequencing as Ser505. [32P]Ser(P)505 in cPLA2 was only dephosphorylated at a slow rate by mammalian tissue homogenates. Protein phosphatases 2A, 2B and 2C all contributed significantly to the overall dephosphorylation of cPLA2. The phosphorylation of cPLA2 by p42 MAP kinase correlated with an approximately 1.5-fold increase in specific enzyme activity which was reversed by dephosphorylation.

Amino Acid Sequence

An integrated YAC map of the human X chromosome.

The human X chromosome is associated with a large number of disease phenotypes, principally because of its unique mode of inheritance that tends to reveal all recessive disorders in males. With the longer term goal of identifying and characterizing most of these genes, we have adopted a chromosome-wide strategy to establish a YAC contig map. We have performed > 3250 inter Alu-PCR product hybridizations to identify overlaps between YAC clones. Positional information associated with many of these YAC clones has been derived from our Reference Library Database and a variety of other public sources. We have constructed a YAC contig map of the X chromosome covering 125 Mb of DNA in 25 contigs and containing 906 YAC clones. These contigs have been verified extensively by FISH and by gel and hybridization fingerprinting techniques. This independently derived map exceeds the coverage of recently reported X chromosome maps built as part of whole-genome YAC maps.

Chromosome Mapping

Construction of genetic maps using distance geometry.

The techniques of distance geometry, which generate coordinates from observed interpoint distances, have been applied to the problem of determining the relative positions of linked genetic loci from observed interlocus distances. Only the most precise data needed to join the loci are used, with missing distances substituted by sums of precise intermediate distances. Good initial positions (and therefore the order) of loci on a linear map are obtained in an operation of complexity O(N3). The method can therefore be used to generate good initial framework maps for the large numbers of markers encountered in current mapping projects. The locus positions can be subsequently refined to maximize the agreement with the originally observed distances, taking account of the weights of individual interlocus distances. By choosing only small distances from which to construct the map, the method reduces any error due to an incorrect choice of mapping function. It also prevents undue expansion of the map due to error-prone markers, since such markers are accommodated in higher dimensions. The method estimates the error in the positions of individual markers on the final map and identifies well- and ill-defined regions of the map.

Algorithms

Efficient high-resolution genetic mapping of mouse interspersed repetitive sequence PCR products, toward integrated genetic and physical mapping of the mouse genome.

The ability to carry out high-resolution genetic mapping at high throughput in the mouse is a critical rate-limiting step in the generation of genetically anchored contigs in physical mapping projects and the mapping of genetic loci for complex traits. To address this need, we have developed an efficient, high-resolution, large-scale genome mapping system. This system is based on the identification of polymorphic DNA sites between mouse strains by using interspersed repetitive sequence (IRS) PCR. Individual cloned IRS PCR products are hybridized to a DNA array of IRS PCR products derived from the DNA of individual mice segregating DNA sequences from the two parent strains. Since gel electrophoresis is not required, large numbers of samples can be genotyped in parallel. By using this approach, we have mapped > 450 polymorphic probes with filters containing the DNA of up to 517 backcross mice, potentially allowing resolution of 0.14 centimorgan. This approach also carries the potential for a high degree of efficiency in the integration of physical and genetic maps, since pooled DNAs representing libraries of yeast artificial chromosomes or other physical representations of the mouse genome can be addressed by hybridization of filter representations of the IRS PCR products of such libraries.

Animals

The properties of a cloned human high-molecular-mass cytosolic phospholipase A2 investigated using a continuous fluorescence displacement assay: evidence for enzyme clustering on phospholipid vesicles.

The 85 kDa human cytosolic phospholipase A2 has been cloned and expressed in insect Sf21 cells. The pure enzyme has been investigated using a fluorescence displacement assay that provides a continuous record of phospholipid hydrolysis [Wilton (1990) Biochem. J. 266, 435-439]. The unusual kinetic properties of this enzyme, previously described using radioactive assays, were readily demonstrated using the continuous fluorescence assay and were examined in detail. It is proposed that the enzyme clusters on the surface of a fixed number of substrate vesicles during the initial stages of catalysis and that the characteristic burst phase of hydrolysis represents the hydrolysis of these vesicles. This clustering produced a molar ratio of total phospholipid substrate to enzyme of about 450:1 at vesicle saturation with enzyme. Under limiting substrate conditions, the lower secondary rate that is observed results eventually in almost complete hydrolysis of the phospholipid; this was confirmed using radioactive substrate. Evidence is presented that during the initial burst phase, equivalent to hydrolysis of the outer monolayer of the vesicle, the enzyme remains tightly bound but is released as the reaction proceeds towards complete hydrolysis of the phospholipid substrate. In the presence of excess substrate, about 370 mol of fatty acid are released per mol of enzyme during the burst phase and it is calculated that this value also approximates to hydrolysis of the outer monolayer of the vesicle. It is proposed that the formation of a stable enzyme-vesicle complex during the burst phase of phospholipid hydrolysis may be due, at least in part, to protein-protein interactions between adjacent enzyme molecules in order to account for the clustering phenomenon.

Animals

Model for a transcript map of human chromosome 21: isolation of new coding sequences from exon and enriched cDNA libraries.

The construction of a transcriptional map for human chromosome 21 requires the generation of a specific catalogue of genes, together with corresponding mapping information. Towards this goal, we conducted a pilot study on a pool of random chromosome 21 cosmids representing 2 Mb of non-contiguous DNA. Exon-amplification and cDNA selection methods were used in combination to extract the coding content from these cosmids, and to derive expressed sequences libraries. These libraries and the source cosmid library were arrayed at high density for hybridisation screening. A strategy was used which related data obtained by multiple hybridisations of clones originating from one library, screened against the other libraries. In this way, it was possible to integrate the information with the physical map and to compare the gene recovery rate of each technique. cDNAs and exons were grouped into bins delineated by EcoRI cosmid fragments, and a subset of 91 cDNAs and 29 exons have been sequenced. These sequences defined 79 non-overlapping potential coding segments distributed in 24 transcriptional units, which were mapped along 21q. Northern blot analysis performed for a subset of cDNAs indicated the existence of a cognate transcript. Comparison to databases indicated three segments matching to known chromosome 21 genes: PFKL, COL6A1 and S100B and six segments matching to unmapped anonymous expressed sequence tags (ESTs). At the translated nucleotide level, strong homologies to known proteins were found with ATP-binding transporters of the ABC family and the dihydroorotase domain of pyrimidine synthetases. These data strongly suggest that bona fide partial genes have been isolated. Several of the newly isolated transcriptional units map to clinically important regions, in particular those involved in Down's syndrome, progressive myoclonus epilepsia and auto-immune polyglandular disease. The study presented here illustrates the complementarity of exon-amplification and cDNA selection techniques for generating a large resource of new expressed landmarks, which contribute to the construction of a chromosome 21 transcript map.

Chromosome Mapping

An algorithm to detect chimeric clones and random noise in genomic mapping.

Experimental noise and noncontiguous clone inserts can pose serious problems in reconstructing genomic maps from hybridization data. We describe an algorithm that easily identifies false positive signals and clones containing chimeric inserts/internal deletions. The algorithm "dechimerizes" clones, splitting them into independent contiguous components and cleaning the initial library into a more consistent data set for further ordering. The effectiveness of the algorithm is demonstrated on both simulated data and the real YAC map of the whole genome of the fission yeast Schizosaccharomyces pombe.

Algorithms

Distribution of trinucleotide repeat sequences across a 2 Mbp region containing the Huntington's disease gene.

The recent observation that the mutation underlying a number of genetic diseases including fragile sites, FRAXA and FRAXE (associated with mental retardation), myotonic dystrophy, spinal and bulbar muscular atrophy (Kennedy's disease), Huntington's disease and spinocerebellar ataxia type 1 are caused by the expansion of a trinucleotide repeat sequence will lead to interest in the identification of such sequences in regions related to other diseases. We report here the identification of all ten classes of trinucleotide repeats within a 2 Mbp region of 4p16.3 containing the Huntington's disease (HD) gene. Fifty one triplet repeats were identified and localised on a high resolution restriction map of a cosmid contig covering this region. This included the triplet repeat (CAG)n, which has subsequently been shown to be expanded in Huntington's disease patients.

Base Sequence

An integrated YAC-overlap and 'cosmid-pocket' map of the human chromosome 21.

We describe here the construction of an ordered clone map of human chromosome 21, based on the identification of ordered sets of YAC clones covering > 90% of the chromosome, and their use to identify groups of cosmid clones (cosmid pockets) localised to subregions defined by the YAC clone map. This is to our knowledge the highest resolution map of one human chromosome to date, localising 530 YAC clones covering both arms of the chromosome, spanning > 36 Mbp, and localising more than 6300 cosmids to 145 intervals on both arms of the chromosome. The YAC contigs have been formed by hybridising a 6.1 equivalents chromosome 21 enriched YAC collection displayed on arrayed nylon membranes to a series of 115 DNA markers and Alu-PCR products from YACs. Forty eight mega-YACs from the previously published CEPH-Genethon map of sequence tagged sites (STS) have also been included in the contig building experiments. A YAC tiling path was then size-measured and confirmed by gel-fingerprinting. A minimal tiling path of 70 YACs were then used as probes against the 7.5 genome equivalents flow sorted chromosome 21 cosmid library in order to identify the lists of cosmids mapping to alternating shared--non-shared intervals between overlapping YACs ('cosmid pockets'). For approximately 1/5 of the minimal tiling path of YACs, locations and non-chimaerism have been confirmed by fluorescence in situ hybridisation (FISH), and approximately 1/5 of all cosmid pocket assignments have independent, confirmatory marker hybridizations in the ICRF cosmid reference library system. We also demonstrate that 'pockets' contain overlapping sets of cosmids (cosmid contigs). In addition to being an important logical intermediate step between the YAC maps published so far and a future map of completely ordered cosmids, this map provides immediately available low-complexity cosmid material for high resolution FISH mapping of chromosomal aberrations on interphase nuclei, and for rapid positional isolation of transcripts in the highly resolved regions of genetic interest.

Chromosome Mapping

Algorithms and software tools for ordering clone libraries: application to the mapping of the genome of Schizosaccharomyces pombe.

A complete set of software tools to aid the physical mapping of a genome has been developed and successfully applied to the genomic mapping of the fission yeast Schizosaccharomyces pombe. Two approaches were used for ordering single-copy hybridisation probes: one was based on the simulated annealing algorithm to order all probes, and another on inferring the minimum-spanning subset of the probes using a heuristic filtering procedure. Both algorithms produced almost identical maps, with minor differences in the order of repetitive probes and those having identical hybridisation patterns. A separate algorithm fitted the clones to the established probe order. Approaches for handling experimental noise and repetitive elements are discussed. In addition to these programs and the database management software, tools for visualizing and editing the data are described. The issues of combining the information from different libraries are addressed. Also, ways of handling multiple-copy probes and non-hybridisation data are discussed.

Algorithms