PubMed Health⌕ Search

Biomedical subjects

David M A Martin

Publications and source records attributed to David M A Martin.

8 recordsLinked to original sources

Identification of multiple distinct Snf2 subfamilies with conserved structural motifs.

The Snf2 family of helicase-related proteins includes the catalytic subunits of ATP-dependent chromatin remodelling complexes found in all eukaryotes. These act to regulate the structure and dynamic properties of chromatin and so influence a broad range of nuclear processes. We have exploited progress in genome sequencing to assemble a comprehensive catalogue of over 1300 Snf2 family members. Multiple sequence alignment of the helicase-related regions enables 24 distinct subfamilies to be identified, a considerable expansion over earlier surveys. Where information is known, there is a good correlation between biological or biochemical function and these assignments, suggesting Snf2 family motor domains are tuned for specific tasks. Scanning of complete genomes reveals all eukaryotes contain members of multiple subfamilies, whereas they are less common and not ubiquitous in eubacteria or archaea. The large sample of Snf2 proteins enables additional distinguishing conserved sequence blocks within the helicase-like motor to be identified. The establishment of a phylogeny for Snf2 proteins provides an opportunity to make informed assignments of function, and the identification of conserved motifs provides a framework for understanding the mechanisms by which these proteins function.

Adenosine Triphosphatases↗

The genome of the African trypanosome Trypanosoma brucei.

African trypanosomes cause human sleeping sickness and livestock trypanosomiasis in sub-Saharan Africa. We present the sequence and analysis of the 11 megabase-sized chromosomes of Trypanosoma brucei. The 26-megabase genome contains 9068 predicted genes, including approximately 900 pseudogenes and approximately 1700 T. brucei-specific genes. Large subtelomeric arrays contain an archive of 806 variant surface glycoprotein (VSG) genes used by the parasite to evade the mammalian immune system. Most VSG genes are pseudogenes, which may be used to generate expressed mosaic genes by ectopic recombination. Comparisons of the cytoskeleton and endocytic trafficking systems with those of humans and other eukaryotic organisms reveal major differences. A comparison of metabolic pathways encoded by the genomes of T. brucei, T. cruzi, and Leishmania major reveals the least overall metabolic capability in T. brucei and the greatest in L. major. Horizontal transfer of genes of bacterial origin has contributed to some of the metabolic differences in these parasites, and a number of novel potential drug targets have been identified.

Amino Acids↗

A preliminary crystallographic analysis of the putative mevalonate diphosphate decarboxylase from Trypanosoma brucei.

Mevalonate diphosphate decarboxylase catalyses the last and least well characterized step in the mevalonate pathway for the biosynthesis of isopentenyl pyrophosphate, an isoprenoid precursor. A gene predicted to encode the enzyme from Trypanosoma brucei has been cloned, a highly efficient expression system established and a purification protocol determined. The enzyme gives monoclinic crystals in space group P2(1), with unit-cell parameters a = 51.5, b = 168.7, c = 54.9 A, beta = 118.8 degrees. A Matthews coefficient VM of 2.5 A3 Da(-1) corresponds to two monomers, each approximately 42 kDa (385 residues), in the asymmetric unit with 50% solvent content. These crystals are well ordered and data to high resolution have been recorded using synchrotron radiation.

Animals↗

Retrotransposon populations of Vicia species with varying genome size.

The (non-LTR) LINE and Ty3-gypsy-type LTR retrotransposon populations of three Vicia species that differ in genome size (Vicia faba, Vicia melanops and Vicia sativa) have been characterised. In each species the LINE retrotransposons comprise a complex, very heterogeneous set of sequences, while the Ty3-gypsy elements are much more homogeneous. Copy numbers of all three retrotransposon groups (Ty1-copia, Ty3-gypsy and LINE) in these species have been estimated by random genomic sequencing and Southern hybridisation analysis. The Ty3-gypsy elements are extremely numerous in all species, accounting for 18-35% of their genomes. The Ty1-copia group elements are somewhat less abundant and LINE elements are present in still lower amounts. Collectively, 20-45% of the genomes of these three Vicia species are comprised of retrotransposons. These data show that the three retrotransposon groups have proliferated to different extents in members of the Vicia genus and high proliferation has been associated with homogenisation of the retrotransposon population.

Blotting, Southern↗

GOtcha: a new method for prediction of protein function assessed by the annotation of seven genomes.

BACKGROUND: The function of a novel gene product is typically predicted by transitive assignment of annotation from similar sequences. We describe a novel method, GOtcha, for predicting gene product function by annotation with Gene Ontology (GO) terms. GOtcha predicts GO term associations with term-specific probability (P-score) measures of confidence. Term-specific probabilities are a novel feature of GOtcha and allow the identification of conflicts or uncertainty in annotation. RESULTS: The GOtcha method was applied to the recently sequenced genome for Plasmodium falciparum and six other genomes. GOtcha was compared quantitatively for retrieval of assigned GO terms against direct transitive assignment from the highest scoring annotated BLAST search hit (TOPBLAST). GOtcha exploits information deep into the 'twilight zone' of similarity search matches, making use of much information that is otherwise discarded by more simplistic approaches. At a P-score cutoff of 50%, GOtcha provided 60% better recovery of annotation terms and 20% higher selectivity than annotation with TOPBLAST at an E-value cutoff of 10(-4). CONCLUSIONS: The GOtcha method is a useful tool for genome annotators. It has identified both errors and omissions in the original Plasmodium falciparum annotation and is being adopted by many other genome sequencing projects.

Animals↗

ELM server: A new resource for investigating short functional sites in modular eukaryotic proteins.

Multidomain proteins predominate in eukaryotic proteomes. Individual functions assigned to different sequence segments combine to create a complex function for the whole protein. While on-line resources are available for revealing globular domains in sequences, there has hitherto been no comprehensive collection of small functional sites/motifs comparable to the globular domain resources, yet these are as important for the function of multidomain proteins. Short linear peptide motifs are used for cell compartment targeting, protein-protein interaction, regulation by phosphorylation, acetylation, glycosylation and a host of other post-translational modifications. ELM, the Eukaryotic Linear Motif server at http://elm.eu.org/, is a new bioinformatics resource for investigating candidate short non-globular functional motifs in eukaryotic proteins, aiming to fill the void in bioinformatics tools. Sequence comparisons with short motifs are difficult to evaluate because the usual significance assessments are inappropriate. Therefore the server is implemented with several logical filters to eliminate false positives. Current filters are for cell compartment, globular domain clash and taxonomic range. In favourable cases, the filters can reduce the number of retained matches by an order of magnitude or more.

Amino Acid Motifs↗

Visual representation of database search results: the RHIMS Plot.

SUMMARY: An algorithm and software are described that provide a fast method to produce a novel, function-oriented visualization of the results of a sequence database search. Text mining of sequence annotations allows position specific plots of potential functional similarity to be compared in a simple compact representation. AVAILABILITY: The application can be accessed via a web server at http://www.compbio.dundee.ac.uk. The RHIMS software may be obtained by request to the authors.

Algorithms↗

Genome sequence of the human malaria parasite Plasmodium falciparum.

The parasite Plasmodium falciparum is responsible for hundreds of millions of cases of malaria, and kills more than one million African children annually. Here we report an analysis of the genome sequence of P. falciparum clone 3D7. The 23-megabase nuclear genome consists of 14 chromosomes, encodes about 5,300 genes, and is the most (A + T)-rich genome sequenced to date. Genes involved in antigenic variation are concentrated in the subtelomeric regions of the chromosomes. Compared to the genomes of free-living eukaryotic microbes, the genome of this intracellular parasite encodes fewer enzymes and transporters, but a large proportion of genes are devoted to immune evasion and host-parasite interactions. Many nuclear-encoded proteins are targeted to the apicoplast, an organelle involved in fatty-acid and isoprenoid metabolism. The genome sequence provides the foundation for future studies of this organism, and is being exploited in the search for new drugs and vaccines to fight malaria.

Animals↗