PubMed Health⌕ Search

Biomedical subjects

Colin A M Semple

Publications and source records attributed to Colin A M Semple.

13 recordsLinked to original sources

Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.

Genome-wide experimental methods to identify disease genes, such as linkage analysis and association studies, generate increasingly large candidate gene sets for which comprehensive empirical analysis is impractical. Computational methods employ data from a variety of sources to identify the most likely candidate disease genes from these gene sets. Here, we review seven independent computational disease gene prioritization methods, and then apply them in concert to the analysis of 9556 positional candidate genes for type 2 diabetes (T2D) and the related trait obesity. We generate and analyse a list of nine primary candidate genes for T2D genes and five for obesity. Two genes, LPL and BCKDHA, are common to these two sets. We also present a set of secondary candidates for T2D (94 genes) and for obesity (116 genes) with 58 genes in common to both diseases.

Computational Biology↗

Genome-wide analysis of mammalian promoter architecture and evolution.

Mammalian promoters can be separated into two classes, conserved TATA box-enriched promoters, which initiate at a well-defined site, and more plastic, broad and evolvable CpG-rich promoters. We have sequenced tags corresponding to several hundred thousand transcription start sites (TSSs) in the mouse and human genomes, allowing precise analysis of the sequence architecture and evolution of distinct promoter classes. Different tissues and families of genes differentially use distinct types of promoters. Our tagging methods allow quantitative analysis of promoter usage in different tissues and show that differentially regulated alternative TSSs are a common feature in protein-coding genes and commonly generate alternative N termini. Among the TSSs, we identified new start sites associated with the majority of exons and with 3' UTRs. These data permit genome-scale identification of tissue-specific promoters and analysis of the cis-acting elements associated with them.

3' Untranslated Regions↗

Heterotachy in mammalian promoter evolution.

We have surveyed the evolutionary trends of mammalian promoters and upstream sequences, utilising large sets of experimentally supported transcription start sites (TSSs). With 30,969 well-defined TSSs from mouse and 26,341 from human, there are sufficient numbers to draw statistically meaningful conclusions and to consider differences between promoter types. Unlike previous smaller studies, we have considered the effects of insertions, deletions, and transposable elements as well as nucleotide substitutions. The rate of promoter evolution relative to that of control sequences has not been consistent between lineages nor within lineages over time. The most pronounced manifestation of this heterotachy is the increased rate of evolution in primate promoters. This increase is seen across different classes of mutation, including substitutions and micro-indel events. We investigated the relationship between promoter and coding sequence selective constraint and suggest that they are generally uncorrelated. This analysis also identified a small number of mouse promoters associated with the immune response that are under positive selection in rodents. We demonstrate significant differences in divergence between functional promoter categories and identify a category of promoters, not associated with conventional protein-coding genes, that has the highest rates of divergence across mammals. We find that evolutionary rates vary both on a fine scale within mammalian promoters and also between different functional classes of promoters. The discovery of heterotachy in promoter evolution, in particular the accelerated evolution of primate promoters, has important implications for our understanding of human evolution and for strategies to detect primate-specific regulatory elements.

Animals↗

The complexity of selection at the major primate beta-defensin locus.

BACKGROUND: We have examined the evolution of the genes at the major human beta-defensin locus and the orthologous loci in a range of other primates and mouse. For the first time these data allow us to examine selective episodes in the more recent evolutionary history of this locus as well as the ancient past. We have used a combination of maximum likelihood based tests and a maximum parsimony based sliding window approach to give a detailed view of the varying modes of selection operating at this locus. RESULTS: We provide evidence for strong positive selection soon after the duplication of these genes within an ancestral mammalian genome. Consequently variable selective pressures have acted on beta-defensin genes in different evolutionary lineages, with episodes both of negative, and more rarely positive selection, during the divergence of primates. Positive selection appears to have been more common in the rodent lineage, accompanying the birth of novel, rodent-specific beta-defensin genes. These observations allow a fuller understanding of the evolution of mammalian innate immunity. In both the rodent and primate lineages, sites in the second exon have been subject to positive selection and by implication are important in functional diversity. A small number of sites in the mature human peptides were found to have undergone repeated episodes of selection in different primate lineages. Particular sites were consistently implicated by multiple methods at positions throughout the mature peptides. These sites are clustered at positions predicted to be important for the specificity of the antimicrobial or chemoattractant properties of beta-defensins. Surprisingly, sites within the prepropeptide region were also implicated as being subject to significant positive selection, suggesting previously unappreciated functional significance for this region. CONCLUSIONS: Identification of these putatively functional sites has important implications for our understanding of beta-defensin function and for novel antibiotic design.

Animals↗

POCUS: mining genomic sequence annotation to predict disease genes.

Here we present POCUS (prioritization of candidate genes using statistics), a novel computational approach to prioritize candidate disease genes that is based on over-representation of functional annotation between loci for the same disease. We show that POCUS can provide high (up to 81-fold) enrichment of real disease genes in the candidate-gene shortlists it produces compared with the original large sets of positional candidates. In contrast to existing methods, POCUS can also suggest counterintuitive candidates.

Autistic Disorder↗

Duplication and selection in the evolution of primate beta-defensin genes.

BACKGROUND: Innate immunity is the first line of defense against microorganisms in vertebrates and acts by providing an initial barrier to microorganisms and triggering adaptive immune responses. Peptides such as beta-defensins are an important component of this defense, providing a broad spectrum of antimicrobial activity against bacteria, fungi, mycobacteria and several enveloped viruses. Beta-defensins are small cationic peptides that vary in their expression patterns and spectrum of pathogen specificity. Disruptions in beta-defensin function have been implicated in human diseases, including cystic fibrosis, and a fuller understanding of the variety, function and evolution of human beta-defensins might form the basis for novel therapies. Here we use a combination of laboratory and computational techniques to characterize the main human beta-defensin locus on chromosome 8p22-p23. RESULTS: In addition to known genes in the region we report the genomic structures and expression patterns of four novel human beta-defensin genes and a related pseudogene. These genes show an unusual pattern of evolution, with rapid divergence between second exon sequences that encode the mature beta-defensin peptides matched by relative stasis in first exons that encode signal peptides. CONCLUSIONS: We conclude that the 8p22-p23 locus has evolved by successive rounds of duplication followed by substantial divergence involving positive selection, to produce a diverse cluster of paralogous genes established before the human-baboon divergence more than 23 million years ago. Positive selection, disproportionately favoring alterations in the charge of amino-acid residues, is implicated as driving second exon divergence in these genes.

Amino Acid Sequence↗

UBA domain containing proteins in fission yeast.

The ubiquitin-proteasome pathway for intracellular proteolysis is involved in a series of cellular and molecular functions, including the degradation of bulk proteins, cell cycle control, DNA repair, antigen presentation, vesicle transport and the regulation of signal transudation pathways and transcription. Considering this variety of cell biological processes, it is puzzling that until recently only very few proteins were known to possess the ability to interact specifically with ubiquitin chains. However, several ubiquitin binding proteins have now been identified and the binding domains have been characterised on both the functional and structural levels. One example of a widespread ubiquitin binding module is the ubiquitin associated (UBA) domain. Here, we discuss the approximately 15 UBA domain containing proteins encoded in the relatively small genome of the fission yeast Schizosaccharomyces pombe. The proteins display remarkable differences in their domain organisation, indicating that these potential ubiquitin binding proteins are involved in various cell activities.

Amino Acid Sequence↗

Signal sequence conservation and mature peptide divergence within subgroups of the murine beta-defensin gene family.

Beta-defensins are two exon genes which encode broad spectrum antimicrobial cationic peptides. We have analyzed the largest murine cluster of these genes which localizes to chromosome 8. Using hidden Markov models, we identified six beta-defensin exon 2-like sequences and subsequently found full-length expressed transcripts for these novel genes. Expression was high in brain and reproductive tissues. Eleven beta-defensins could be grouped into two clear subgroups by virtue of their position and high signal sequence (exon 1 encoded) identity. In contrast, however, there was a very low level of sequence conservation in the exon 2 region encoding the mature antimicrobial peptide. Examination of the gene sequences of orthologs in other rodents also revealed an excess of nucleotide changes that altered amino acids in the mature peptide region. Evolutionary analysis revealed strong evidence that following gene duplication, exon 1 and surrounding noncoding DNA show little divergence within subgroups. The focus for rapid sequence divergence is localized in the DNA encoding the mature peptide and this is driven by accelerated positive selection. This mechanism of evolution is consistent with the role of this gene family as defense against bacterial pathogens and the sequence changes have implications for novel antibiotic design.

Animals↗

The comparative proteomics of ubiquitination in mouse.

Ubiquitination is a common posttranslational modification in eukaryotic cells, influencing many fundamental cellular processes. Defects in ubiquitination and the processes it mediates are involved in many human disease states. The ubiquitination of a substrate involves four classes of enzymes:a ubiquitin-activating enzyme (E1), a ubiquitin-conjugating enzyme (E2), a ubiquitin protein ligase (E3), and a de-ubiquitinating enzyme (DUB). A substantial number of E1s (four), E2s (13), E3s (97), and DUBs (six) that were previously unknown in the mouse are included in the FANTOM2 Representative Transcript and Protein Set (RTPS). Many of the genes encoding these proteins will constitute promising candidates for involvement in disease. In addition, the RTPS provides the basis for the most comprehensive survey of ubiquitination-associated proteins across eukaryotes undertaken to date. Comparisons of these proteins across human and other organisms suggest that eukaryotic evolution has been associated with an increase in the number and diversity of E3s (possessing either zinc-finger RING, F-box, or HECT domains) and DUBs (containing the ubiquitin thiolesterase family 2 domain). These increases in numbers are too large to be accounted for by the presence of fragmentary proteins in the data sets examined. Much of this innovation appears to have been associated with the emergence of multicellular organisms, and subsequently of vertebrates, increasing the opportunity for complex regulation of ubiquitination-mediated cellular and developmental processes.

Animals↗

SNP genotyping on pooled DNAs: comparison of genotyping technologies and a semi automated method for data storage and analysis.

We have compared the accuracy, efficiency and robustness of three methods of genotyping single nucleotide polymorphisms on pooled DNAs. We conclude that (i) the frequencies of the two alleles in pools should be corrected with a factor for unequal allelic amplification, which should be estimated from the mean ratio of a set of heterozygotes (k); (ii) the repeatability of an assay is more important than pinpoint accuracy when estimating allele frequencies, and assays should therefore be optimised to increase the repeatability; and (iii) the size of a pool has a relatively small effect on the accuracy of allele frequency estimation. We therefore recommend that large pools are genotyped and replicated a minimum of four times. In addition, we describe statistical approaches to allow rigorous comparison of DNA pool results. Finally, we describe an extension to our ACeDB database that facilitates management and analysis of the data generated by association studies.

Automation↗

Computational comparison of human genomic sequence assemblies for a region of chromosome 4.

Much of the available human genomic sequence data exist in a fragmentary draft state following the completion of the initial high-volume sequencing performed by the International Human Genome Sequencing Consortium (IHGSC) and Celera Genomics (CG). We compared six draft genome assemblies over a region of chromosome 4p (D4S394-D4S403), two consecutive releases by the IHGSC at University of California, Santa Cruz (UCSC), two consecutive releases from the National Centre for Biotechnology Information (NCBI), the public release from CG, and a hybrid assembly we have produced using IHGSC and CG sequence data. This region presents particular problems for genomic sequence assembly algorithms as it contains a large tandem repeat and is sparsely covered by draft sequences. The six assemblies differed both in terms of their relative coverage of sequence data from the region and in their estimated rates of misassembly. The CG assembly method attained the lowest level of misassembly, whereas NCBI and UCSC assemblies had the highest levels of coverage. All assemblies examined included <60% of the publicly available sequence from the region. At least 6% of the sequence data within the CG assembly for the D4S394-D4S403 region was not present in publicly available sequence data. We also show that even in a problematic region, existing software tools can be used with high-quality mapping data to produce genomic sequence contigs with a low rate of rearrangements.

Chromosomes, Human, Pair 4↗