PubMed Health⌕ Search

Biomedical subjects

Scott J Tebbutt

Publications and source records attributed to Scott J Tebbutt.

6 recordsLinked to original sources

Dynamic variable selection in SNP genotype autocalling from APEX microarray data.

BACKGROUND: Single nucleotide polymorphisms (SNPs) are DNA sequence variations, occurring when a single nucleotide--adenine (A), thymine (T), cytosine (C) or guanine (G)--is altered. Arguably, SNPs account for more than 90% of human genetic variation. Our laboratory has developed a highly redundant SNP genotyping assay consisting of multiple probes with signals from multiple channels for a single SNP, based on arrayed primer extension (APEX). This mini-sequencing method is a powerful combination of a highly parallel microarray with distinctive Sanger-based dideoxy terminator sequencing chemistry. Using this microarray platform, our current genotype calling system (known as SNP Chart) is capable of calling single SNP genotypes by manual inspection of the APEX data, which is time-consuming and exposed to user subjectivity bias. RESULTS: Using a set of 32 Coriell DNA samples plus three negative PCR controls as a training data set, we have developed a fully-automated genotyping algorithm based on simple linear discriminant analysis (LDA) using dynamic variable selection. The algorithm combines separate analyses based on the multiple probe sets to give a final posterior probability for each candidate genotype. We have tested our algorithm on a completely independent data set of 270 DNA samples, with validated genotypes, from patients admitted to the intensive care unit (ICU) of St. Paul's Hospital (plus one negative PCR control sample). Our method achieves a concordance rate of 98.9% with a 99.6% call rate for a set of 96 SNPs. By adjusting the threshold value for the final posterior probability of the called genotype, the call rate reduces to 94.9% with a higher concordance rate of 99.6%. We also reversed the two independent data sets in their training and testing roles, achieving a concordance rate up to 99.8%. CONCLUSION: The strength of this APEX chemistry-based platform is its unique redundancy having multiple probes for a single SNP. Our model-based genotype calling algorithm captures the redundancy in the system considering all the underlying probe features of a particular SNP, automatically down-weighting any 'bad data' corresponding to image artifacts on the microarray slide or failure of a specific chemistry. In this regard, our method is able to automatically select the probes which work well and reduce the effect of other so-called bad performing probes in a sample-specific manner, for any number of SNPs.

Algorithms↗

MACGT: multi-dimensional automated clustering genotyping tool for analysis of microarray-based mini-sequencing data.

SUMMARY: Multi-dimensional Automated Clustering Genotyping Tool (MACGT) is a Java application that clusters complex multi-dimensional vector data derived from single nucleotide polymorphism (SNP) genotyping experiments using mini-sequencing based microarray chemistries such as arrayed primer extension (APEX). Spot intensity output files from microarray experiments across multiple samples are imported into MACGT. The datasets can include four channels of intensity data for each spot, replica spots for each SNP probe and multiple probe types (APEX and allele-specific APEX probes) on both DNA strands for each SNP. MACGT automatically clusters these multi-dimensionality datasets for each SNP across multiple samples. Incorporation of additional array datasets from known samples that have previously validated SNP genotype calls allows unknown samples to be automatically assigned a genotype based on the clustering, along with numerical measures of confidence for each genotype call. Calling accuracy by MACGT exceeds 98% when applied to genotyping data from APEX microarrays, and can be increased to >99.5% by applying thresholds to the confidence measures.

Algorithms↗

Deoxynucleotides can replace dideoxynucleotides in minisequencing by arrayed primer extension.

Scientific literature describing arrayed primer extension and other array-based minisequencing technologies consistently cite the requirement for four fluorescent dideoxynucleotides (with concomitant absence/inactivation of deoxynucleotides) to ensure single-base extension and thus sequence-specific intensity data that can be interpreted as a base call or genotype. We present compelling evidence that fluorescent deoxynucleotides can reliably be used in microarray minisequencing experiments, generating fluorescent sequence extension intensity profiles that are homologous to the single-base extensions obtained with terminator dideoxynucleotides. Due to the almost 10-fold higher costs (and limited fluorophore choice) of many commercially available fluorescent dideoxynucleotides, compared to fluorescent deoxynucleotides, as well as other potentially constraining intellectual property and licensing issues, this hitherto dismissed microarray chemistry represents an important reevaluation in the field of array-based genotyping and related enzymology.

Chromosome Mapping↗

SNP Chart: an integrated platform for visualization and interpretation of microarray genotyping data.

UNLABELLED: SNP Chart is a Java application for the visualization and interpretation of microarray genotyping data primarily derived from arrayed primer extension-based chemistries. Spot intensity output files from microarray analysis tools are imported into SNP Chart, together with a multi-channel TIFF image of the original array experiment and a list of the actual single nucleotide polymorphisms (SNPs) being tested. Data from different and/or replicate probes that interrogate the same SNP, but that are scattered across the array grid, can be reassembled into a single chart format, specific for the SNP. This allows a quick and very effective 'visualization'/'quality control' of the data from multiple probes for the same SNP that can be easily interpreted and manually scored as a genotype. AVAILABILITY: http://www.snpchart.ca.

Computer Graphics↗

Microarray genotyping resource to determine population stratification in genetic association studies of complex disease.

We have developed a robust microarray genotyping chip that will help advance studies in genetic epidemiology. In population-based genetic association studies of complex disease, there could be hidden genetic substructure in the study populations, resulting in false-positive associations. Such population stratification may confound efforts to identify true associations between genotype/haplotype and phenotype. Methods relying on genotyping additional null single nucleotide polymorphism (SNP) markers have been proposed, such as genomic control (GC) and structured association (SA), to correct association tests for population stratification. If there is an association of a disease with null SNPs, this suggests that there is a population subset with different genetic background plus different disease susceptibility. Genotyping over 100 null SNPs in the large numbers of patient and control DNA samples that are required in genetic association studies can be prohibitively expensive. We have therefore developed and tested a resequencing chip based on arrayed primer extension (APEX) from over 2000 DNA probe features that facilitate multiple interrogations of each SNP, providing a powerful, accurate, and economical means to simultaneously determine the genotypes at 110 null SNP loci in any individual. Based on 1141 known genotypes from other research groups, our GC SNP chip has an accuracy of 98.5%, including non-calls.

Animals↗

Evaluation of gene targeting by homologous recombination in ovine somatic cells.

Mouse models for some human genetic diseases are limited in their applications since they do not accurately reproduce the phenotype of the human disease. It has been suggested that larger animals, for example sheep, might produce more useful models, as some aspects of sheep physiology and anatomy are more similar to those of humans. The development of methods to clone animals from somatic cells provides a potential novel route to generate such large animal models following gene targeting. Here, we assess targeting of the cystic fibrosis transmembrane conductance regulator (CFTR) gene in ovine somatic cells using homologous recombination (HR) of targeting constructs with extensive (>11 kb) homology. Electroporation of these constructs into ovine fetal and post-natal fibroblasts generated G418-resistant clones, but none analyzed had undergone HR, suggesting that at least for this locus, it is an extremely inefficient process. Karyotyping of targeted ovine fetal fibroblasts showed them to be less chromosomally stable than post-natal fibroblasts, and, moreover, extended culture periods caused them to senesce, adversely affecting their viability for use as nuclear transfer donor cells. These data stress the importance of donor cell choice in somatic cell cloning and suggest that culture time be kept to a minimum prior to nuclear transfer in order to maximize cell viability.

Animals↗