PubMed Health⌕ Search

Biomedical subjects

Paul Hardenbol

Publications and source records attributed to Paul Hardenbol.

6 recordsLinked to original sources

Optimal genotype determination in highly multiplexed SNP data.

High-throughput genotyping technologies that enable large association studies are already available. Tools for genotype determination starting from raw signal intensities need to be automated, robust, and flexible to provide optimal genotype determination given the specific requirements of a study. The key metrics describing the performance of a custom genotyping study are assay conversion, call rate, and genotype accuracy. These three metrics can be traded off against each other. Using the highly multiplexed Molecular Inversion Probe technology as an example, we describe a methodology for identifying the optimal trade-off. The methodology comprises: a robust clustering algorithm and assessment of a large number of data filter sets. The clustering algorithm allows for automatic genotype determination. Many different sets of filters are then applied to the clustered data, and performance metrics resulting from each filter set are calculated. These performance metrics relate to the power of a study and provide a framework to choose the most suitable filter set to the particular study.

Algorithms↗

Large-scale characterization of public database SNPs causing non-synonymous changes in three ethnic groups.

Single nucleotide polymorphisms (SNPs) that lead to non-synonymous changes in proteins may have functional effects and be subject to selection. Hence they are of particular interest in the study of genetic diseases. We have genotyped approximately 28,000 such SNPs in three ethnic populations (the HapMap plates) and ten primate species and analyzed these data for evidence of selection. We find SNPs predicted by PolyPhen to be damaging, have lower allele frequencies, and are particularly likely to be population-specific. We have also grouped SNPs by molecular function or biological process of the associated genes and find evidence that selection may be acting in concert on classes of genes.

Animals↗

Population structure, differential bias and genomic control in a large-scale, case-control association study.

The main problems in drawing causal inferences from epidemiological case-control studies are confounding by unmeasured extraneous factors, selection bias and differential misclassification of exposure. In genetics the first of these, in the form of population structure, has dominated recent debate. Population structure explained part of the significant +11.2% inflation of test statistics we observed in an analysis of 6,322 nonsynonymous SNPs in 816 cases of type 1 diabetes and 877 population-based controls from Great Britain. The remainder of the inflation resulted from differential bias in genotype scoring between case and control DNA samples, which originated from two laboratories, causing false-positive associations. To avoid excluding SNPs and losing valuable information, we extended the genomic control method by applying a variable downweighting to each SNP.

Adolescent↗

Positive selection of a pre-expansion CAG repeat of the human SCA2 gene.

A region of approximately one megabase of human Chromosome 12 shows extensive linkage disequilibrium in Utah residents with ancestry from northern and western Europe. This strikingly large linkage disequilibrium block was analyzed with statistical and experimental methods to determine whether natural selection could be implicated in shaping the current genome structure. Extended Haplotype Homozygosity and Relative Extended Haplotype Homozygosity analyses on this region mapped a core region of the strongest conserved haplotype to the exon 1 of the Spinocerebellar ataxia type 2 gene (SCA2). Direct DNA sequencing of this region of the SCA2 gene revealed a significant association between a pre-expanded allele [(CAG)8CAA(CAG)4CAA(CAG)8] of CAG repeats within exon 1 and the selected haplotype of the SCA2 gene. A significantly negative Tajima's D value (-2.20, p < 0.01) on this site consistently suggested selection on the CAG repeat. This region was also investigated in the three other populations, none of which showed signs of selection. These results suggest that a recent positive selection of the pre-expansion SCA2 CAG repeat has occurred in Utah residents with European ancestry.

Ataxins↗

Highly multiplexed molecular inversion probe genotyping: over 10,000 targeted SNPs genotyped in a single tube assay.

Large-scale genetic studies are highly dependent on efficient and scalable multiplex SNP assays. In this study, we report the development of Molecular Inversion Probe technology with four-color, single array detection, applied to large-scale genotyping of up to 12,000 SNPs per reaction. While generating 38,429 SNP assays using this technology in a population of 30 trios from the Centre d'Etude Polymorphisme Humain family panel as part of the International HapMap project, we established SNP conversion rates of approximately 90% with concordance rates >99.6% and completeness levels >98% for assays multiplexed up to 12,000plex levels. Furthermore, these individual metrics can be "traded off" and, by sacrificing a small fraction of the conversion rate, the accuracy can be increased to very high levels. No loss of performance is seen when scaling from 6,000plex to 12,000plex assays, strongly validating the ability of the technology to suppress cross-reactivity at high multiplex levels. The results of this study demonstrate the suitability of this technology for comprehensive association studies that use targeted SNPs in indirect linkage disequilibrium studies or that directly screen for causative mutations.

Chromosome Inversion↗

Multiplexed genotyping with sequence-tagged molecular inversion probes.

We report on the development of molecular inversion probe (MIP) genotyping, an efficient technology for large-scale single nucleotide polymorphism (SNP) analysis. This technique uses MIPs to produce inverted sequences, which undergo a unimolecular rearrangement and are then amplified by PCR using common primers and analyzed using universal sequence tag DNA microarrays, resulting in highly specific genotyping. With this technology, multiplex analysis of more than 1,000 probes in a single tube can be done using standard laboratory equipment. Genotypes are generated with a high call rate (95%) and high accuracy (>99%) as determined by independent sequencing.

Cells, Cultured↗