PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Haplotype structures”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics↗

Molecular haplotyping of genomic DNA for multiple single-nucleotide polymorphisms located kilobases apart using long-range polymerase chain reaction and intramolecular ligation.

Genetic polymorphisms are well-recognized causes of interindividual differences in disease risk and treatment response in humans. For genes containing multiple single-nucleotide polymorphisms (SNPs), haplotype structure is often the principal determinant of phenotypic consequences, and haplotype distribution represents the best approach for assessing patterns of linkage disequilibrium. To permit more widespread molecular determination of haplotypes, we developed a simple yet robust method to determine haplotype structure for multiple SNPs located up to 30 kb apart in genomic DNA using long-range polymerase chain reaction (LR-PCR) and intramolecular ligation. Complete concordance was shown between the new method and conventional approaches, such as family pedigree analysis or cloning and sequencing. The availability of a simple method to directly determine haplotype structure using genomic DNA, without family pedigree analysis, cloning or complex instrumentation, provides an important new tool for elucidating the genetic determinants of drug disposition and effects, disease risk, and molecular evolution.

Alleles↗

Haplotype block structure and its applications to association studies: power and study designs.

Recent studies have shown that the human genome has a haplotype block structure, such that it can be divided into discrete blocks of limited haplotype diversity. In each block, a small fraction of single-nucleotide polymorphisms (SNPs), referred to as "tag SNPs," can be used to distinguish a large fraction of the haplotypes. These tag SNPs can potentially be extremely useful for association studies, in that it may not be necessary to genotype all SNPs; however, this depends on how much power is lost. Here we develop a simulation study to quantitatively assess the power loss for a variety of study designs, including case-control designs and case-parental control designs. First, a number of data sets containing case-parental or case-control samples are generated on the basis of a disease model. Second, a small fraction of case and control individuals in each data set are genotyped at all the loci, and a dynamic programming algorithm is used to determine the haplotype blocks and the tag SNPs based on the genotypes of the sampled individuals. Third, the statistical power of tests was evaluated on the basis of three kinds of data: (1) all of the SNPs and the corresponding haplotypes, (2) the tag SNPs and the corresponding haplotypes, and (3) the same number of randomly chosen SNPs as the number of tag SNPs and the corresponding haplotypes. We study the power of different association tests with a variety of disease models and block-partitioning criteria. Our study indicates that the genotyping efforts can be significantly reduced by the tag SNPs, without much loss of power. Depending on the specific haplotype block-partitioning algorithm and the disease model, when the identified tag SNPs are only 25% of all the SNPs, the power is reduced by only 4%, on average, compared with a power loss of approximately 12% when the same number of randomly chosen SNPs is used in a two-locus haplotype analysis. When the identified tag SNPs are approximately 14% of all the SNPs, the power is reduced by approximately 9%, compared with a power loss of approximately 21% when the same number of randomly chosen SNPs is used in a two-locus haplotype analysis. Our study also indicates that haplotype-based analysis can be much more powerful than marker-by-marker analysis.

Algorithms↗

Robustness of inference of haplotype block structure.

In this report, we examine the validity of the haplotype block concept by comparing block decompositions derived from public data sets by variants of several leading methods of block detection. We first develop a statistical method for assessing the concordance of two block decompositions. We then assess the robustness of inferred haplotype blocks to the specific detection method chosen, to arbitrary choices made in the block-detection algorithms, and to the sample analyzed. Although the block decompositions show levels of concordance that are very unlikely by chance, the absolute magnitude of the concordance may be low enough to limit the utility of the inference. For purposes of SNP selection, it seems likely that methods that do not arbitrarily impose block boundaries among correlated SNPs might perform better than block-based methods.

Algorithms↗

Direct determination of MUC5B promoter haplotypes based on the method of single-strand conformation polymorphism and their statistical estimation.

Haplotype-based human genome research is important in identifying disease susceptibility genes efficiently. Although haplotype reconstruction by statistical methods is widely used, direct haplotype determination by molecular techniques has also been developed as a complementary method for statistical estimation. In this study, we demonstrate a molecular haplotyping method making use of single-strand conformation polymorphism (SSCP) gels. We identified 10 common SNPs and a dinucleotide insertion/deletion polymorphism within 2-kb region upstream of the transcription initiation site of MUC5B and determined haplotype structure, dividing the region into two DNA fragments. Real haplotypes were determined unambiguously by our SSCP-based analysis with fragments longer than 1 kb. Haplotypes reconstructed from diploid genotypes in the same region by the statistical methods including EM algorithm were also evaluated. Direct comparison between statistical estimation and direct determination of haplotypes revealed that major haplotypes containing multiple marker sites showing strong LD are estimated in great accuracy but that a variety of haplotypes reflecting weak LD are not reconstructed precisely enough. Our data can be helpful in implementing molecular haplotyping or statistical estimation, since usage of these methods may be determined depending on the haplotype structures.

Base Sequence↗

Four novel defective alleles and comprehensive haplotype analysis of CYP2C9 in Japanese.

Genetic variations in cytochrome P450 2C9 (CYP2C9) are known to contribute to interindividual and interethnic variability in response to clinical drugs such as warfarin. In the present study, CYP2C9 from 263 Japanese subjects was resequenced, resulting in the discovery of 62 variations including 32 novel ones. In addition to the two known non-synonymous single nucleotide polymorphisms (SNPs), Ile359Leu (*3; allele frequency=0.030) and Leu90Pro (*13; 0.002), seven novel non-synonymous SNPs, Leu17Ile (0.002), Lys118ArgfsX9 (*25; 0.002), Thr130Arg (*26; 0.002), Arg150Leu (*27; 0.004), Gln214Leu (*28; 0.002), Pro279Thr (*29; 0.002) and Ala477Thr (*30; 0.002), were found. Functional characterization of novel alleles using a mammalian cell expression system in vitro revealed that *25 was a null allele and that *26, *28 and *30 were defective alleles. The *26 product showed a 90% decrease in the Vmax value but little change in the Km value towards diclofenac. Both *28 and *30 products showed two-fold higher Km values and three-fold lower Vmax values than the *1 allele, suggesting the importance of Gln214 and Ala477 for substrate recognition. Linkage disequilibrium and haplotype analyses were performed using the detected variations. Only five haplotypes (frequency >0.02) accounted for most (>87%) of the inferred haplotypes, and they were closely associated with the haplotypes of CYP2C19 in Japanese. Although the haplotype structure of CYP2C9 was rather simple in Japanese, the haplotype distribution was quite different from those previously reported in Caucasians and Africans. Taken together, novel defective alleles and detailed haplotype structures would be useful for determining metabolic phenotypes of CYP2C9 substrate drugs in Japanese and probably Asians.

Alleles↗

Inverted triplications formed by iterative template switches generate structural variant diversity at genomic disorder loci.

The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. The point of template switching in 4 samples was shown to be a segment of ∼2.2-5.5 kb of 100% nucleotide similarity within inverted repeat pairs. These data provide experimental evidence that inverted low-copy repeats act as recombinant substrates. This type of CGR can result in multiple conformers generating diverse SV haplotypes in susceptible dosage-sensitive loci.

Humans↗

Significant variation in haplotype block structure but conservation in tagSNP patterns among global populations.

The initial belief that haplotype block boundaries and haplotypes were largely shared across populations was a foundation for constructing a haplotype map of the human genome using common SNP markers. The HapMap data document the generality of a block-like pattern of linkage disequilibrium (LD) with regions of low and high haplotype diversity but differences among the populations. Studies of many additional populations demonstrate that LD patterns can be highly variable among populations both across and within geographic regions. Because of this variation, emphasis has shifted to the generalizability of tagSNPs, those SNPs that capture the bulk of variation in a region. We have examined the LD and tagSNP patterns based upon over 2000 individual samples in 38 populations and 134 SNPs in 10 genetically independent loci for a total of 517 kb with an average density of 1 SNP/5 kb. Four different 'block' definitions and the pairwise LD tagSNP selection algorithm have been applied. Our results not only confirm large variation in block partition among populations from different regions (agreeing with previous studies including the HapMap) but also show that significant variation can occur among populations within geographic regions. None of the block-defining algorithms produces a consistent pattern within or across all geographic groups. In contrast, tagSNP transferability is much greater than the similarity of LD patterns and, although not perfect, some generalizations of transferability are possible. The analyses show an asymmetric pattern of tagSNP transferability coinciding with the subsetting of variation attributed to the spread of modern humans around the world.

Genetic Variation↗

Genome-wide definitive haplotypes determined using a collection of complete hydatidiform moles.

We present genome-wide definitive haplotypes, determined using a collection of 74 Japanese complete hydatidiform moles, each carrying a genome derived from a single sperm. The haplotypes incorporate 281,439 common SNPs, genotyped with a high throughput array-based oligonucleotide hybridization technique. Comparison of haplotypes inferred from pseudoindividuals (constructed from randomized mole pairs) with those of moles showed some switch errors in resolution of phases by the computational inference method. The effects of these errors on local haplotype structure and selection of tag SNPs are discussed. We also show that definitive haplotypes of moles may be useful for elucidation of long-range haplotype structure, and should be more effective for detecting extended haplotype homozygosity indicative of positive selection.

Female↗

Recurrent structural variation and recent turnover at the 17q21.31 locus in humans and great apes.

The 17q21.31 locus in humans harbors several complex structural haplotypes including a ~970kb inversion. Different inversion haplotypes have been associated with susceptibility to microdeletions causing Koolen-de Vries syndrome and variation in fecundity and recombination rates. Here, using 210 haplotype-resolved human genome assemblies and pangenome graph-based approaches we characterize 11 distinct structural haplotypes, several of which have not been previously described. Extending our analyses to a set of haplotype-resolved great-ape genomes, we characterize the structure of an independent inversion in chimpanzees which extends an additional 650kb, encompasses 5 additional genes, and is ~2 million years younger than the human inversion. We further determine that gorillas exhibit an independent duplication of the KANSL1 gene which may predispose them to Koolen-de Vries syndrome causing microdeletions. Using short read sequencing data we characterize 17q21.31 haplotype diversity worldwide in ~5174 individuals from 107 populations finding increased frequencies of KANSL1 duplication-containing haplotypes in both European and South Asian populations as well as 8 double recombination events between inverted and non-inverted haplotypes ranging in size from 20-180kb. Finally, using 626 ancient Eurasian human genomes we show the frequency of haplotypes containing KANSL1 duplications has increased ~6-fold over the past 12 thousand years in Europe. Together, our results highlight the dynamics, complexity, and recurrent, independent evolution of a medically relevant locus across humans and great apes.

Journal Article↗

The architecture of the tau haplotype block in different ethnicities.

We have assessed the pattern of the extended haplotype block over the tau gene which covers a region of approximately 2 Mb in different ethnicities. This analysis shows that the pattern of linkage disequilibrium over the tau region is shared by different ethnic groups indicating that haplotype structure in human is ancient. We discuss this observation in terms of the establishment of the haplotype structure and the possible impact of the tau haplotype on neurodegeneration in humans.

Ethnicity↗

Accounting for haplotype uncertainty in matched association studies: a comparison of simple and flexible techniques.

Population-based case-control studies measuring associations between haplotypes of single nucleotide polymorphisms (SNPs) are increasingly popular, in part because haplotypes of a few "tagging" SNPs may serve as surrogates for variation in relatively large sections of the genome. Due to current technological limitations, haplotypes in cases and controls must be inferred from unphased genotypic data. Using individual-specific inferred haplotypes as covariates in standard epidemiologic analyses (e.g., conditional logistic regression) is an attractive analysis strategy, as it allows adjustment for nongenetic covariates, provides omnibus and haplotype-specific tests of association, and can estimate haplotype and haplotype x environment interaction effects. In principle, some adjustment for the uncertainty in inferred haplotypes should be made. Via simulation, we compare the performance (bias and mean squared error of haplotype and haplotype x environment interaction effect estimates) of several analytic strategies using inferred haplotypes in the context of matched case-control data. These strategies include using only the most likely haplotype assignment, the expectation substitution approach described by Stram et al. ([2003b] Hum. Hered. 55:179-190) and others, and an improper version of multiple imputation. For relatively uncomplicated haplotype structures and moderate haplotype relative risks (</=2), all methods performed comparably well (small bias with appropriately-sized confidence intervals). For larger relative risks, the most likely haplotype and multiple imputation strategies showed noticeable bias towards the null; the expectation substitution strategy still performed well. When there was more uncertainty in the inferred haplotypes, the most likely and multiple imputation strategies showed even more bias towards the null, while the expectation substitution method had slightly smaller than nominal confidence intervals for larger relative risks (>/=5). An application to progesterone-receptor haplotypes and endometrial cancer further illustrates that the performance of all these methods depends on how well the observed haplotypes "tag" the unobserved causal variant.

Algorithms↗

Multiple local PfDHFR I164L haplotype expansions drive Plasmodium falciparum antifolate resistance in Uganda.

Mutations in the Plasmodium falciparum genes, pfdhfr and pfdhps, drive antifolate resistance and threaten malaria control in regions where sulfadoxine-pyrimethamine (SP) is the primary chemoprevention strategy. The spatial patterns and evolutionary dynamics of these mutations in high-transmission settings remain incompletely understood. Here we genotyped 11 resistance-associated mutations in pfdhfr and pfdhps in 4,725&#x2009;P. falciparum isolates collected from 16 Ugandan health facilities as part of annual surveillance between 2016 and 2022. Notably, we show that the frequency of PfDHFR I164L, which confers higher pyrimethamine resistance, increased over time from 19.4% to 32.4%. Using identity-by-descent, haplotype structure, and extended haplotype homozygosity analyses, we show that PfDHFR I164L is present on multiple haplotype backgrounds and undergoes localised expansions, without detectable signatures of recent positive selection at all but one site. Our results suggest that the evolution of antifolate resistance, driven by PfDHFR I164L, is spatially heterogeneous and complex in regions that primarily use SP chemoprevention programmes.

Plasmodium falciparum↗

Sequence analysis of the mannose-binding lectin (MBL2) gene reveals a high degree of heterozygosity with evidence of selection.

Human mannose-binding protein (MBL) is a component of innate immunity. To capture the common genetic variants of MBL2, we resequenced a 10.0 kb region that includes MBL2 in 102 individuals representing four major US ethnic groups. In all, 87 polymorphic sites were observed, indicating a high level of heterozygosity (total pi=18.3 x 10(-4)). Estimates of linkage disequilibrium across MBL2 indicate that it is divided into two blocks, with a probable recombination hot spot in the 3' end. Three non-synonymous SNPs in exon 1 of the encoding MBL2 gene and three upstream SNPs form common 'secretor haplotypes' that can predict circulating levels. Common variants have been associated with increased susceptibility to infection and autoimmune diseases. The high frequencies of B, C and D alleles in certain populations suggest a possible selective advantage for heterozygosity. There is limited diversity of haplotype structure; the 'secretor haplotypes' lie on a restricted number of extended haplotypes, which could include additional linked SNPs, which might also have possible functional implications. There is evidence for gene conversion in the region between the two blocks, in the last exon. Our data should form the basis for conducting MBL2 candidate gene association studies using a locus-wide approach.

Ethnicity↗

Algorithms for association study design using a generalized model of haplotype conservation.

There is considerable interest in computational methods to assist in the use of genetic polymorphism data for locating disease-related genes. Haplotypes, contiguous sets of correlated variants, may provide a means of reducing the difficulty of the data analysis problems involved. The field to date has been dominated by methods based on the "haplotype block" hypothesis, which assumes discrete population-wide boundaries between conserved genetic segments, but there is strong reason to believe that haplotype blocks do not fully capture true haplotype conservation patterns. In this paper, we address the computational challenges of using a more flexible, block-free representation of haplotype structure called the "haplotype motif" model for downstream analysis problems. We develop algorithms for htSNP selection and missing data inference using this more generalized model of sequence conservation. Application to a dataset from the literature demonstrates the practical value of these block-free methods.

Algorithms↗

Comparative analysis of haplotype association mapping algorithms.

BACKGROUND: Finding the genetic causes of quantitative traits is a complex and difficult task. Classical methods for mapping quantitative trail loci (QTL) in miceuse an F2 cross between two strains with substantially different phenotype and an interval mapping method to compute confidence intervals at each position in the genome. This process requires significant resources for breeding and genotyping, and the data generated are usually only applicable to one phenotype of interest. Recently, we reported the application of a haplotype association mapping method which utilizes dense genotyping data across a diverse panel of inbred mouse strains and a marker association algorithm that is independent of any specific phenotype. As the availability of genotyping data grows in size and density, analysis of these haplotype association mapping methods should be of increasing value to the statistical genetics community. RESULTS: We describe a detailed comparative analysis of variations on our marker association method. In particular, we describe the use of inferred haplotypes from adjacent SNPs, parametric and nonparametric statistics, and control of multiple testing error. These results show that nonparametric methods are slightly better in the test cases we study, although the choice of test statistic may often be dependent on the specific phenotype and haplotype structure being studied. The use of multi-SNP windows to infer local haplotype structure is critical to the use of a diverse panel of inbred strains for QTL mapping. Finally, because the marginal effect of any single gene in a complex disease is often relatively small, these methods require the use of sensitive methods for controlling family-wise error. We also report our initial application of this method to phenotypes cataloged in the Mouse Phenome Database. CONCLUSION: The use of inbred strains of mice for QTL mapping has many advantages over traditional methods. However, there are also limitations in comparison to the traditional linkage analysis from F2 and RI lines. Application of these methods requires careful consideration of algorithmic choices based on both theoretical and practical factors. Our findings suggest general guidelines, though a complete evaluation of these methods can only be performed as more genetic data in complex diseases becomes available.

Algorithms↗

The effect of haplotype-block definitions on inference of haplotype-block structure and htSNPs selection.

It has been recently suggested that the human genome is organized as a series of haplotype blocks, and efforts to create a genome-wide haplotype map are already underway. Several computational algorithms have been proposed to partition the genome. However, little is known about their behaviors in relation to the haplotype-block partitioning and haplotype-tagging SNPs selection. Here, we present a systematic comparison of three classes of haplotype-block partition definitions, a diversity-based method, a linkage-disequilibrium (LD)-based method, and a recombination-based method. The data used were derived from a coalescent simulation under both a uniform recombination model and one that assumes recombination hotspots. There were considerable differences in haplotype information loss in the measure of entropy when the partition methods were compared under different population-genetics scenarios. Under both recombination models, the results from the LD-based definition and the recombination-based definition were more similar to each other than were the results from the diversity-based definition. This work demonstrates that when undertaking haplotype-based association mapping, the choice of haplotype-block definition and SNP selection requires careful consideration.

Algorithms↗

Durability of marker-quantitative trait loci haplotypes in structured populations.

Given the relative ease of identifying genetic markers linked to QTL (compared to finding the loci themselves), it is natural to ask whether linked markers can be used to address questions concerning the contemporary dynamics and recent history of the QTL. In particular, can a marker allele found associated with a QTL allele in a QTL mapping study be used to track population dynamics or the history of the QTL allele? For this strategy to succeed, the marker-QTL haplotype must persist in the face of recombination over the relevant time frame. Here we investigate the dynamics of marker-QTL haplotype frequencies under recombination, population structure, and divergent selection to assess the potential utility of linked markers for a population genetic study of QTL. For two scenarios, described as "secondary contact" and "novel allele," we use both deterministic and stochastic methods to describe the influence of gene flow between habitats, the strength of divergent selection, and the genetic distance between a marker and the QTL on the persistence of marker-QTL haplotypes. We find that for most reasonable values of selection on a locus (s < or = 0.5) and migration (m > 1%) between differentially selected populations, haplotypes of typically spaced markers (5 cM) and QTL do not persist long enough (>100 generations) to provide accurate inference of the allelic state at the QTL.

Alleles↗