PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Haplotype structures”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Optimal haplotype block-free selection of tagging SNPs for genome-wide association studies.

It is widely hoped that the study of sequence variation in the human genome will provide a means of elucidating the genetic component of complex diseases and variable drug responses. A major stumbling block to the successful design and execution of genome-wide disease association studies using single-nucleotide polymorphisms (SNPs) and linkage disequilibrium is the enormous number of SNPs in the human genome. This results in unacceptably high costs for exhaustive genotyping and presents a challenging problem of statistical inference. Here, we present a new method for optimally selecting minimum informative subsets of SNPs, also known as "tagging" SNPs, that is efficient for genome-wide selection. We contrast this method to published methods including haplotype block tagging, that is, grouping SNPs into segments of low haplotype diversity and typing a subset of the SNPs that can discriminate all common haplotypes within the blocks. Because our method does not rely on a predefined haplotype block structure and makes use of the weaker correlations that occur across neighboring blocks, it can be effectively applied across chromosomal regions with both high and low local linkage disequilibrium. We show that the number of tagging SNPs selected is substantially smaller than previously reported using block-based approaches and that selecting tagging SNPs optimally can result in a two- to threefold savings over selecting random SNPs.

Algorithms↗

Contribution of the LRP5 gene to normal variation in peak BMD in women.

UNLABELLED: The role of the LRP5 gene in rare BMD-related traits has recently been shown. We tested whether variation in this gene might play a role in normal variation in peak BMD. Association between SNPs in LRP5 and hip and spine BMD was measured in 1301 premenopausal women. Only a small proportion of the BMD variation was attributable to LRP5 in our sample. INTRODUCTION: Mutations in the low-density lipoprotein receptor-related protein 5 (LRP5) gene have been implicated as the cause of multiple distinct BMD-related rare Mendelian phenotypes. We sought to examine whether the LRP5 gene contributes to the observed variation in peak BMD in the normal population. MATERIALS AND METHODS: We genotyped 12 single nucleotide polymorphisms (SNPs) in LRP5 using allele-specific PCR and mass spectrometry methods. Linkage disequilibrium between the genotyped LRP5 SNPs was measured. We tested for association between these SNPs and both hip and spine BMD (adjusted for age and body weight) in 1301 healthy premenopausal women who took part in a sibling pair study aimed at identifying the genes underlying peak bone mass. Our study used both population-based (ANOVA) and family-based (quantitative transmission disequilibrium test) association methodology. RESULTS AND CONCLUSIONS: The linkage disequilibrium pattern and haplotype block structure within the LRP5 gene were consistent with that observed in other studies. Although significant evidence of association was found between LRP5 SNPs and both hip and spine BMD, only a small proportion of the total variation in these phenotypes was accounted for. The genotyped SNPs accounted for approximately 0.8% of the variation in femoral neck BMD and 1.1% of the variation in spine BMD. Results from our sample suggest that natural variation in and around LRP5 is not a major contributor to the observed variability in peak BMD at either the femoral neck or lumbar spine in white women.

Adult↗

Molecular genetic analysis among subspecies of two Eurasian sturgeon species, Acipenser baerii and A. stellatus.

Two species, the Siberian sturgeon, Acipenser baerii, and stellate sturgeon, A. stellatus, were studied using mitochondrial DNA (mtDNA) (D-loop, cytochrome b (cyt-b) and ND5/6 genes) sequencing to determine whether traditionally defined subspecies correspond to taxonomic entities and conservation management units. Initially, several mtDNA regions for each taxon (A. baerii: 737 bp D-loop, 750 bp ND5, 200 bp ND6, and 790 bp cyt-b; A. stellatus: 737 bp D-loop and 600 bp ND5) were examined. The D-loop was the most variable region and was sequenced for 35 A. baerii and 82 A. stellatus individuals. No fixed, diagnostic differences were found between any of the subspecies. Geographical structuring of haplotypes was observed within A. baerii, and gene flow estimates suggest isolation of the A. baerii baicalensis subspecies and the Yenisie and Lena River populations. No intraspecific subdivisioning was found within the genetic data for A. stellatus. The use of the phylogenetic criterion (fixed diagnostic differences) for identifying conservation units is compared to the rationale and results of other methods. Overall, morphologically and geographically based subspecies designations within Acipenseridae may not directly correspond to the biological entities appropriate for management and should not be used for conservation programmes without genetic support.

Animals↗

A block-free hidden Markov model for genotypes and its application to disease association.

We present a new stochastic model for genotype generation. The model offers a compromise between rigid block structure and no structure altogether: It reflects a general blocky structure of haplotypes, but also allows for "exchange" of haplotypes at nonboundary SNP sites; it also accommodates rare haplotypes and mutations. We use a hidden Markov model and infer its parameters by an expectation-maximization algorithm. The algorithm was implemented in a software package called HINT (haplotype inference tool) and tested on 58 datasets of genotypes. To evaluate the utility of the model in association studies, we used biological human data to create a simple disease association search scenario. When comparing HINT to three other models, HINT predicted association most accurately.

Algorithms↗

Data-mining methods as useful tools for predicting individual drug response: application to CYP2D6 data.

OBJECTIVES: Selecting a maximally informative subset of polymorphisms to predict a clinical outcome, such as drug response, requires appropriate search methods due to the increased dimensionality associated with looking at multiple genotypes. In this study, we investigated the ability of several pattern recognition methods to identify the most informative markers in the CYP2D6 gene for the prediction of CYP2D6 metabolizer status. METHODS: Four data-mining tools were explored: decision trees, random forests, artificial neural networks, and the multifactor dimensionality reduction (MDR) method. Marker selection was performed separately in eight population samples of different ethnic origin to evaluate to what extent the most informative markers differ across ethnic groups. RESULTS: Our results show that the number of polymorphisms required to predict CYP2D6 metabolic phenotype with a high accuracy can be dramatically reduced owing to the strong haplotype block structure observed at CYP2D6. MDR and neural networks provided nearly identical results and performed the best. CONCLUSION: Data-mining methods, such as MDR and neural networks, appear as promising tools to improve the efficiency of genotyping tests in pharmacogenetics with the ultimate goal of pre-screening patients for individual therapy selection with minimum genotyping effort.

Cytochrome P-450 CYP2D6↗

Positive selection of a pre-expansion CAG repeat of the human SCA2 gene.

A region of approximately one megabase of human Chromosome 12 shows extensive linkage disequilibrium in Utah residents with ancestry from northern and western Europe. This strikingly large linkage disequilibrium block was analyzed with statistical and experimental methods to determine whether natural selection could be implicated in shaping the current genome structure. Extended Haplotype Homozygosity and Relative Extended Haplotype Homozygosity analyses on this region mapped a core region of the strongest conserved haplotype to the exon 1 of the Spinocerebellar ataxia type 2 gene (SCA2). Direct DNA sequencing of this region of the SCA2 gene revealed a significant association between a pre-expanded allele [(CAG)8CAA(CAG)4CAA(CAG)8] of CAG repeats within exon 1 and the selected haplotype of the SCA2 gene. A significantly negative Tajima's D value (-2.20, p < 0.01) on this site consistently suggested selection on the CAG repeat. This region was also investigated in the three other populations, none of which showed signs of selection. These results suggest that a recent positive selection of the pre-expansion SCA2 CAG repeat has occurred in Utah residents with European ancestry.

Ataxins↗

Linkage disequilibrium testing when linkage phase is unknown.

Linkage disequilibrium, the nonrandom association of alleles from different loci, can provide valuable information on the structure of haplotypes in the human genome and is often the basis for evaluating the association of genomic variation with human traits among unrelated subjects. But, linkage phase of genetic markers measured on unrelated subjects is typically unknown, and so measurement of linkage disequilibrium, and testing whether it differs significantly from the null value of zero, requires statistical methods that can account for the ambiguity of unobserved haplotypes. A common method to test whether linkage disequilibrium differs significantly from zero is the likelihood-ratio statistic, which assumes Hardy-Weinberg equilibrium of the marker phenotype proportions. We show, by simulations, that this approach can be grossly biased, with either extremely conservative or liberal type I error rates. In contrast, we use simulations to show that a composite statistic, proposed by Weir and Cockerham, maintains the correct type I error rates, and, when comparisons are appropriate, has similar power as the likelihood-ratio statistic. We extend the composite statistic to allow for more than two alleles per locus, providing a global composite statistic, which is a strong competitor to the usual likelihood-ratio statistic.

Alleles↗

Glutathione S-transferase M1 status and breast cancer risk: a meta-analysis.

It is not yet clear whether Glutathione S-transferase M1 (GSTM1) polymorphisms affect the risk of breast cancer. The aim of this study is to provide a comprehensive meta-analysis of all the available, published case-control studies on the extent of the possible association between GSTM1 polymorphisms and susceptibility to breast cancer. Twenty case-control studies on GSTM1 and breast cancer were identified using both PUBMED and a manual search. Meta-analysis was conducted by the Peto method. Subgroup analyses were undertaken, in order to explore the relationship between effect sizes and the study characteristics. The overall odds ratio (OR) was found to be 1.06 (95% CI, 0.99-1.14). The OR for post-menopausal women with GSTM1 deficiency was determined to be 1.19 (95% CI, 1.05-1.34). In populations with a low frequency of GSTM1 deficiency, a greater increase was observed (OR, 1.20; 95% CI, 1.08-1.34). Furthermore, the highest associations were found in post-menopausal women with a low frequency of GSTM1 deficiency (OR, 1.44; 95% CI, 1.20-1.73). The fact that GSTM1 deficiency is not rare in the general population implies that the attributable risk for breast cancer could be sizable. Further studies focusing on the structure of haplotype blocks of GSTM1 are required in order to find a specific haplotype with a predisposing breast cancer susceptibility allele.

Breast Neoplasms↗

Molecular mapping of the HLA class II region in HLA-DR3 associated idiopathic membranous nephropathy.

Susceptibility to IMN is associated, in European Caucasoids, with the extended HLA haplotype in A1, B8, and DR3. It is unclear from previous investigations of HLA class II genes whether the association with A1, B8, DR3 is due to an HLA-DR or -DQ locus, or both, or to another locus linked to HLA class II. To examine genetic polymorphism over a more extensive area of DNA than previously, we carried out long range mapping of the HLA class II region of A1, B8, DR3 patients and healthy controls to discover if new markers of disease could be identified at this level of organization. Large fragments of genomic DNA were cut using enzymes with infrequent restriction sites, and were separated by pulsed field gel electrophoresis (PFGE) and analyzed using a series of probes which cover the HLA class II region. In several different DR3 haplotypes examined, the overall content of DNA and organization of the class II region were similar. However, both patient and control B8, DR3 haplotypes contained an extra Pvul site in the DRB region, compared to the disease-neutral B18, DR3 haplotype. Further, the DP region of the patient B8, DR3 haplotypes contained an additional partial BssHII cutting site which was not identified in the control B8, DR3 haplotypes. This structural heterogeneity in the vicinity of DP could have implications for genetic susceptibility to IMN and for linkage disequilibrium.

Chromosomes, Human, Pair 6↗

Contrasting signals of selection at the EDAR gene in global and Latin American populations.

The EDAR gene is a classic target of positive selection in humans, mainly through the nonsynonymous variant rs3827760 (EDARV370A) associated with ectodermal traits. Using high-resolution data from the 1000 Genomes Project, we combined sliding-window F_ST, BayeScan, and extended haplotype homozygosity (EHH) analyses to examine global and Latin American patterns of differentiation. Globally, a strong signal of positive selection was confirmed at EDAR, dominated by the rs3827760 haplotype background and its extended linkage disequilibrium structure. In contrast, within Latin America, differentiation reflected admixture-driven haplotype persistence rather than contemporary selection. A genome-wide FST scan comparing individuals from the upper and lower quartiles of Native American ancestry showed that EDAR lies among the most highly differentiated regions in this contrast, consistent with ancestry-driven haplotype structure rather than post-admixture adaptive evolution. These results indicate that EDAR retains its evolutionary signature globally but not within recently admixed populations, where demographic history rather than selection shapes its genetic landscape.

Humans↗

Comparative analysis of the disease-associated complement C4 gene from the HLA-A1, B8, DR3 haplotype.

Complement component C4 is an important protein of the classical, or antibody-mediated pathway of complement activation. Human C4 is located within the central region of the major histocompatibility complex on chromosome 6. Partial C4 deficiency has been associated with an increased susceptibility to immune complex disease. The strongest association with partial C4 deficiency is with systemic lupus erythematosus (SLE) and has been shown in most racial groups studied. Interestingly, Caucasian population studies have demonstrated an increased prevalence of C4A null alleles in SLE patients, in particular in association with the haplotype HLA-A1, B8, BfS, C4AQ0, C4B1, DR3. To investigate whether the C4 gene on this haplotype had any structural irregularities which may explain disease association, we sequenced the entire C4B gene from this haplotype. The results revealed that the gene encoded on the disease-associated haplotype carried major structural differences (when compared to C4A3) at the exonic level only in the C4d region. A high degree of conservation in both the 5' and 3' untranslated regions imply that disease associations will not be due to differential C4 expression as a result of regulatory differences between C4 genes. It appears likely that protein clearance mechanisms may account for the altered levels of C4 seen between different isotypes.

Amino Acid Sequence↗

Phylogeography of Arcterica nana (Ericaceae) suggests another range expansion history of Japanese alpine plants.

We conducted a phylogeographic study on the alpine plant Arcterica nana based on haplotypes of chloroplast DNA. Using a sequence of approximately 1,071 bp of intergenic spacers of chloroplast DNA (trnT-L, psbB-psbF), we detected 13 haplotypes among 193 individuals sampled from 22 populations. Two dominant haplotypes were distributed over the entire range of this species in Japan, and we found several local private haplotypes. An analysis of molecular variance (AMOVA) indicated no geographic structure within the haplotype distribution. In addition, the genetic distance was not related to its corresponding geographic distance (Mantel test: r=-0.049, P=0.66), indicating a homogeneous geographic structure throughout the entire distribution range in the Japanese archipelago. The most parsimonious explanation for this geographic structure is that A. nana spread across its extant distribution range in the Japanese archipelago through a recent range expansion event. However, this pattern is inconsistent with the previous phylogeography of Japanese alpine plants, which reveals that haplotypes in central Honshu are differentiated from those in more northern regions. Arcterica nana may have experienced a different history from other alpine plants, suggesting that the history of Japanese alpine flora may include at least two different geographic radiation patterns.

DNA, Chloroplast↗

High diversity of alpha-globin haplotypes in a Senegalese population, including many previously unreported variants.

RFLP haplotypes at the alpha-globin gene complex have been examined in 190 individuals from the Niokolo Mandenka population of Senegal: haplotypes were assigned unambiguously for 210 chromosomes. The Mandenka share with other African populations a sample size-independent haplotype diversity that is much greater than that in any non-African population: the number of haplotypes observed in the Mandenka is typically twice that seen in the non-African populations sampled to date. Of these haplotypes, 17.3% had not been observed in any previous surveys, and a further 19.1% have previously been reported only in African populations. The haplotype distribution shows clear differences between African and non-African peoples, but this is on the basis of population-specific haplotypes combined with haplotypes common to all. The relationship of the newly reported haplotypes to those previously recorded suggests that several mutation processes, particularly recombination as homologous exchange or gene conversion, have been involved in their production. A computer program based on the expectation-maximization (EM) algorithm was used to obtain maximum-likelihood estimates of haplotype frequencies for the entire data set: good concordance between the unambiguous and EM-derived sets was seen for the overall haplotype frequencies. Some of the low-frequency haplotypes reported by the estimation algorithm differ greatly, in structure, from those haplotypes known to be present in human populations, and they may not represent haplotypes actually present in the sample.

Genetic Variation↗

Hybridization and population genetics of two macaque species in Sulawesi, Indonesia.

This study investigates hybridization and population genetics of two species of macaque monkey in Sulawesi, Indonesia, using molecular markers from mitochondrial, autosomal, and Y-chromosome DNA. Hybridization is the interbreeding of individuals from different parental taxa that are distinguishable by one or more heritable characteristics. Because hybridization can affect population structure of the parental taxa, it is an important consideration for conservation management. On the Indonesian island of Sulawesi an explosive diversification of macaques has occurred; seven of 19 species in the genus Macaca live on this island. The contact zone of the subjects of this study, M. maura and M. tonkeana, is located at the base of the southwestern peninsula of Sulawesi. Land conversion in Sulawesi is occurring at an alarming pace; currently two species of Sulawesi macaque, one of which is M. maura, are classified as endangered species. Results of this study indicate that hybridization among M. maura and M. tonkeana has led to different distributions of molecular variation in mitochondrial DNA and nuclear DNA in the contact zone; mitochondrial DNA shows a sharp transition from M. maura to M. tonkeana haplotypes, but nuclear DNA from the parental taxa is homogenized in a narrow hybrid zone. Similarly, within M. maura divergent mitochondrial DNA haplotypes are geographically structured but population subdivision in the nuclear genome is low or absent. In M. tonkeana, mitochondrial DNA haplotypes are geographically structured and a high level of nuclear DNA population subdivision is present in this species. These results are largely consistent with a macaque behavioral paradigm of female philopatry and obligate male dispersal, suggest that introgression between M. maura and M. tonkeana is restricted to the hybrid zone, and delineate one conservation management unit in M. maura and at least two in M. tonkeana.

Animals↗

Immunogenetics of disease susceptibility: new perspectives in HLA.

Structural analysis of HLA class II molecular variation occurring within haplotypes implicated in specific HLA-associated diseases now provides more specific and sensitive mechanisms for investigation of genetic susceptibility to disease. Using the HLA-DR4 association with two distinct diseases, IDDM and JRA, as a model, we can conclude the following: There are at least seven distinct haplotypes which share the HLA DR4 specificity; these haplotypes include six alleles at the DR-beta genetic locus. These allelic differences are subtle, encompassing a very few clustered amino acid changes, but are sufficient to generate different patterns of T cell alloreactivity; there are at least three different alleles of DQ-beta genes associated with DR4+ haplotypes, with major structural differences recognized by biochemical analysis and by specific antibodies; different DR4-associated diseases are associated with different specific allelic variants of DR and DQ genes. DR4+ IDDM is most closely associated with the DQ 3.2 allele at DQ-beta; DR4+ JRA, on the other hand, appears to be highly associated with rare alleles at DR-beta, but not DQ. Notably, there are many alleles, and therefore DR4+ haplotypes, which are not implicated in 'HLA-DR4-associated' diseases.

Arthritis, Juvenile↗

Bayes estimates of haplotype effects.

We describe a Markov chain Monte Carlo implementation of a Bayesian approach to estimating associations of a trait with a large set of haplotypes recently introduced by Clayton and Jones [Am J Hum Genet 65:1161-9, 2000]. The model uses the length of the longest segment in common between any two haplotypes to define the prior correlation structure for the set of haplotype effects, using an intrinsic autocorrelation model. When applied to the Genetic Analysis Workshop 12 data for trait Q1, we found highly significant variation between haplotypes, using either a structured or unstructured covariance matrix.

Bayes Theorem↗

Further studies on using multiple-cross mapping (MCM) to map quantitative trait loci.

We have completed whole-genome scans for quantitative trait loci (QTLs) associated with acute ethanol-induced activation in the six F(2) intercrosses that can be formed from the C57BL/6J (B6), DBA/2J (D2) , BALB/cJ (C), and LP/J (LP) inbred strains. The goal was to test the hypothesis that given the relatively simple structure of the laboratory mouse genome, the same QTLs will be detected in multiple crosses which in turn will provide support for the strategy of multiple-cross mapping (MCM). QTLs with LOD scores greater than 4 were detected on Chrs 1, 2, 3, 8, 9, 13, 14, and 16. Only for the QTL on distal Chr 1 was there convincing evidence that the same or at least a very similar QTL was detected in multiple crosses. We also mapped the Chr 2 QTL directly in heterogeneous stock (HS) animals derived from the four inbred strains. At G(19) the QTL was mapped to an approximately 3-Mbp interval and this interval was associated with a haplotype block with a largely biallelic structure: B6-L:C-D2. We conclude that mapping in HS animals not only provides significantly greater QTL resolution, at least in some cases it provides significantly more information about the QTL haplotype structure.

Animals↗

Familial clustering of IGHC deletions and duplications: functional and molecular analysis.

The human immunoglobulin heavy chain constant region locus (IGHC) comprises nine genes and two pseudogenes clustered in a 350 kilobase (kb) region on chromosome 14q32. Several IGHC haplotypes with single or multiple gene deletions and duplications have been characterized. The most likely mechanism accounting for these unusual haplotypes is the unequal crossing-over between homologous regions within the locus. Here we report the analysis of an unusual case of familial clustering of deletions/duplications. In the two branches of the BON family, three duplicated and two deleted haplotypes, all probably independent in origin, have been characterized. The structure of the haplotypes, one of which is described here for the first time, supports the hypothesis of homologous unequal crossing-over as the origin of recombinant haplotypes. The analysis of serological markers in a subject carrying one deleted and one duplicated haplotype allowed us the first direct inferences concerning the functions of the duplicated IGHC haplotypes.

Adolescent↗