PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Haplotype structures”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Selection and evaluation of tagging SNPs in the neuronal-sodium-channel gene SCN1A: implications for linkage-disequilibrium gene mapping.

Association studies are widely seen as the most promising approach for finding polymorphisms that influence genetically complex traits, such as common diseases and responses to their treatment. Considerable interest has therefore recently focused on the development of methods that efficiently screen genomic regions or whole genomes for gene variants associated with complex phenotypes. One key element in this search is the use of linkage disequilibrium to gain maximal information from typing a selected subset of highly informative single-nucleotide polymorphism (SNP) markers, now often called "tagging SNPs" (tSNPs). Probably the most common approach to linkage-disequilibrium gene mapping involves a three-step program: (1) characterization of the haplotype structure in candidate genes or genomic regions of interest, (2) identification of tSNPs sufficient to represent the most common haplotypes, and (3) typing of tSNPs in clinical material. Early definitions of tSNPs focused on the amount of haplotype diversity that they explained. To select tSNPs that would have maximal power in a genetic association study, however, we have developed optimization criteria based on the r2 measure of association and have compared these with other criteria based on the haplotype diversity. To evaluate the full program and to assess how well the selected tags are likely to perform, we have determined the haplotype structure and have assessed tSNPs in the SCN1A gene, an important candidate gene for sporadic epilepsy. We find that as few as four tSNPs are predicted to maintain a consistently high r2 value with all other common SNPs in the gene, indicating that the tags could be used in an association study with only a modest reduction in power relative to direct assays of all common SNPs. This implies that very large case-control studies can be screened for variation in hundreds of candidate genes with manageable experimental effort, once tSNPs are identified. However, our results also show that tSNPs identified in one population may not necessarily perform well in another, indicating that the preliminary study to identify tSNPs and the later case-control study should be performed in the same population. Our results also indicate that tSNPs will not easily identify discrepant SNPs, which lie on importantly discriminating but apparently short genealogical branches. This could significantly complicate tagging approaches for phenotypes influenced by variants that have experienced positive selection.

Adult↗

Linkage disequilibrium and haplotype diversity in the genes of the renin-angiotensin system: findings from the family blood pressure program.

Association studies of candidate genes with complex traits have generally used one or a few single nucleotide polymorphisms (SNPs), although variation in the extent of linkage disequilibrium (LD) within genes markedly influences the sensitivity and precision of association studies. The extent of LD and the underlying haplotype structure for most candidate genes are still unavailable. We sampled 193 blacks (African-Americans) and 160 whites (European-Americans) and estimated the intragenic LD and the haplotype structure in four genes of the renin-angiotensin system. We genotyped 25 SNPs, with all but one of the pairs spaced between 1 and 20 kb, thus providing resolution at small scale. The pattern of LD within a gene was very heterogeneous. Using a robust method to define haplotype blocks, blocks of limited haplotype diversity were identified at each locus; between these blocks, LD was lost owing to the history of recombination events. As anticipated, there was less LD among blacks, the number of haplotypes was substantially larger, and shorter haplotype segments were found, compared with whites. These findings have implications for candidate-gene association studies and indicate that variation between populations of European and African origin in haplotype diversity is characteristic of most genes.

Adult↗

Haplotypic analysis of the TNF locus by association efficiency and entropy.

BACKGROUND: To understand the causal basis of TNF associations with disease, it is necessary to understand the haplotypic structure of this locus. We genotyped 12 single-nucleotide polymorphisms (SNPs) distributed over 4.3 kilobases in 296 healthy, unrelated Gambian and Malawian adults. We generated 592 high-quality haplotypes by integrating family- and population-based reconstruction methods. RESULTS: We found 32 different haplotypes, of which 13 were shared between the two populations. Both populations were haplotypically diverse (gene diversity = 0.80, Gambia; 0.85, Malawi) and significantly differentiated (p < 10-5 by exact test). More than a quarter of marker pairs showed evidence of intragenic recombination (29% Gambia; 27% Malawi). We applied two new methods of analyzing haplotypic data: association efficiency analysis (AEA), which describes the ability of each SNP to detect every other SNP in a case-control scenario; and the entropy maximization method (EMM), which selects the subset of SNPs that most effectively dissects the underlying haplotypic structure. AEA revealed that many SNPs in TNF are poor markers of each other. The EMM showed that 8 of 12 SNPs (Gambia) and 7 of 12 SNPs (Malawi) are required to describe 95% of the haplotypic diversity. CONCLUSIONS: The TNF locus in the Gambian and Malawi sample is haplotypically diverse and has a rich history of intragenic recombination. As a consequence, a large proportion of TNF SNPs must be typed to detect a disease-modifying SNP at this locus. The most informative subset of SNPs to genotype differs between the two populations.

Adult↗

SNPs, haplotypes, and model selection in a candidate gene region: the SIMPle analysis for multilocus data.

Modern molecular techniques make discovery of numerous single nucleotide polymorphims (SNPs) in candidate gene regions feasible. Conventional analysis relies on either independent tests with each variant or the use of haplotypes in association analysis. The first technique ignores the dependencies between SNPs. The second, though it may increase power, often introduces uncertainty by estimating haplotypes from population data. Additionally, as the number of loci expands for a haplotype, ambiguity in interpretation increases for determining the underlying genetic components driving a detected association. Here, we present a genotype-level analysis to jointly model the SNPs via a SNP interaction model with phase information (SIMPle) to capture the underlying haplotype structure. This analysis estimates both the risk associated with each variant and the importance of phase between pairwise combinations of SNPs. Thus, rather than selecting between genotype- or haplotype-level approaches, the SIMPle method frames the analysis of multilocus data in a model selection paradigm, the aim to determine which SNPs, phase terms, and linear combinations best describe the relation between genetic variation and a trait of interest. To avoid unstable estimation due to sparse data and to incorporate both the dependencies among terms and the uncertainty in model selection, we propose a Bayes model averaging procedure. This highlights key SNPs and phase terms and yields a set of best representative models. Using simulations, we demonstrate the utility of the SIMPle model to identify crucial SNPs and underlying haplotype structures across a variety of causal models and genetic architectures.

Bayes Theorem↗

Haplotype analysis in population genetics and association studies.

Several studies of haplotype structures in the human genome in various populations have been published recently. Such knowledge may provide valuable information on human evolutionary history and lead to the development of more efficient strategies to identify genetic variants that increase susceptibility to human diseases. In this review, we summarize the current understanding of haplotype structure, diversity, and distribution in the human genome, with a focus on statistical issues in using haplotypes for studies of population genetics and evolutionary history, as well as to identify genetic variants underlying complex human traits.

Evolution, Molecular↗

Multi-locus selection and the structure of variation at the white gene of Drosophila melanogaster.

We surveyed sequence variation and divergence for the entire 5972-bp transcriptional unit of the white gene in 15 lines of Drosophila melanogaster and one line of D. simulans. We found a very high degree of haplotypic structuring for the polymorphisms in the 3' half of the gene, as opposed to the polymorphisms in the 5' half. To determine the evolutionary mechanisms responsible for this pattern we sequenced a 1612-bp segment of the white gene from an additional 33 lines of D. melanogaster from a European and a North American population. This 1612-bp segment encompasses an 834-bp region of the white gene in which the polymorphisms form high frequency haplotypes that cannot be explained by a neutral equilibrium model of molecular evolution. The small number of recombinants in the 834-bp region suggests epistatic selection as the cause of the haplotypic structuring, while an investigation of nucleotide diversity supports a directional selection hypothesis. A multi-locus selection model that combines features from both hypotheses and takes the recent history of D. melanogaster into account may be the best explanation for these data.

ATP-Binding Cassette Transporters↗

Haplotype analysis of the human collectin placenta 1 (hCL-P1) gene.

Collectins are a family of C-type lectins found in vertebrates. These proteins have four regions, a relatively short N-terminal region, a collagen-like region, an alpha-helical coiled coil, and a carbohydrate recognition domain. Collectins are involved in host defense through their ability to bind carbohydrate antigens on microorganisms. Type A scavenger receptors are classical-type scavenger receptors that also have collagen-like domains. We previously described a new scavenger receptor, collectin from placenta [collectin placenta 1 (CL-P1)]. CL-P1 is a type II membrane protein with all four regions. We found that CL-P1 can bind and phagocytize both bacteria and yeast. In addition to that, it reacts with oxidized low-density lipoprotein (LDL) but not with acetylated LDL. These results suggest that CL-P1 might play important roles in host defenses and/or atherosclerosis formation. One rational strategy to study the role of CL-P1 in these pathological conditions would be to perform a haplotype association study using human samples. As a first step for this strategy, we analyzed the haplotype structure of the CL-P1gene. By sequencing the CL-P1 gene in ten Japanese volunteers, we identified five single-nucleotide polymorphisms (SNPs) with a minor allele frequency of at least 29%. To obtain SNPs in the 5'-upstream region of the gene, we screened a total of 20 SNPs described in the database and finally picked up one SNP for the present study. Thus, a total of six SNPs, one in the 5'-upstream region, two in intron 2, one in exon 5, and two in exon 6, were used to analyze the haplotype structure of the gene, with DNAs derived from 54 individuals (108 alleles). The analysis revealed that only two of six SNPs showed significant linkage disequilibrium ( r(2) > 0.5) with each other. This haplotype information may be useful in disease-association studies in which a contribution of the CL-P1 gene has been suspected, especially in immunological disturbance or atherosclerosis. Two SNPs in exon 6, both leading to amino acid substitutions, could be candidates for influencing disease susceptibility.

Amino Acid Substitution↗

The APOA1/C3/A4/A5 gene cluster, lipid metabolism and cardiovascular disease risk.

PURPOSE OF REVIEW: APOA1/C3/A4/A5 are key components modulating lipoprotein metabolism and cardiovascular disease risk. This review examines the evidence regarding linkage disequilibrium and haplotype structure within the A1/C3/A4/A5 cluster, and assesses its association with plasma lipids and cardiovascular disease risk. In addition, we use genomic information from several species to draw inferences about the location of functional variants within this cluster. RECENT FINDINGS: The close physical distance of these genes and the interrelated functions of these apolipoproteins have encumbered attempts to determine the role of individual variants on lipid metabolism. Therefore, current research aims to define linkage disequilibrium and haplotype structure within this cluster. Functional variants in regulatory regions are most interesting as they are potentially amenable to therapy. Comparative genomics can contribute to the identification of such functional variants. SUMMARY: Genetic variability at the APOA1/C3/A4/A5 cluster has been examined in relation to lipid metabolism and cardiovascular disease risk. However, the findings are inconsistent. This is partly due to the classic approach of studying single and mostly nonfunctional polymorphisms. Moreover, allelic expression may depend on the concurrent presence of environmental factors. Association studies using haplotypes should increase the power to detect true associations and interactions. We hypothesize that phenotypes observed in association with transcriptional regulatory variants can be readily modified by environmental factors. Therefore, studies focusing on regulatory variants may be more fruitful to locate/define future therapeutic targets.

Apolipoprotein A-I↗

Association of PDCD1 with susceptibility to systemic lupus erythematosus: evidence of population-specific effects.

OBJECTIVE: The A allele of the PD1.3 single-nucleotide polymorphism (SNP) on the programmed cell death gene PDCD1 was markedly more frequent in patients with systemic lupus erythematosus (SLE) than in unaffected controls in a recent study involving large sets of Swedish, European American, and Mexican families. This study sought to determine the role of PDCD1 in susceptibility to SLE in the Spanish population. METHODS: Seven PDCD1 SNPs were studied in 518 SLE patients and 800 healthy control subjects who had been recruited in 5 distant towns spanning continental Spain. Patients and controls were of Spanish ancestry. The diagnosis of SLE was in accordance with the American College of Rheumatology updated classification criteria. RESULTS: The A allele of the PD1.3 polymorphism was significantly less frequent in Spanish female patients with SLE than in Spanish female controls (9.0% versus 13.0%, odds ratio 0.67, 95% confidence interval 0.50-0.89). This difference was consistent across the 5 sets of samples grouped by town of recruitment. The other PDCD1 SNPs were not associated with SLE susceptibility. The haplotype structure of PDCD1 in the Spanish controls was different from that reported in other healthy control populations. CONCLUSION: Our results confirm the association of PDCD1 with susceptibility to SLE, but the findings show a lack of involvement of the PD1.3 SNP, which is contrary to the role of the PD1.3 A allele observed previously. These contradictory results probably reflect population differences in the haplotype structure of the PDCD1 locus. More research focusing on new polymorphisms and identifying associations in other populations will be needed to clarify the role of PDCD1 in SLE susceptibility.

Alleles↗

Scrutiny of the glutamine-fructose-6-phosphate transaminase 1 (GFPT1) locus reveals conserved haplotype block structure not associated with diabetic nephropathy.

Glutamine-fructose-6-phosphate transaminase 1 (GFAT) is the rate-limiting enzyme of the hexosamine pathway that has been implicated in the pathogenesis of diabetic nephropathy. As such, we hypothesized that GFPT1, which encodes for GFAT, may confer genetic susceptibility to this complication among Caucasians. Screening of all known functional regions of GFPT1 revealed six single nucleotide polymorphisms (SNPs) that were located in the promoter, introns, and 3' untranslated region. The approximately 60 kb GFPT1 locus was encompassed in a single conserved haplotype block, and two tagging SNPs were sufficient to capture >90% of the haplotype diversity. Analysis of these SNPs in a case-control study made up of type 1 diabetic subjects (324 case subjects with diabetic nephropathy and 289 control subjects with normoalbuminuria despite >15 years of diabetes) revealed no significant association even after stratification by sex, diabetes duration, glucose control, and blood pressure. Similar results were obtained among type 2 diabetic subjects (202 case and 114 control subjects). Genetic variation in GFPT1 is thus unlikely to have a major impact on susceptibility to diabetic nephropathy.

3' Untranslated Regions↗

Interethnic variability of ERCC2 polymorphisms.

Excision Repair Cross-Complementing Rodent Repair Group 2 (ERCC2) plays an important role in DNA repair by eliminating bulky DNA adducts produced by platinum agents during the nucleotide excision repair pathway. Several studies have associated polymorphisms in ERCC2 with response to platinum therapy, lung cancer risk, and DNA repair capacity. This study examined ERCC2 polymorphisms and haplotype structure across 18.9 kb in 95 European, 95 African, and 95 Asian individuals. Single-nucleotide polymorphisms (SNPs) (ERCC2 -9164 A>T, -1989 A>G, -516 G>A, 468 C>A [Arg156Arg], 1737 C>T [Val579Val], 2133 C>T [Asp711Asp], and 2251 T>G [Lys751Gln]) were mined and mapped using Golden Path, PolyMAPr, and Promolign. Genotyping was performed using PCR and pyrosequencing. Allele frequencies ranged from 0 to 0.47 (Europeans), 0.05 to 0.72 (Africans), and 0 to 0.47 (Asians). The synonymous cSNP at codon 579 could not be confirmed in our populations. There were significant differences in haplotype structure and frequency between populations. This information on ERCC2 genomic structure will allow the construction of definitive studies to clarify the clinical role of this important gene.

Adult↗

The structure of haplotype blocks in the human genome.

Haplotype-based methods offer a powerful approach to disease gene mapping, based on the association between causal mutations and the ancestral haplotypes on which they arose. As part of The SNP Consortium Allele Frequency Projects, we characterized haplotype patterns across 51 autosomal regions (spanning 13 megabases of the human genome) in samples from Africa, Europe, and Asia. We show that the human genome can be parsed objectively into haplotype blocks: sizable regions over which there is little evidence for historical recombination and within which only a few common haplotypes are observed. The boundaries of blocks and specific haplotypes they contain are highly correlated across populations. We demonstrate that such haplotype frameworks provide substantial statistical power in association studies of common genetic variation across each region. Our results provide a foundation for the construction of a haplotype map of the human genome, facilitating comprehensive genetic association studies of human disease.

Africa↗

The SRY-1532 site of the human Y chromosome is subject to recurrent single nucleotide mutations.

Haplotype determination based on three Y-linked polymorphic sites, 92R7 (C/T), SRY-1532 (A/G), and YAP (-/+), in 127 males belonging to three caste Hindu populations of South India (Vizag Brahmins, Peruru Brahmins, and Kammas) and 13 males belonging to a migrant group (the Siddis) showed the existence of all four haplotypes (CA-, CG-, TG-, and TA-) under the YAP- background. This finding suggests that the reverse mutation (G-->A) at the SRY-1532 site, described earlier in the literature, is present in South Indian populations as well. The YAP+ mutation was seen in only five Siddi individuals. Four of these were of the CG+ haplotype structure, but a novel haplotype (CA+) was found in one male. To explain the occurrence of the six haplotypes found within these three sites, a haplotype tree is constructed that introduces a new reverse mutation at the SRY-1532 site (G-->A), occurring under the CG+ background after the migrant Siddi population arrived in India.

Gene Frequency↗

HaploBlockFinder: haplotype block analyses.

UNLABELLED: Recent studies have unveiled discrete block-like structures of linkage disequilibrium (LD) in the human genome. We have developed a set of computer programs to analyze the block-like LD structures (haplotype blocks) based on haplotype data. Three definitions of haplotype block are supported, including minimal LD range, no historic recombination, and chromosome coverage. Tagged SNPs that uniquely distinguish common haplotypes are identified. A greedy algorithm was used to improve the efficiency. Two separate utilities were also provided to assist visual inspection of haplotype block structure and pattern of linkage disequilibrium. AVAILABILITY: A web interface for the HaploBlockFinder is available at http://cgi.uc.edu/cgi-bin/kzhang/haploBlockFinder.cgi the source codes are also freely available on the web site.

Algorithms↗

Molecular population genetics of herbivore-induced protease inhibitor genes in European aspen (Populus tremula L., Salicaceae).

Plants defend themselves against the attack of natural enemies by using an array of both constitutively expressed and induced defenses. Long-lived woody perennials are overrepresented among plant species that show strong induced defense responses, whereas annual plants and crop species are underrepresented. However, most studies of plant defense genes have been performed on annual or short-lived perennial weeds or crop species. Here I use molecular population genetic methods to survey six wound-inducible protease inhibitors (PIs) in a long-lived woody, perennial plant species, the European aspen (Populus tremula), to evaluate the likelihood of either recurrent selective sweeps or balancing selection maintaining amino acid polymorphisms in these genes. The results show that none of the six PI genes have reduced diversities at synonymous sites, as would be expected in the presence of recurrent selective sweeps. However, several genes show some evidence of nonneutral evolution such as enhanced linkage disequilibrium and a large number of high-frequency-derived mutations. A group of at least four Kunitz trypsin inhibitor genes appear to have experienced elevated levels of nonsynonymous substitutions, indicating allelic turnover on an evolutionary timescale. One gene, TI1, has enhanced levels of intraspecific polymorphism at nonsynonymous sites and also has an unusual haplotype structure characterized by two divergent haplotypes occurring at roughly equal frequencies in the sample. One haplotype has very low levels of intraallelic nucleotide diversity, whereas the other haplotype has levels of diversity comparable to other genes in P. tremula. Patterns of sequence diversity at TI1 do not fit a simple model of either balancing selection or recurrent selective sweeps. This suggests that selection at TI1 is more complex, possibly involving allelic cycling.

Base Sequence↗

Blocks of limited haplotype diversity revealed by high-resolution scanning of human chromosome 21.

Global patterns of human DNA sequence variation (haplotypes) defined by common single nucleotide polymorphisms (SNPs) have important implications for identifying disease associations and human traits. We have used high-density oligonucleotide arrays, in combination with somatic cell genetics, to identify a large fraction of all common human chromosome 21 SNPs and to directly observe the haplotype structure defined by these SNPs. This structure reveals blocks of limited haplotype diversity in which more than 80% of a global human sample can typically be characterized by only three common haplotypes.

Algorithms↗

htSNPer1.0: software for haplotype block partition and htSNPs selection.

BACKGROUND: There is recently great interest in haplotype block structure and haplotype tagging SNPs (htSNPs) in the human genome for its implication on htSNPs-based association mapping strategy for complex disease. Different definitions have been used to characterize the haplotype block structure in the human genome, and several different performance criteria and algorithms have been suggested on htSNPs selection. RESULTS: A heuristic algorithm, generalized branch-and-bound algorithm, is applied to the searching of minimal set of haplotype tagging SNPs (htSNPs) according to different htSNPs performance criteria. We develop a software htSNPer1.0 to implement the algorithm, and integrate three htSNPs performance criteria and four haplotype block definitions for haplotype block partitioning. It is a software with powerful Graphical User Interface (GUI), which can be used to characterize the haplotype block structure and select htSNPs in the candidate gene or interested genomic regions. It can find the global optimization with only a fraction of the computing time consumed by exhaustive searching algorithm. CONCLUSION: htSNPer1.0 allows molecular geneticists to perform haplotype block analysis and htSNPs selection using different definitions and performance criteria. The software is a powerful tool for those focusing on association mapping based on strategy of haplotype block and htSNPs.

Algorithms↗

Dispersion of human Y chromosome haplotypes based on five microsatellites in global populations.

We have analyzed five microsatellite loci from the nonrecombining portion of the human Y chromosome in 15 diverse human populations to evaluate their usefulness in the reconstruction of human evolution and early male migrations. The results show that, in general, most populations have the same set of the most frequent alleles at these loci. Hypothetical ancestral haplotypes, reconstructed on the basis of these alleles and their close derivatives, are shared by multiple populations across racial and geographical boundaries. A network of the observed haplotypes is characterized by a lack of clustering of geographically proximal populations. In spite of this, few distinct clusters of closely related populations emerged in the network, which are associated with population-specific alleles. A tree based on allele frequencies also shows similar results. Lack of haplotypic structure associated with the presumed ancestral haplotypes consisting of individuals from almost all populations indicate a recent common ancestry and/or extensive male migration during human evolutionary history. The convergent nature of microsatellite mutation confounds population relationships. Optimum resolution of Y chromosome evolution will require the use of additional microsatellite loci and diallelic genetic markers with lower mutation rates.

Alleles↗