PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Haplotype structures”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article↗

Assessing the performance of the haplotype block model of linkage disequilibrium.

Several recent studies have suggested that linkage disequilibrium (LD) in the human genome has a fundamentally "blocklike" structure. However, thus far there has been little formal assessment of how well the haplotype block model captures the underlying structure of LD. Here we propose quantitative criteria for assessing how blocklike LD is and apply these criteria to both real and simulated data. Analyses of several large data sets indicate that real data show a partial fit to the haplotype block model; some regions conform quite well, whereas others do not. Some improvement could be obtained by genotyping higher marker densities but not by increasing the number of samples. Nonetheless, although the real data are only moderately blocklike, our simulations indicate that, under a model of uniform recombination, the structure of LD would actually fit the block model much less well. Simulations of a model in which much of the recombination occurs in narrow hotspots provide a much better fit to the observed patterns of LD, suggesting that there is extensive fine-scale variation in recombination rates across the human genome.

Computer Simulation↗

Marine population structure in an anadromous fish: life-history influences patterns of mitochondrial DNA variation in the eulachon, Thaleichthys pacificus.

Due to the apparent decline in size of a number of populations, eulachon, Thaleichthys pacificus, have recently become the focus of a conservation movement in the northeast Pacific. Little is known of the marine life-history phase of this anadromous fish, and although it has been suggested that eulachon spawning in different rivers may form distinct populations, nothing is known of their population structure. Molecular genetic data were used to investigate population structure and possible management schemes. Mitochondrial DNA genotypes, determined through restriction fragment length polymorphisms (RFLP) analysis, were resolved in fish from several rivers throughout the geographical range of eulachon. Our data support the idea that extant eulachon populations result from postglacial dispersal from a single Wisconsinan glacial refuge. Further, while three of the 37 haplotypes recovered account for approximately 79% of the samples, many private haplotypes were observed, suggesting possible regional population structure. While a great deal of genetic variation was observed (37 haplotypes in 315 samples), an AMOVA showed that > 97% of the total variation was detected within populations. As yet, it is unclear whether genetically distinct populations of eulachon exist, or if these fish may be treated as one or a few large populations. Results were tested against predictions made from hypotheses concerning the origin and persistence of subdivided populations in marine species, and seem to be more consistent with the Member-Vagrant hypothesis than isolation by distance. Eulachon present an interesting situation that illustrates the difficulties involved in defining management units in organisms with high levels of gene flow.

Animals↗

Siberian population of the New Stone Age: mtDNA haplotype diversity in the ancient population from the Ust'-Ida I burial ground, dated 4020-3210 BC by 14C.

On the basis of analysis of mtDNA from skeletal remains, dated by 14C 4020-3210 BC, from the Ust'-Ida I Neolithic burial ground in Cis-Baikal area of Siberia, we obtained genetic characteristics of the ancient Mongoloid population. Using the 7 restriction enzymes for the analysis of site's polymorphism in 16,106-16,545 region of mtDNA, we studied the structure of the most frequent DNA haplotypes, and estimated the intrapopulational nucleotide diversity of the Neolithic population. Comparison of the Neolithic and modern indigeneous populations from Siberia, Mongolia and Ural showed, that the ancient Siberian population is one of the ancestors of the modern population of Siberia. From genetic distance, in the assumption of constant nucleotide substitution rate, we estimated the divergence time between the Neolithic and the modern Siberian population. This divergence time (5572 years ago) is conformed to the age of skeletal remains (5542-5652 years). With use of the 14C dates of the skeletal remains, nucleotide substitution rate in mtDNA was estimated as 1% sequence divergence for 8938-9115 years.

Asian People↗

Evidence for positive selection and population structure at the human MAO-A gene.

We report the analysis of human nucleotide diversity at a genetic locus known to be involved in a behavioral phenotype, the monoamine oxidase A gene. Sequencing of five regions totaling 18.8 kb and spanning 90 kb of the monoamine oxidase A gene was carried out in 56 male individuals from seven different ethnogeographic groups. We uncovered 41 segregating sites, which formed 46 distinct haplotypes. A permutation test detected substantial population structure in these samples. Consistent with differentiation between populations, linkage disequilibrium is higher than expected under panmixia, with no evidence of a decay with distance. The extent of linkage disequilibrium is not typical of nuclear loci and suggests that the underlying population structure may have been accentuated by a selective sweep that fixed different haplotypes in different populations, or by local adaptation. In support of this suggestion, we find both a reduction in levels of diversity (as measured by a Hudson-Kreitman-Aguade test with the DMD44 locus) and an excess of high frequency-derived variants, as expected after a recent episode of positive selection.

Animals↗

DNA sequence analysis of the HLA-DRw12 allele.

The complete DNA sequence of a DR beta chain cDNA encoding the DRw12 allele has been determined. The sequence of this DRB1 allele reveals a structural relationship to the group of other DRB1 genes found on DRw52 haplotypes, such as DR3, -w11, -w13, -w14, and -w8. The structural similarities among this group of alleles are particularly evident in the first hypervariable as well as in the 3' untranslated region. The second hypervariable region contains a unique sequence not identified in any other DRB1 allele. The third hypervariable region appears to have arisen by gene conversion events involving two DRB1 chain genes, DR7 and DR1 or DR2/Dw21.

Alleles↗

Relationship between complement components C4A and C4B diversities and two TNFA promoter polymorphisms in two healthy Caucasian populations.

The RP-C4-CYP21-TNX (RCCX) modules and the tumor necrosis factor (TNF) gene cluster are probably the most polymorphic genomic regions in the human central major histocompatibility complex (MHC). Using definitive methods for genotypic and phenotypic analyses of complement components C4A and C4B, determination of the RCCX length variants, and SSP-PCR/RFLP analyses of TNFA promoter polymorphisms at positions -308 and -238, we studied the complex relationships between the C4 and TNFA polymorphisms in two normal Caucasian populations. The patterns of the RCCX modular structures and the allelic frequency of -308A TNFA (TNF2) were similar between the Budapest (n = 125) and the Ohio (n = 80) Caucasians. However, the frequency of the -238A allele was significantly higher in the Ohio (11.3%) than in the Budapest (1.6%) study population (p < 0.0001). Marked features were found in the RCCX length variants in the TNF2 carriers and noncarriers. Strong associations were found between the C4AQ0 B1 haplotype from the monomodular short (mono-S) RCCX structure and the TNF2 allele, and between the C4A6 B1 haplotype from the bimodular long-short (LS) structure of the RCCX and the TNFA -238A allele. However, 36%-46% of the TNF2 carriers did not associate with a mono-S in both study cohorts, and 57.1% of the TNFA -238A carriers in Ohio did not associate with C4A6, which has a defective complement C5 convertase activity. The carriers of TNF2 allele had significantly lower C4A serum concentration (0.17 +/- 0.08 g/l) than noncarriers (0.23 +/- 0.09 g/l) (p < 0.001). The lowest C4A serum levels were found in TNF2 carriers with mono-S structures (0.14 +/- 0.06 g/l). In essence, our results demonstrated the heterogeneities of the TNFA promoter polymorphisms, and the linkage disequilibrium of TNFA -308A and -238A alleles with complement C4A deficiency and impaired C4A protein function, respectively.

Blotting, Southern↗

Genetic structure of the star sea squirt, Botryllus schlosseri, introduced in southern European harbours.

The introduction of new genetic variants or species is often caused by maritime transport between harbours. Botryllus schlosseri is a cosmopolitan ascidian species that is found in both harbours and open shore habitats. In order to determine the influence of ship traffic on the genetic structure and phylogeography of B. schlosseri in southern Europe, we analyzed the variability of a fragment of the mitochondrial gene cytochrome c oxidase subunit I (COI). We sampled seven Atlanto-Mediterranean harbour populations and three open-shore populations. In addition, we sequenced some colonies from the US-Atlantic coast and from other Mediterranean localities to perform phylogenetic analyses. Although the number of polymorphic sites recorded (25.8%) was within the range observed in other population studies based on ascidian COI sequences, the haplotypic diversity (16 haplotypes out of 181 sequences) was much lower. Moreover, a lack of intermediate haplotypes was observed. This pattern of high nucleotide diversity and low haplotype diversity was consistent with introduction events of a few divergent haplotypes. We found a strong genetic structure in the study populations. Gene flow was only appreciable between some harbour populations. Harbour- and open-shore populations were well differentiated, although there was no evidence for isolation by distance. A nested clade analysis pointed to long-distance colonization, possibly coupled with subsequent fragmentation, as the underlying process. Our results suggest that B. schlosseri entered the study area via harbour-hopping, possibly through recurrent introduction events. The haplotypes from North America and most of the European ones were grouped in the same phylogenetic clade. This suggests occasional gene flow between both continents, probably through ship transport.

Animals↗

The H-2Kk MHC peptide-binding groove anchors the backbone of an octameric antigenic peptide in an unprecedented mode.

A wealth of data has accumulated on the structure of mouse MHC class I (MHCI) molecules encoded by the H-2(b) and H-2(d) haplotypes. In contrast, there is a dearth of structural data regarding H-2(k)-encoded molecules. Therefore, the structures of H-2K(k) complexed to an octameric peptide from influenza A virus (HA(259-266)) and to a nonameric peptide from SV40 (SV40(560-568)) have been determined by x-ray crystallography at 2.5 and 3.0 A resolutions, respectively. The structure of the H-2K(k)-HA(259-266) complex reveals that residues located on the floor of the peptide-binding groove contact directly the backbone of the octameric peptide and force it to lie deep within the H-2K(k) groove. This unprecedented mode of peptide binding occurs despite the presence of bulky residues in the middle of the floor of the H-2K(k) peptide-binding groove. As a result, the Calpha atoms of peptide residues P5 and P6 are more buried than the corresponding residues of H-2K(b)-bound octapeptides, making them even less accessible to TCR contact. When bound to H-2K(k), the backbone of the SV40(560-568) nonapeptide bulges out of the peptide-binding groove and adopts a conformation reminiscent of that observed for peptides bound to H-2L(d). This structural convergence occurs despite the totally different architectures of the H-2L(d) and H-2K(k) peptide-binding grooves. Therefore, these two H-2K(k)-peptide complexes provide insights into the mechanisms through which MHC polymorphism outside primary peptide pockets influences the conformation of the bound peptides and have implications for TCR recognition and vaccine design.

Animals↗

Bayesian analysis of haplotypes for linkage disequilibrium mapping.

Haplotype analysis of disease chromosomes can help identify probable historical recombination events and localize disease mutations. Most available analyses use only marginal and pairwise allele frequency information. We have developed a Bayesian framework that utilizes full haplotype information to overcome various complications such as multiple founders, unphased chromosomes, data contamination, and incomplete marker data. A stochastic model is used to describe the dependence structure among several variables characterizing the observed haplotypes, for example, the ancestral haplotypes and their ages, mutation rate, recombination events, and the location of the disease mutation. An efficient Markov chain Monte Carlo algorithm was developed for computing the estimates of the quantities of interest. The method is shown to perform well in both real data sets (cystic fibrosis data and Friedreich ataxia data) and simulated data sets. The program that implements the proposed method, BLADE, as well as the two real datasets, can be obtained from http://www.fas.harvard.edu/~junliu/TechRept/01folder/diseq_prog.tar.gz.

Bayes Theorem↗

Separating population structure from population history: a cladistic analysis of the geographical distribution of mitochondrial DNA haplotypes in the tiger salamander, Ambystoma tigrinum.

Nonrandom associations of alleles or haplotypes with geographical location can arise from restricted gene flow, historical events (fragmentation, range expansion, colonization), or any mixture of these factors. In this paper, we show how a nested cladistic analysis of geographical distances can be used to test the null hypothesis of no geographical association of haplotypes, test the hypothesis that significant associations are due to restricted gene flow, and identify patterns of significant association that are due to historical events. In this last case, criteria are given to discriminate among contiguous range expansion, long-distance colonization, and population fragmentation. The ability to make these discriminations depends critically upon an adequate geographical sampling design. These points are illustrated with a worked example: mitochondrial DNA haplotypes in the salamander Ambystoma tigrinum. For this example, prior information exists about restricted gene flow and likely historical events, and the nested cladistic analyses were completely concordant with this prior information. This concordance establishes the plausibility of this nested cladistic approach, but much future work will be necessary to demonstrate robustness and to explore the power and accuracy of this procedure.

Ambystoma↗

Natural variation in the PmbHLH162 promoter regulates anthocyanin biosynthesis and accumulation in Prunus mume.

Anthocyanin accumulation is a vital agronomic and ornamental trait, as it not only contributes to adaptation to environmental stress but also enhances ornamental value. In this study, a genome-wide association study (GWAS) was conducted using 328 accessions of mei (Prunus mume) to identify single-nucleotide polymorphisms (SNPs) associated with red pigmentation in petals, filaments, and xylem. Based on these significant SNPs, we defined 2 haplotypes (bHLH162hap1 and bHLH162hap2) and identified PmbHLH162, a bHLH transcription factor gene responsible for anthocyanin biosynthesis regulation. Transient silencing of PmbHLH162 in mei petals via Agrobacterium-mediated transformation resulted in significant color fading, whereas its overexpression dramatically elevated anthocyanin levels. Haplotype analysis showed that 2 promoter variants in bHLH162hap2 (Chr03_2669885 A/C and Chr03_2670272 A/G) alter the binding affinity of transcription factors PmWRKY18 and PmWRKY70. Stronger binding to the G/C alleles gave rise to higher PmbHLH162 expression in bHLH162hap2, thereby promoted red pigmentation in multiple tissues. By contrast, accessions carrying bHLH162hap1 displayed light/colorless phenotype without accumulation of red pigment. Furthermore, PmbHLH162 interacted respectively with PmMYC2, PmTT8, and PmEGL1 to form heterodimers, and markedly enhanced PmMYC2-mediated transcriptional activation of the anthocyanin biosynthetic structural genes PmCHS and PmANS. Geographic haplotype analysis revealed that bHLH162hap2 was predominantly enriched in high-latitude northern populations but was declining markedly at lower latitudes. Collectively, our study reveals the genetic and molecular basis underlying anthocyanin accumulation in mei and identifies a PmbHLH162-PmMYC2 regulatory module in which PmbHLH162 enhances PmMYC2-mediated activation of key anthocyanin biosynthetic genes. The additional interactions of PmbHLH162 with the MBW-associated bHLH factors PmTT8 and PmEGL1 further suggest potential crosstalk between this module and the canonical anthocyanin regulatory network.

Anthocyanins↗

Haplotype-resolved telomere-to-telomere genome assembly of Populus lasiocarpa unveils retrotransposon-driven centromere evolution.

Centromeres, essential for chromosome segregation, exhibit remarkable evolutionary dynamism in sequence composition and structural organization. Here, we report the first haplotype-resolved, telomere-to-telomere genome assembly of Populus lasiocarpa (PLAS) and precisely map all 38 functional centromeres through CENH3 ChIP-Seq. Unlike classical satellite-rich centromeres in model plants, PLAS centromeres lack abundant satellite arrays but are dominated by retrotransposons, particularly RLG and RIL elements, which form intricate nested TE arrays within the functional centromeric regions, disrupting their structural integrity and driving their evolution. Comparative analysis with P. trichocarpa reveals a conserved retrotransposon-dominated architecture, despite minimal sequence conservation. We propose a cyclic model of centromere evolution in which autonomous retrotransposons destabilize functional centromeres through epigenetic erosion, triggering neocentromere formation at pericentromeric sites enriched in transposable elements (TEs) and tandem repeats (TRs). These neocentromeres either succumb to recurrent retrotransposon invasions or stabilize through KARMA-mediated TR expansion, ultimately giving rise to satellite-rich centromeres. Our work redefines centromeres as dynamic, epigenetically plastic domains shaped by retrotransposon-TR antagonism, challenging the satellite-centric paradigm and offering novel insights into plant genome evolution.

Retroelements↗

The "Sardinian" HLA-A30,B18,DR3,DQw2 haplotype constantly lacks the 21-OHA and C4B genes. Is it an ancestral haplotype without duplication?

The C4 and 21-OH loci of the class III HLA have been studied by specific DNA probes and the restriction enzyme Taq 1 in 24 unrelated Sardinian individuals selected from completely HLA-typed families. All 24 individuals had the HLA extended haplotype A30,Cw5,B18, BfF1,DR3,DRw52,DQw2, named "Sardinian" in the present paper because of its frequency of 15% in the Sardinian population. Eighteen of these were homozygous for the entire haplotype, and six were heterozygous at the A locus and blank (or homozygous) at all the other loci. In all completely homozygous cells and in four heterozygous cells at the A locus, the restriction fragments of the 21-OHA (3.2 kb) and C4B (5.8 kb or 5.4 kb) genes were absent, and the fragments of the C4A (7.0 kb) and 21-OHB (3.7 kb) genes were present. It is suggested that the "Sardinian" haplotype is an ancestral haplotype without duplication of the C4 and 21-OH genes, practically always identical in its structure, also in unrelated individuals. The diversity of this haplotype in the class III region (about 30 kb less) may be at least partially responsible for its misalignment with most haplotypes, which have duplicated C4 and 21-OH genes, and therefore also for its decreased probability to recombine. This can help explain its high stability and frequency in the Sardinian population. The same conclusion can be suggested for the Caucasian extended haplotype A1,B8,DR3 that always seems to lack the C4A and 21-OHA genes.

Blotting, Southern↗

Familial anticardiolipin antibodies and C4 deficiency genotypes that coexist with MHC DQB1 risk factors.

OBJECTIVE: To investigate the familial basis of antiphospholipid antibodies by studying putative risk factors at the C4 and MHC class II loci. METHODS: Autoimmune diseases, anticardiolipin (aCL) and other autoantibodies were studied in 38 first and 2nd degree family members of 3 index cases selected for primary antiphospholipid syndrome (APS) and 33 controls. C4 protein phenotyping and restriction fragment length polymorphism analysis of C4 and MHC class II loci were performed. RESULTS: Nineteen family members (46%) had autoimmune diseases or autoantibodies; aCL were present in 10 family members, 4 of whom had primary APS. Each family had 2 or more subjects with aCL. Among 22 independent haplotypes in family members, there was a high frequency of C4A and C4B deficiency alleles (0.41 vs 0.18 in 66 controls, p = 0.03) and a strong trend toward an increase in MHC DQB1 putative risk factors that share the TRAELDT structural domain. This DQB1 structural domain was present in 4/5 different haplotypes that contained a C4B deficiency genotype; however, neither of 2 different haplotypes with a C4A deletion (one being a common ancestral haplotype) contained this DQB1 putative risk factor. Among the 10 family members who had aCL, 10/20 haplotypes contained a C4 deficiency genotype; moreover, the DQB1 putative risk factor was present in all 16 MHC haplotypes that did not contain a C4A deletion. CONCLUSION: In these families, expression of an autoimmunity trait as aCL antibody appears to be associated with the coexistence of C4 deficiency alleles with DQB1 alleles that contain the TRAELDT structural domain.

Adolescent↗

A sparse marker extension tree algorithm for selecting the best set of haplotype tagging single nucleotide polymorphisms.

Single nucleotide polymorphisms (SNPs) play a central role in the identification of susceptibility genes for common diseases. Recent empirical studies on human genome have revealed block-like structures, and each block contains a set of haplotype tagging SNPs (htSNPs) that capture a large fraction of the haplotype diversity. Herein, we present an innovative sparse marker extension tree (SMET) algorithm to select optimal htSNP set(s). SMET reduces the search space considerably (compared to full enumeration strategy), and therefore improves computing efficiency. We tested this algorithm on several datasets at three different genomic scales: (1) gene-wide (NOS3, CRP, IL6 PPARA, and TNF), (2) region-wide (a Whitehead Institute inflammatory bowel disease dataset and a UK Graves' disease dataset), and (3) chromosome-wide (chromosome 22) levels. SMET offers geneticists with greater flexibilities in SNP tagging than lossless methods with adjustable haplotype diversity coverage (phi). In simulation studies, we found that (1) an initial sample size of 50 individuals (100 chromosomes) or more is needed for htSNP selection; (2) the SNP tagging strategy is considerably more efficient when the underlying block structure is taken into account; and (3) htSNP sets at 80-90% phi are more cost-effective than the lossless sets in term of relative power, relative risk ratio estimation, and genotyping efforts. Our study suggests that the novel SMET algorithm is a valuable tool for association tests.

Algorithms↗

Genomic organization of the S-locus region of Brassica.

To gain some insights into the structure of the S-locus and the mechanisms that have kept its diversity, a 75-kb genomic fragment containing the self-incompatibility (S) locus region was isolated from the S12-haplotype of Brassica rapa and compared with those of other S-haplotypes. The region around the S determinant genes was highly polymorphic and filled with S-haplotype-specific intergenic sequences. The diverse genomic structure must contribute to the suppression of recombination at the S-locus.

Base Sequence↗

ALOX5AP gene and the PDE4D gene in a central European population of stroke patients.

BACKGROUND AND PURPOSE: Recent evidence has implicated the genes for 5-lipoxygenase activating protein (ALOX5AP) and phosphodiesterase 4D (PDE4D) as susceptibility genes for stroke in the Icelandic population. The aim of the present study was to explore the role of these genes in a central European population of stroke patients. METHODS: A total of 639 consecutive stroke patients and 736 unrelated population-based controls that had been matched for age and sex were examined using a case-control design. Twenty-two single-nucleotide polymorphisms (SNPs) covering ALOX5AP were genotyped. For PDE4D, microsatellite AC008818-1 and 12 SNPs, which tag all common haplotypes in previously identified linkage disequilibrium (LD) blocks, were analyzed. RESULTS: A nominally significant association with stroke was observed with several SNPs from ALOX5AP, including SNP SG13S114, which had been part of the Icelandic at-risk haplotype. Associations were stronger in males than in females, with SG13S114 (odds ratio, 1.24; 95% CI, 1.04 to 1.55; P=0.017) and SG13S100 (odds ratio, 1.26; 95% CI 1.03 to 1.54; P=0.024) showing the strongest associations. No significant associations were detected with single markers and haplotypes in PDE4D. The frequencies of single-marker alleles and haplotypes differed largely from those in the Icelandic population. CONCLUSIONS: The present study suggests that sequence variants in the ALOX5AP gene are significantly associated with stroke, particularly in males. Variants in the PDE4D gene are not a major risk factor for stroke in individuals from central Europe. Population differences in allele and haplotype frequencies as well as LD structure may contribute to the observed differences between populations.

3',5'-Cyclic-AMP Phosphodiesterases↗