PubMed HealthSearch

SEARCH · PubMed Health

Results for “Haplotype structures”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Branching plasticity and candidate gene-hormone networks associated with shade responses in soybean under relay strip intercropping.

BACKGROUND: Branching is a key determinant of high-yield plant architecture in soybean, particularly in maize- soybean relay strip intercropping where plants experience an "initially shaded-then fully illuminated" light regime. However, the genetic regulation of branching responses to shading remains poorly understood. METHODS: We evaluated 11 branching-related traits across 202 soybean accessions grown under monoculture (SS) and relay strip intercropping (RI). Branch number (BN), branching incidence (BI), and total branch length (TBL) were assessed together with stress tolerance indices (STI) and relative distance plasticity index (RDPI). Genome-wide association studies (GWAS) using mixed linear model (MLM) and three-variance-component MLM (3VmrMLM) were combined with haplotype and protein structural analyses to refine candidate genes. RESULTS: Based on Pearson correlation analysis of all 11 traits, BN, BI, and TBL measured before maize harvest showed the strongest and most consistent associations with branch seed weight within the corresponding cropping system (BSW_SS under SS and BSW_RI under RI), whereas other traits showed weaker or environment-dependent associations. Higher STI values calculated from these traits during the co-growth phase were negatively associated with BSW_RI, suggesting weaker compensatory recovery after light restoration in genotypes with more stable early branching patterns between SS and RI. In contrast, mediation analysis indicated that RDPI was positively associated with BSW_RI mainly through improved mature branching architecture (MB_index), which accounted for approximately 70% of the total positive effect. GWAS identified 57 and 74 significant QTNs using MLM and 3VmrMLM, respectively, and LD-window genes were filtered for exonic nonsynonymous or premature stop-codon variants, yielding 883 genes with putative functional variants. Two high-confidence genes emerged: Glyma.02G058600 (PP2C55), exhibiting shading-specific haplotype effects likely linked to GA-mediated branch-stem balance, and Glyma.02G059900 (DA1-related protein), showing stable effects across environments and implicated in ABA-mediated suppression of axillary meristems. CONCLUSIONS: These results provide insight into the genetic and physiological basis of soybean branching responses under relay strip intercropping, clarify that branching plasticity and relative shade tolerance represent distinct response dimensions in this system, and identify putative loci that may be useful for breeding soybean cultivars with improved shade adaptation and yield stability.

Glycine max

Mapping of the immune response genes in the major histocompatibility complex of the Rhesus monkey.

Interest in the Ir genes of rheus monkeys stems from their phylogenetic relationship to man and the extensive data already available on the major histocompatibility complex of the monkey. At least two independent dominant H-linked Ir genes have been identified in the rhesus. These genes control the ability of monkeys to respond to the random linear copolymer of glutamyl alanine (GA), or the dinitrophenyl conjugate of glutamyl lysine (DNP-GL). These synthetic polymers can elicit weak delayed-type skin reactions and strong humoral responses in some monkeys. In a series of unrelated monkeys phenotyped for the serologically defined RhL-A specificities of both segregant series, there were no correlations between any RhL-A specificity and responder status to the GA or DNP-GL polymers. However, segregation analysis of 21 rhesus families sired by 3 fathers indicated the capacity of the offspring to form antibodies was associated with genes coded for in the RhL-A complex. In three monkeys, verified recombination within the RhL-A complex between the genes coding for the serologically defined determinants (SD loci) and the gene(s) controlling the lymphocyte-activating determinants (Lad loci) responsible for mixed lymphocyte reactivity was established. In two of these monkeys the immune response genes controlling the DNP-GL response segregated with the Lad genes, while in the third case the Ir-GL gene segregated with the SD loci, tentatively localizing the Ir-GL gene between the SD and Lad loci. In addition, we have shown that genetically distinct genes control responsiveness to DNP-GL and GA. These genes were separated by recombination, thus one monkey inherited the Lad, Ir-GL, and SD loci from one paternal haplotype and by crossing over inherited the gene controlling GA responsiveness from the other paternal haplotype. The fine structure mapping of the RhL-A gene complex is compared with the H-2 and HL-A gene complexes. Several striking similarities were noted.

Alanine

Allelic variation and light-responsive regulation of FaMYB10-2 underlie tissue-specific anthocyanin accumulation in strawberry.

Anthocyanins critically determine fruit color, nutrition, and stress resilience in cultivated strawberry (Fragaria × ananassa), directly influencing consumer preference. Despite complex genetic and environmental regulation of their biosynthesis, the basis for tissue-specific pigmentation, notably the widespread occurrence of red skin and pale flesh, remains poorly understood. We integrated genomic, transcriptomic, and functional analyses across 200 cultivars to dissect receptacle pigmentation regulation. Approaches included FaMYB10-2 allele mining, promoter structural variant (SV) identification, expression profiling, regulatory interaction assays, and characterization of upstream light-responsive factors. FaMYB10-2 was identified as the key R2R3-MYB regulator of fruit anthocyanin biosynthesis. Alleles FaMYB10-2.2 and FaMYB10-2.3 encode truncated proteins retaining bHLH-binding capacity but lacking activation domains, functioning as dominant-negative repressors. A promoter SV 986 bp upstream of FaMYB10-2 was associated with reduced pale fruit due to cis-regulatory divergence. The SV (Alt) allele is prevalent in Asian cultivars, while the Ref allele is enriched in Western germplasm. Crucially, a light-responsive FaHYH-FaWRKY71 cascade activates FaMYB10-2 and structural genes haplotype-dependently, compensating for weak MYB activity in the skin. Our findings reveal a multilayered regulatory system integrating allelic variation, cis-regulatory divergence, and environmental signals, advancing anthocyanin understanding and providing engineering targets for polyploid crop color improvement.

Fragaria

Genetic epidemiology of beta-thalassemia in Sicily: do sequences 5' to the G gamma gene and 5' to the beta gene interact to enhance HbF expression in beta-thalassemia?

The present epidemiological study of the molecular characteristics of beta-thalassemia in Sicily was prompted by the disparate phenotypic expression (in clinical status and absolute HbF level) observed in two beta-thalassemic homozygotes who were also homozygous for the beta-like globin gene cluster haplotype III. We suspected that polymorphisms within haplotype III could be the cause for the discrepancy. Based on the association of particular conformations of the (AT)xT(y) motif (-540 5' to the beta gene) with milder forms of thalassemia and sickle cell anemia, 38 homozygous beta-thalassemia patients were studied to define their haplotypes, the -158 site 5' to the G gamma gene (linked to haplotype III) and the structure of the (AT)xT(y) motif. We found that the patient who was phenotypically mild and homozygous for beta-thalassemia, haplotype III, and the -158 C----T mutation was homozygous for the rare (AT)9T5 motif. In contrast, the patient homozygous for beta-thalassemia, haplotype III, and the -158 mutation, but exhibiting a severe clinical course, was homozygous for the (AT)7T7 configuration. Others have suggested that (AT)9T5 is a negative regulatory protein binding sequence, and it is a silent carrier state for beta-thalassemia. The usual configuration (AT)7T7, has considerably less affinity for regulatory protein binding, and it is the most common configuration in Sicilian beta-thalassemics (67 of the 78 chromosomes studied). Within the 38 patients studied, seven were informative because they had various combinations of the (AT)9T5 and (AT)7T7 motif, and the -158 C----T mutation. The results in these patients suggest that only the co-presence of the (AT)9T5 configuration and a C----T change at -158 5' to the G gamma gene is associated with high HbF expression and a mild clinical phenotype. We postulate that these two regions of the beta-like globin gene cluster interact, when endowed with the proper sequences, to enhance the expression of HbF secondary to anemia.

Adolescent

Polyploidy-mediated variations in glutamate receptor proteins linked to Fusarium wilt resistance in upland cotton.

Cotton production in the US faces a serious threat from Fusarium oxysporum f. sp. vasinfectum race 4 (FOV4), a soil-borne fungus causing Fusarium wilt by infecting the roots and vascular system of susceptible cotton, leading to rapid wilting and death. Here, we investigate genetic mechanisms of resistance to FOV4 in the highly resistant upland cotton genotype "U1" using an early-generation segregating biparental population ("U1" × "CSX8308") with comprehensive genomic resources. Reference-grade genomic assemblies of the parents revealed minor structural variations between "U1" haplotypes, a high degree of collinearity at chromosome synteny and micro-synteny levels, and significant divergence from "CSX8308" with 8.9 million SNPs. QTL analysis identified significant markers on chromosomes D03 and A02 linked to reduced Fusarium wilt severity. Within these regions, two glutamate-receptor-like (GLR) genes showed structural variation and overlapped between translocated segments on A02 and D03, suggesting a rare but important reinforcing effect of parallel evolution between susceptible and resistant genotypes. Transcriptome profiles of "U1" under FOV4 infection reveal activation of calcium-binding proteins and transcription factors regulating plant hormones (ethylene, abscisic acid, jasmonic acid, and salicylic acid), along with enzymes involved in cell wall remodeling and phytoalexin production. Advancing cotton improvement depends on incorporating durable genetic disease resistance into high-yielding, high-quality cultivars.

Fusarium

Rhesus blood group haplotype determination by nanopore sequencing and adaptive sampling enables the precise determination of complex allele combinations that could not be accurately determined by standard methods.

BACKGROUND: Patients with chronic transfusion needs such as those with sickle cell disease face a high risk of developing antibodies against high-prevalence antigens in the RH blood group system, complicating transfusion therapy and potentially necessitating stem cell transplantation. Molecular characterization of the RH system is hindered by hybrid alleles and high sequence homology between RHD and RHCE, limiting the effectiveness of conventional short-read sequencing. STUDY DESIGN AND METHODS: We analyzed 11 control and 20 patient samples, some of which could not be reliably genotyped by standard methods. RESULTS: Nanopore sequencing with adaptive sampling enables targeted, amplification-free long-read sequencing of the RH locus, resolving homologous and complex hybrid structures and enabling complete haplotype phasing for all samples, including samples that could not be accurately determined by standard methods like serology and short-read sequencing. Four new alleles were identified and for 13 out of 20 patients the results led to a change in the transfusion regimen. DISCUSSION: These findings show that nanopore sequencing with adaptive sampling allows unambiguous genotyping of the RH system, improves detection of complex variants, and supports better-matched transfusion strategies for chronically transfused patients.

Rh-Hr Blood-Group System

Shared chemical properties of different murine thymus-leukemia antigens.

Immunochemical studies of murine thymus-leukemia antigens (TLA) have confirmed that the subunit structure consists of a 45,000-dalton heavy chain and a beta 2 microglobulin (beta 2m) light chain. Similar structural features are exhibited by the TLA from thymocytes of Tlaa, Tlac, Tlad, and a leukemia cell derived from C57BL/6, a Tlab strain. In addition to the similar subunit structure from the four haplotypes, each TLA shows a similar pattern of trypsin proteolysis. This procedure yields a major heavy chain cleavage product of approximately 37,000 daltons that remains associated with beta 2m and retains most or all of the antigenic determinants of the intact TLA. Evidence is presented that TLA do not exhibit Fc receptor properties, nor do they adsorb to murine leukemia virus antigens under the conditions of isolation for analysis on polyacrylamide gel electrophoresis (PAGE) in sodium dodecyl sulfate (SDS). Taken together these findings strongly support the hypothesis that TLA comprise a family of chemically similar antigens belonging to a structurally and genetically related group that includes H-2D, H-2K, and Qa-2,3.

Animals

Comparison of haplotypes of the major histocompatibility complex in the rat. I. The Ag-B7 (H-1g) and Ag-B8 (H-1k) haplotypes.

Two haplotypes of the major histocompatibility complex of the rat, Ag-B7 and Ag-B8, have been compared with known H-1 haplotypes using the F1 skin-graft test and the dextran haemagglutination test. Both of these Ag-B haplotypes were different from the known H-1 haplotypes and determined different private specificities. The Ag-B7 haplotype was denoted as H-1g and the Ag-B8 haplotype as H-1k. The complex structure of the serologically detected antigenic products of these haplotypes was determined by means of H-1 congenic lines.

Alleles

Genetic control of antibody response to bovine rhodopsin in mice: epitope mapping of rhodopsin structure.

Inbred strains of mice of independent haplotype were immunized with bovine rhodopsin. All mice tested except SJL developed significant titers of specific antibodies 21 days after a single immunization. Anti-rhodopsin antibody level differed among conventional inbred strains. Comparison of the immune response to rhodopsin of congenic mice on two different genetic backgrounds showed that animals with an A background typically produced higher levels of specific antibody than mice with a B10 background. Titer of specific antibodies in antisera of mice of the same H-2 haplotype but different Igh haplotype differed; e.g. for H-2d haplotype, NZB (Ighn) generated the highest level of antibody with BALB/c (Igha), DBA/2 (Ighc), and B10.D2 (Ighb) strains giving successively lower responses. The location of immunodominant regions of bovine rhodopsin was investigated in primary sera among strains of mice. Sera were tested for their binding of anti-rhodopsin antibodies to synthetic peptides covering the entire primary structure of rhodopsin. From direct binding studies with hydrophilic rhodopsin peptides, the majority of the antigenic binding sites were localized in the sequence of the amino terminus, the II-III loop and the carboxyl terminus. Binding to these antigenic peptides was not strain restricted. Application of the overlapping synthetic peptide strategy of Geysen enabled refinement of these epitopes and determination of an additional major epitope in the hydrophobic sequence 304-310.

Amino Acid Sequence

A systematic strategy for identifying causal single nucleotide polymorphisms and their target genes on Juvenile arthritis risk haplotypes.

BACKGROUND: Although genome-wide association studies (GWAS) have identified multiple regions conferring genetic risk for juvenile idiopathic arthritis (JIA), we are still faced with the task of identifying the single nucleotide polymorphisms (SNPs) on the disease haplotypes that exert the biological effects that confer risk. Until we identify the risk-driving variants, identifying the genes influenced by these variants, and therefore translating genetic information to improved clinical care, will remain an insurmountable task. We used a function-based approach for identifying causal variant candidates and the target genes on JIA risk haplotypes. METHODS: We used a massively parallel reporter assay (MPRA) in myeloid K562 cells to query the effects of 5,226 SNPs in non-coding regions on JIA risk haplotypes for their ability to alter gene expression when compared to the common allele. The assay relies on 180 bp oligonucleotide reporters ("oligos") in which the allele of interest is flanked by its cognate genomic sequence. Barcodes were added randomly by PCR to each oligo to achieve > 20 barcodes per oligo to provide a quantitative read-out of gene expression for each allele. Assays were performed in both unstimulated K562 cells and cells stimulated overnight with interferon gamma (IFNg). As proof of concept, we then used CRISPRi to demonstrate the feasibility of identifying the genes regulated by enhancers harboring expression-altering SNPs. RESULTS: We identified 553 expression-altering SNPs in unstimulated K562 cells and an additional 490 in cells stimulated with IFNg. We further filtered the SNPs to identify those plausibly situated within functional chromatin, using open chromatin and H3K27ac ChIPseq peaks in unstimulated cells and open chromatin plus H3K4me1 in stimulated cells. These procedures yielded 42 unique SNPs (total = 84) for each set. Using CRISPRi, we demonstrated that enhancers harboring MPRA-screened variants in the TRAF1 and LNPEP/ERAP2 loci regulated multiple genes, suggesting complex influences of disease-driving variants. CONCLUSION: Using MPRA and CRISPRi, JIA risk haplotypes can be queried to identify plausible candidates for disease-driving variants. Once these candidate variants are identified, target genes can be identified using CRISPRi informed by the 3D chromatin structures that encompass the risk haplotypes.

Humans

Analysis of murine major histocompatibility complex class II-restricted T-cell responses to the flavivirus Kunjin by using vaccinia virus expression.

The present paper analyzes the influence of major histocompatibility complex (MHC) class II (Ir) genes on MHC class II-restricted T-cell responses to West Nile virus (WNV) and recombinant vaccinia virus-derived Kunjin virus antigens and identifies the immunodominant Kunjin virus antigens. Generally, mice were primed by intravenous infection with WNV or Kunjin virus, and their CD4+ T cells were stimulated in vitro 14 days later with WNV or Kunjin virus antigens to pulse macrophage or B-cell antigen-presenting cells (APC). WNV-specific in vitro T-cell responses from H-2b mice were higher than those from H-2d, H-2k, and H-2q mice. When recombinant vaccinia virus-derived Kunjin virus antigen preparations were tested in vitro, Kunjin virus-immune T cells of H-2b haplotype responded most strongly to structural (prM, C, E) and membrane-associated nonstructural (NS1) proteins encoded by VKV 1031 and showed weaker responses to cytosolic nonstructural protein NS5 (VKV 1022), whereas the responders of H-2k haplotype responded most strongly to the antigens encoded by VKV 1022 and gave lesser responses to VKV 1031. H-2d T cells gave weaker responses than either H-2b or H-2k cells, with responses to VKV 1031 generally being higher than those to VKV 1022. Responses to VKV 1023 or VKV 1024 encoding all of the NS3 to NS5 gene sequence or to VKV 1023 encoding all of NS3 were weak or absent. Within a given inbred strain, B cells and macrophages differed in their abilities to present recombinant vaccinia virus-derived Kunjin virus antigens, both in terms of magnitude of T-cell responses induced and the particular Kunjin virus protein presented. T cells from different non-MHC genetic backgrounds varied in their requirements of macrophage numbers as APC for maximum reactivity, suggesting that the concentration of class II MHC antigens and other molecules affecting APC-T-cell interaction varied in mice with different genetic backgrounds. Regardless of MHC haplotype, responses to VKV 1024, which encompasses VKV 1023 and VKV 1022, were either absent or lower than those to VKV 1022, possibly reflecting differences in the processing requirements of these two proteins. When mice were primed intravenously with recombinant vaccinia virus and when their CD4+ T cells were stimulated in vitro with native Kunjin virus antigens, VKV 1031 primed more efficiently than Kunjin virus and VKV 1022 primed similarly to Kunjin virus.

Animals

T cell determinant structure: cores and determinant envelopes in three mouse major histocompatibility complex haplotypes.

T lymphocytes recognize discrete regions on an antigen. The specificity of the T cell responses in three mouse strains of differing major histocompatibility complex (MHC) haplotype to a protein antigen, lysozyme, was analyzed using a series of peptides that walk the antigen in single amino acid steps. These peptide series were synthesized using the pin synthesis system, which was modified to allow the peptides to be cleaved from the pins into a physiological buffer free of toxic compounds. This methodology overcomes many of the problems associated with the production of peptides for screening proteins for antigenic determinants. The T cell determinants for the three strains were markedly different. This result points out the limitations of algorithms predicting determinants without reference to the MHC, and the importance of the empirical methodology. This analysis of the T cell response to lysozyme constitutes the most complete study of reactivity to a foreign protein to date and illustrates many important features of antigen recognition by T cells, e.g., presence of major and minor determinant regions. The outer boundaries of each immunogenic region, the determinant envelope, are difficult to define from recently immunized lymph nodes because of the heterogeneity in T cell recognition. However, core sequences common to all the immunogenic peptides in a continuous sequence can be easily defined.

Amino Acid Sequence

Hidden genomic structure and widespread structural polymorphism across environmental gradients in the spiny sea star Marthasterias glacialis.

Genomic regions of reduced recombination can preserve linkage among co-adapted alleles, facilitating local adaptation despite high connectivity. Such regions-often generated by chromosomal inversions-may be especially important in highly dispersive marine taxa yet remain poorly documented in echinoderms. Here, we combined a chromosome-level reference genome with genome-wide ddRAD-seq from 296 Marthasterias glacialis individuals across 19 Atlantic-Mediterranean locations to quantify population structure and scan for recombination-suppressed haploblocks. Genome-wide neutral markers showed significant population differentiation together with evidence of high connectivity, revealed by the presence of inter-ecoregion migrants. Additionally, we identified 16 polymorphic haploblocks with patterns consistent with putative chromosomal inversions spanning 18.6% of the genome. Haploblock haplotypes were strongly environmentally and geographically structured and contained genes with key functions in stress response, osmoregulation and thermal tolerance. Haplotype distributions also paralleled previously described mitochondrial lineages despite nuclear gene flow, consistent with a model of ancient divergence followed by secondary contact. Overall, our results suggest a role for widespread structural polymorphism in adaptive differentiation in Echinodermata, providing a framework for linking echinoderm genome rearrangements to ecological divergence. Marthasterias glacialis thus emerges as a promising system to explore how structural variation contributes to adaptation and genome evolution in highly dispersive organisms.

Animals

Haplotype-resolved genome assembly and implementation of VitExpress, an open interactive transcriptomic platform for grapevine.

Haplotype-resolved genome assemblies were produced for Chasselas and Ugni Blanc, two heterozygous Vitis vinifera cultivars by combining high-fidelity long-read sequencing and high-throughput chromosome conformation capture (Hi-C). The telomere-to-telomere full coverage of the chromosomes allowed us to assemble separately the two haplo-genomes of both cultivars and revealed structural variations between the two haplotypes of a given cultivar. The deletions/insertions, inversions, translocations, and duplications provide insight into the evolutionary history and parental relationship among grape varieties. Integration of de novo single long-read sequencing of full-length transcript isoforms (Iso-Seq) yielded a highly improved genome annotation. Given its higher contiguity, and the robustness of the IsoSeq-based annotation, the Chasselas assembly meets the standard to become the annotated reference genome for V. vinifera. Building on these resources, we developed VitExpress, an open interactive transcriptomic platform, that provides a genome browser and integrated web tools for expression profiling, and a set of statistical tools (StatTools) for the identification of highly correlated genes. Implementation of the correlation finder tool for MybA1, a major regulator of the anthocyanin pathway, identified candidate genes associated with anthocyanin metabolism, whose expression patterns were experimentally validated as discriminating between black and white grapes. These resources and innovative tools for mining genome-related data are anticipated to foster advances in several areas of grapevine research.

Vitis

Genome-wide analysis of FATA associated with drought tolerance in tetraploid potato (Solanum tuberosum).

The cuticle represents the outer most protective barrier against biotic and abiotic stresses. It is composed of cutin and waxes and protects plants from desiccation, UV, cold, mechanical stresses, and pathogens. GWAS/BSAseq combined with SeqSNP analyses in an association panel of 34 potato cultivars had revealed that the acyl-ACP thioesterase FATA (Soltu.DM.06G033680.1) is significantly associated with drought tolerance in potato. Apart from three FATB genes, only one FATA gene is present in potato that has the highest homology to FATA2 in Arabidopsis. FATA is responsible for the export of C18:1 fatty acid from chloroplast into cytosol, which is necessary for the biosynthesis of cutin. A knockout mutant of AtFATA2 was analyzed with regard to the cuticle permeability and to drought tolerance as well as recovery. Loss of FATA function leads to higher sensibility to water deficit in Arabidopsis, but to no change in recovery. The increased permeability of the cuticle in the fata2 knockout mutant as shown indirectly by higher chlorophyll leaching might play a role in this. Haplotypes for FATA were identified for the two potato cultivars Albatros and Désirée. All Désirée haplotypes and Albatros haplotypes 1, 3 and 4 were also revealed by former potato pan genome studies, while Albatros haplotype 2 is unique and has not been described before. Protein models were developed to investigate the influence of different SNPs in the haplotypes on the predicted protein structure and especially the substrate cavity. In potato, protein modeling suggests that only the hypothetical isoform B of FATA might be able to process oleoyl-ACP, but not hypothetical isoform A. However, this hypothesis needs to be verified by enzyme activity assays.

FATA

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article

Haplotype-resolved reconstruction and functional interrogation of cancer karyotypes.

Complex karyotype changes are widespread in cancer genomes. A major gap in cancer genome characterization is the resolution of rearranged chromosomes with chromosome-length continuity. Here, we describe a two-tiered approach to determine the segmental composition of rearranged chromosomes with haplotype resolution. First, we present refLinker, a bioinformatic method for robust determination of chromosomal haplotypes using cancer Hi-C data. By contrast with existing methods, refLinker is insensitive to the presence of large-scale DNA deletions, duplications, and high-level amplification in cancer genomes. Second, we demonstrate a computational strategy to determine the segmental structure of rearranged chromosomes using haplotype-specific Hi-C contacts. We apply these methods to breast cancer genomes and provide direct evidence for long-range transcriptional changes associated with rearrangements of the inactive X chromosome. Together, these results highlight refLinker's broad utility for studying the functional consequences of chromosomal rearrangements.

Humans

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article