PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Integrated multi-omics analyses provide new insights into genomic variation landscape and regulatory network candidate genes associated with walnut endocarp.

Persian walnut (Juglans regia) is an economically important nut oil tree; the fruit has a hard endocarp/shell to protect seeds, thus playing a key role in its evolution, and the shell thickness is an important trait for walnut breeding. However, the genomic landscape and the gene regulatory networks associated with walnut shell development remain to be systematically elucidated. Here, we report a high-quality genome assembly of the walnut cultivar 'Xiangling' and construct a graphic structure pan-genome of eight Juglans species to reveal the genetic variations at the genome level. We re-sequence 285 accessions to characterize the genomic variation landscape. Through genome-wide association studies (GWAS), we identified 19 loci associated with more than 268 loci that underwent selection during walnut domestication and improvement. Multi-omics analyses, including transcriptomics, metabolomics, DNA methylation, and spatial transcriptomics across eleven developmental stages, revealed several candidate genes related to secondary cell biosynthesis and lignin accumulation. This integrated multi-omics approach revealed several candidate genes associated with secondary cell biosynthesis and lignin accumulation, such as UGP, MYB308, MYB83, NAC043, NAC073, CCoAOMT1, CCoAOMT7, CHS2, CESA7, LAC7, COBL4, and IRX12. Overexpression of JrUGP and JrMYB308 in Arabidopsis thaliana confirmed their roles in lignin biosynthesis and cell wall thickening. Consequently, our comprehensive multi-omics findings offer novel insights into walnut genetic variation and network regulation of endocarp development and shell thickness, which enable further genome-informed breeding strategies for walnut cultivar improvement.

Juglans↗

Polygenic and monogenic adaptation drive evolutionary rescue at different magnitudes of environmental change.

Understanding the genetic basis of rapid adaptation is key to predicting species' evolutionary responses to environmental change. However, it is still debatable whether many small-effect mutations or a few large-effect mutations underlie rapid adaptation, and how this knowledge can predict population survival or extinction. To address this question, we performed a series of ecologically grounded forward-in-time genetic simulations to study rapid adaptation and extinction with increasing magnitudes of environmental change. These simulations were seeded with genomic variation of the plant Arabidopsis thaliana to have a realistic genomic structure, with one (monogenic) to 1,000 (polygenic) variants with varying heritabilities contributing to an environmental adaptive trait. Our results revealed two distinct scenarios of rapid adaptation and population rescue. Under small-to-moderate environmental shifts, high polygenic traits increased evolutionary rescue probability. Under extreme environmental shifts, high polygenic traits lead predictably to extinction, yet monogenic traits sometimes produce one-off winning adaptive genotypes. We interpret our rapid evolutionary rescue findings in terms of the fundamental theorem of natural selection, where trait polygenicity shapes the distribution of genetic variance in fitness across replicates and, in turn, the probability of population survival, with polygenic architectures producing more stable and predictable fitness variance and monogenic architectures generating highly skewed and variable outcomes. These results highlight the insights genomics gives us into the (un)predictability of species' evolutionary responses to global change, with management implications for assisted adaptation and conservation.

Arabidopsis↗

Conservation and variation of gene regulation in embryonic stem cells assessed by comparative genomics.

We have examined the gene structure and regulatory regions of octamer-binding transcription factor 3/4 (Oct 3/4), sex determining region Y box 2 (Sox2), signal transducer and activator of transcription 3 (Stat3), embryonal stem cell-specific gene 1 (ESG), Nanog homeobox (Nanog), and several other genes highly expressed in embryonic stem (ES) cells across different species. Our analysis showed that ES cell-expressed Ras (ERAS) was orthologous to a human pseudogene Harvey Ras (HRASP) and that the promoter and other regulatory sequences were highly divergent. No ortholog of (ES) cell-derived homeobox containing gene (Ehox) could be identified in human, and the closest paralogs PEPP gene subfamily 1 (PEPP1), PEPP2, and extraembryonic, spermatogenesis, homeobox 1 (Esx1) were not expressed by ES cells and shared little homology. The Sox2 promoter was the most conserved across species and the Oct3/4 promoter region showed significant homology particularly in the distal enhancer active in ES cells. Analysis suggested common and divergent pathways of regulation. Conserved Oct3/4 and Sox2 co-binding domains were identified in most ES expressed genes, highlighting the importance of this transcriptional pathway. Conserved fibroblast growth factor response element sites were identified in regulatory regions, suggesting a potential parallel pathway for regulation by FGFs. A central role of Stat3 activation in self-renewal and in a regulatory feedback loop was suggested by the identification of the conserved binding sites in most pathways. Although most pathways were evolutionarily conserved, promoters and genomic structure of the leukemia inhibitory factor (LIF) pathway components were divergent, likely explaining the differential requirement of LIF for human and rodent cells. Our analysis further suggested that the Nanog regulatory pathway was relatively independent of the LIF/Oct pathway and may interact with the Nodal/transforming growth factor-beta pathway. These results provide a framework for examining the current reported differences between rodent and human ES cells and define targets for future perturbation studies.

Amino Acid Sequence↗

Heavy chain joining region segments of the channel catfish. Genomic organization and phylogenetic implications.

The JH locus of the channel catfish has been characterized to determine the organization and structural diversity of JH segments. These analyses indicate that there are a total of nine JH segments tightly clustered within a region spanning about 2.2 kb. The JH locus is closely linked to the CH 1 domain of the expressed catfish H chain; the distance between the CH proximal JH segment (JH9) and the CH 1 domain is about 1.8 kb. Each JH segment has an upstream recombination sequence, which includes a T-rich nonamer, a 22- to 24-bp spacer, and a phylogenetically conserved heptamer. Each JH segment also has an open reading frame that encodes the conserved framework region 4 tryptophan (Trp-103) and terminates with a RNA donor splice site. The catfish JH locus contains an internal repetitive sequence region characterized by a short (183-188 bp) repeat that occurs sequentially five times. Strong sequence homology as well as the unified length of the repeated sequences indicate that JH segments JH3-JH7 probably arose as the result of a series of homologous but unequal crossover events. Sequence alignments of the duplicated JH segments indicates that there is diversity within the 5-11 nucleotides located immediately downstream from the heptamer, an observation which indicates that closely related JH segments can serve to enhance CDR3 diversity in the expressed H chain. Comparisons of the genomic JH sequences with different cDNA clones indicate that each JH segment is probably functional and that junctional diversity serves an important role in the generation of CDR3 diversity. In addition, single base differences observed in comparisons of JH-encoded regions indicate that there is probably somatic mutation or allelic variation of genomic JH segments. These studies suggest that the characteristic structure and organizational pattern of JH segments in higher vertebrates may have evolved early in vertebrate phylogeny at the level of the bony fish.

Amino Acid Sequence↗

Study of correlations in DNA sequences.

We present a method for unified statistical analysis of short and long range correlations between various nucleotides in genomic DNA strands. The approach is based on the mutual study of Fourier structure factor spectra and pair correlation functions. The analysis of cross correlations in the different ranges of structural spectra permits identification of the main sources of correlations, namely, the coherent point mutations, coincident periodicities or large scale density variations. The technique for assessment of structural coupling between various genes in the genome is also described.

Animals↗

Evolution and domestication-trait associations of ultra-long centromere haplotypes in pepper plants.

Centromeric and pericentromeric regions of most eukaryotic genomes are highly repetitive and strongly recombination-suppressed, confounding efforts to resolve genetic variation, population structure and phenotypic associations. Pepper (Capsicum annuum) centromeres are nearly devoid of satellite repeats, facilitating assembly and population-level comparison of centromeric regions. Here we integrate 9 near-complete genome assemblies, CENH3 ChIP-seq profiles from 26 diverse accessions, and resequencing and phenotypic data from ~400 cultivated and wild accessions to investigate population-level diversity and phenotypic relevance of pepper peri/centromeric regions. Functional centromere positions are largely fixed on 8 of 12 chromosomes, whereas the remaining 4 carry distinct centromeric epialleles shaped mainly by centromere repositioning and pericentromeric inversions. Pepper centromeres are embedded within ultra-long centromere-spanning haplotype (cenhap) blocks, ranging from 29.8 to 112.9 Mb and collectively covering 23.96% of the genome; each block contains only 1-4 major haplotypes. Some cenhaps may act as supergene-like units and are strongly associated with fruit traits, probably because recombination-suppressed intervals harbour multiple fruit-related genes, including OFP and F-box genes. F2 segregation assays further reveal transmission distortion of chromosomes carrying alternative cenhaps. Together, these findings highlight peri/centromeric regions as underrecognized reservoirs of agronomically important variation.

Centromere↗

Ultrastructure meets reproductive success: performance of a sphecid wasp is correlated with the fine structure of the flight-muscle mitochondria.

Organisms show a remarkable inter-individual variation in reproductive success. The proximate causes of this variation are not well understood. We hypothesized that the ultrastructure of costly or complex tissues or organelles might affect reproductive performance. We tested this hypothesis in females of a sphecid wasp, the European beewolf, Philanthus triangulum (Hymenoptera, Sphecidae), that show considerable variation in reproductive success. The most critical component of reproduction in beewolf females is flying with paralysed honeybees, which more than double their weight. Because of the high energetic requirements for flight, we predicted that the ultrastructure of the flight-muscle mitochondria might influence female success. We determined the density of mitochondria and the density of the inner mitochondrial membranes (DIMM) of the flight muscles as well as age, body size and fat content. Only DIMM had a significant influence on female reproductive success, which might be mediated by an elevated adenosine triphosphate (ATP) supply. The variation in DIMM might result from differences in larval provisions or from an accumulation of mutations in the mitochondrial genome. Our results support the hypothesis that the organization of complex structures contributes to inter-individual variation in reproductive success.

Adenosine Triphosphate↗

Widespread heteroplasmy in schistosomes makes an mtVNTR marker "nearsighted".

Mitochondrial markers are often hailed as the preferred DNA elements for analyses of population subdivision. To this end we have employed a mitochondrial repeat element to examine the population structure in Schistosoma mansoni (human blood flukes). Schistosome isolates were collected from each of 21 different patients representing seven different areas of a Brazilian village. These parasite isolates demonstrate substantial genetic polymorphism, with an average of 10 genotypes infecting each patient, which is more readily detected because of high levels of heteroplasmy (i.e., 72.5% of the individual worms exhibit multiple versions of this repeat region with different numbers of repeats). Due to the high number of common haplotypes in the population, this repeat element from S. mansoni has a large proportion (47%) of its genetic variation described by differences among mitochondrial genomes within individual worms. However, when only rare haplotypes are considered, population structure can be detected. It seems that heteroplasmy in the schistosome population of Melquiades is both the source of plentiful genetic variation and a confounding factor in the analysis of that variation. Thus the schistosome population in Melquiades may actually be more strongly subdivided than we are able to detect using this mitochondrial marker.

Animals↗

High expression of UDP-N-acetylglucosamine: beta-D mannoside beta-1,4-N-acetylglucosaminyltransferase III (GnT-III) in chronic myelogenous leukemia in blast crisis.

The activity and mRNA expression of UDP-N-acetylglucosamine: beta-D mannoside beta-1,4-N-acetylglucosaminyl transferase III (GnT-III: EC 2.4.1.144) were investigated in hematological malignancies. GnT-III activity was elevated in patients with chronic myelogenous leukemia (CML) in blast crisis and patients with multiple myeloma (MM), as compared to normal healthy subjects and patients with other hematological malignancies including CML in chronic phase. The GnT-III transcript was the same size in leukemic cells from various hematological diseases and cell lines, while expression of the transcript was not found to correlate significantly with enzyme activity, implying that post-translational modification might regulate the activity of GnT-III. Southern-blot analysis showed no significant variation in the structure and position of the GnT-III genome, indicating that the gene is present as a single copy without isoforms. Furthermore, analyses by immunoprecipitation and Western blot revealed that high GnT-III activity in KU812 cell, a CML cell line, resulted in an increase in E4-PHA binding to CD45, a major surface glycoprotein of the leukocyte, indicating that more bisecting GlcNAc was added to CD45 catalyzed by elevated GnT-III.

Acetylglucosamine↗

Innovation from reduction: gene loss, domain loss and sequence divergence in genome evolution.

Analyses of genome sequences have revealed a surprisingly variable distribution of genes, reflecting the generation of novel genes, lateral gene transfer and gene loss. The impact of gene loss on organisms has been difficult to examine, but the loss of protein coding genes, the loss of domains within proteins and the divergence of genes have made surprising contributions to the differences among organisms. This paper reviews surveys of gene loss and divergence in fungal and archaeal genomes that indicate suites of functionally related genes tend to undergo loss and divergence. Instances of fungal gene loss highlighted here suggest that specific cellular systems have changed, such as Ca 2+ biology in Saccharomyces cerevisiae and peroxisome function in Schizosaccharomyces pombe. Analyses of loss and divergence can provide specific predictions regarding protein-protein interactions, and the relationship between networks of protein interactions and loss may form a part of a parametric model of genome evolution.

Chromosome Mapping↗

Chromosomal inversion polymorphism leads to extensive genetic structure: a multilocus survey in Drosophila subobscura.

The adaptive character of inversion polymorphism in Drosophila subobscura is well established. The O(ST) and O(3+4) chromosomal arrangements of this species differ by two overlapping inversions that arose independently on O(3) chromosomes. Nucleotide variation in eight gene regions distributed along inversion O(3) was analyzed in 14 O(ST) and 14 O(3+4) lines. Levels of variation within arrangements were quite similar along the inversion. In addition, we detected (i) extensive genetic differentiation between arrangements in all regions, regardless of their distance to the inversion breakpoints; (ii) strong association between nucleotide variants and chromosomal arrangements; and (iii) high levels of linkage disequilibrium in intralocus and also in interlocus comparisons, extending over distances as great as approximately 4 Mb. These results are not consistent with the higher genetic exchange between chromosomal arrangements expected in the central part of an inversion from double-crossover events. Hence, double crossovers were not produced or, alternatively, recombinant chromosomes were eliminated by natural selection to maintain coadapted gene complexes. If the strong genetic differentiation detected along O(3) extends to other inversions, nucleotide variation would be highly structured not only in D. subobscura, but also in the genome of other species with a rich chromosomal polymorphism.

Animals↗

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Bioinformatics↗

Structural variation in the Waxy gene and differentiation in foxtail millet [Setaria italica (L.) P. Beauv.]: implications for multiple origins of the waxy phenotype.

The origin and evolution of the waxy type of foxtail millet [Setaria italica (L.) P. Beauv] were studied by analyzing structural variation in the Waxy gene. Initially, the Waxy gene was amplified by RT-PCR, RACE and genomic PCR from a non-waxy strain to determine the structure of the wild-type gene. Secondly, we screened by PCR for polymorphisms at the Waxy locus in 79 strains with various waxy phenotypes. We then carried out genomic Southern analysis on 67 strains and identified seven RFLP classes which were designated as types I-VII. RFLP type was correlated with phenotype, such that types I and II corresponded to non-waxy, types III and VI to low-amylose, and types IV, V and VII to waxy phenotypes. The differences between RFLP types could be attributed to insertions in the Waxy gene. Types II and VI were caused by the insertion of a Tourist element into intron 1 and a SINE-like sequence into intron 12, respectively. Types III, IV, V and VII were characterized by the insertion of large sequences into the Waxy gene that may alter the expression of the gene. Thus, multiple, independent insertions in the Waxy gene appear to have caused the loss-of-function waxy phenotypes. Furthermore, the geographical distributions of the three RFLP types associated with the waxy phenotype (types IV, V and VII) were distinct, with type IV being found mainly in Taiwan and Japan, type V in Korea, and type VII in Myanmar. These results indicate a polyphyletic origin for the waxy phenotype in landraces of foxtail millet.

Base Sequence↗

A family of retrotransposons and associated genomic variation in wheat.

A family of related retroelements was characterized in the genomes of some Graminease species. The structure of these retroelements indicates that they are retrotransposons containing reading frames with sequence similarity to the polyproteins of copia and Ty. This family of retroelements (termed WIS-2) occurs in the genomes of barley, wheat, rye, oats, and Aegilops species. Ongoing genomic variation both within individual plants of a wheat variety and within and between varieties of wheat is associated with some members of the WIS-2 family.

Amino Acid Sequence↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

Haplotype and linkage disequilibrium architecture for human cancer-associated genes.

To facilitate association-based linkage studies we have studied the linkage disequilibrium (LD) and haplotype architecture around five genes of interest for cancer risk: ATM, BRCA1, BRCA2, RAD51, and TP53. Single nucleotide polymorphisms (SNPs) were identified and used to construct haplotypes that span 93-200 kb per locus with an average SNP density of 12 kb. These markers were genotyped in four ethnically defined populations that contained 48 each of African Americans, Asian Americans, Hispanic Americans, and European Americans. Haplotypes were inferred using an expectation maximization (EM) algorithm, and the data were analyzed using D', R(2), Fisher's exact P-values, and the four-gamete test for recombination. LD levels varied widely between loci from continuously high LD across 200 kb to a virtual absence of LD across a similar length of genome. LD structure also varied at each gene and between populations studied. This variation indicates that the success of linkage-based studies will require a precise description of LD at each locus and in each population to be studied. One striking consistency between genes was that at each locus a modest number of haplotypes present in each population accounted for a high fraction of the total number of chromosomes. We conclude that each locus has its own genomic profile with regard to LD, and despite this there is the widespread trend of relatively low haplotype diversity. As a result, a low marker density should be adequate to identify haplotypes that represent the common variation at a locus, thereby decreasing costs and increasing efficacy of association studies.

Alleles↗

Polymorphic variations in the ori sequences from the mitochondrial genomes of different wild-type yeast strains.

We determined the restriction maps and primary structures of two as yet poorly characterized regions of the mitochondrial genomes of different wild-type strains of Saccharomyces cerevisiae. These regions respectively comprised the ori1 sequence and the newly identified ori8 sequence. Ori1 and ori8, together with their flanking sequences, exhibit a large polymorphism, resulting from specific variations due to insertions or deletions of optional GC clusters at different locations. The mechanisms underlying such sequence rearrangements are discussed.

Base Sequence↗