PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Symmetrical arrangement of the heterologous regions of rabbit poxvirus and vaccinia virus DNA.

Cleavage sites for the restriction endonucleases EcoRI, KpnI and XhoI were mapped on rabbit poxvirus and vaccinia virus DNA. These physical maps were used to analyse the structural variations between the two DNAs. Two specific heterologous regions, symmetrically arranged at each end of the genomes, have been identified. Region 1, representing the exterior part of the terminal repetition, appears to contain unrelated sequences in each DNA and accounts for the difference in length of the two genomes. Region 2, separated from region 1 by a conserved part of the terminal repetition, is located at the transition from repeated to unique DNA sequences. Its overall length of about 4 megadaltons is well conserved and it contains individual DNA-specific as well as conserved restriction sites. The major central part of the genomes (over 100 megadaltons) contains very few, widely dispersed restriction site variations.

Base Sequence↗

Surprisingly frequent chromosomal instability in cultivated peanut.

This study, the third in a three-part series, investigates whether chromosomal instability persists in cultivated peanut. The allotetraploid peanut (Arachis hypogaea; genome type AABB) originated from the hybridization and polyploidization of A. duranensis (AA) and A. ipaënsis (BB). Our first study established that this was an extremely narrow genetic origin, likely from a single hybridization event. This raised a paradox: how did such narrow genetics give rise to the phenotypic diversity seen in cultivated peanut? The second study addressed this, showing that a single neoallotetraploid spontaneously generates striking diversity, and that homoeologous exchanges-abundant in early generations following polyploidy-are a key mechanism in creating this diversity. In contrast to this early-generation instability, cultivated peanut is generally considered to be genetically stable, presumably due to selection. This third study tests whether residual instability still occurs in modern peanut. From a single plant of the highly selfed 'genome stock' of the cultivar 'Tifrunner', we advanced lineages through seven generations in a pollinator-free greenhouse. Among 233 plants, we identified three new large-scale chromosomal instability events: a large deletion on chromosome B01, associated with reduced pod width and seed weight, and two ABBB compositions involving chromosomes A02/B02 and A05/B05. With these observations in hand, we reinterpreted previously published data from two recombinant inbred populations. Together, these results indicate that at least 1% of pure pedigree A. hypogaea plants exhibit spontaneous large-scale chromosomal changes-a surprising frequency of instability that likely contributes to peanut's long-term adaptability and evolution.

Arachis↗

Evidence of structural genomic region recombination in Hepatitis C virus.

BACKGROUND/AIM: Hepatitis C virus (HCV) has been the subject of intense research and clinical investigation as its major role in human disease has emerged. Although homologous recombination has been demonstrated in many members of the family Flaviviridae, to which HCV belongs, there have been few studies reporting recombination on natural populations of HCV. Recombination break-points have been identified in non structural proteins of the HCV genome. Given the implications that recombination has for RNA virus evolution, it is clearly important to determine the extent to which recombination plays a role in HCV evolution. In order to gain insight into these matters, we have performed a phylogenetic analysis of 89 full-length HCV strains from all types and sub-types, isolated all over the world, in order to detect possible recombination events. METHOD: Putative recombinant sequences were identified with the use of SimPlot program. Recombination events were confirmed by bootscaning, using putative recombinant sequence as a query. RESULTS: Two crossing over events were identified in the E1/E2 structural region of an intra-typic (1a/1c) recombinant strain. CONCLUSION: Only one of 89 full-length strains studied resulted to be a recombinant HCV strain, revealing that homologous recombination does not play an extensive roll in HCV evolution. Nevertheless, this mechanism can not be denied as a source for generating genetic diversity in natural populations of HCV, since a new intra-typic recombinant strain was found. Moreover, the recombination break-points were found in the structural region of the HCV genome.

Evolution, Molecular↗

Linkage disequilibrium testing when linkage phase is unknown.

Linkage disequilibrium, the nonrandom association of alleles from different loci, can provide valuable information on the structure of haplotypes in the human genome and is often the basis for evaluating the association of genomic variation with human traits among unrelated subjects. But, linkage phase of genetic markers measured on unrelated subjects is typically unknown, and so measurement of linkage disequilibrium, and testing whether it differs significantly from the null value of zero, requires statistical methods that can account for the ambiguity of unobserved haplotypes. A common method to test whether linkage disequilibrium differs significantly from zero is the likelihood-ratio statistic, which assumes Hardy-Weinberg equilibrium of the marker phenotype proportions. We show, by simulations, that this approach can be grossly biased, with either extremely conservative or liberal type I error rates. In contrast, we use simulations to show that a composite statistic, proposed by Weir and Cockerham, maintains the correct type I error rates, and, when comparisons are appropriate, has similar power as the likelihood-ratio statistic. We extend the composite statistic to allow for more than two alleles per locus, providing a global composite statistic, which is a strong competitor to the usual likelihood-ratio statistic.

Alleles↗

Polymorphic simple sequence repeat regions in chloroplast genomes: applications to the population genetics of pines.

Simple sequence repeats (SSRs), consisting of tandemly repeated multiple copies of mono-, di-, tri-, or tetranucleotide motifs, are ubiquitous in eukaryotic genomes and are frequently used as genetic markers, taking advantage of their length polymorphism. We have examined the polymorphism of such sequences in the chloroplast genomes of plants, by using a PCR-based assay. GenBank searches identified the presence of several (dA)n.(dT)n mononucleotide stretches in chloroplast genomes. A chloroplast (cp) SSR was identified in three pine species (Pinus contorta, Pinus sylvestris, and Pinus thunbergii) 312 bp upstream of the psbA gene. DNA amplification of this repeated region from 11 pine species identified nine length variants. The polymorphic amplified fragments were isolated and the DNA sequence was determined, confirming that the length polymorphism was caused by variation in the length of the repeated region. In the pines, the chloroplast genome is transmitted through pollen and this PCR assay may be used to monitor gene flow in this genus. Analysis of 305 individuals from seven populations of Pinus leucodermis Ant. revealed the presence of four variants with intrapopulational diversities ranging from 0.000 to 0.629 and an average of 0.320. Restriction fragment length polymorphism analysis of cpDNA on the same populations previously failed to detect any variation. Population subdivision based on cpSSR was higher (Gst = 0.22, where Gst is coefficient of gene differentiation) than that revealed in a previous isozyme study (Gst = 0.05). We anticipate that SSR loci within the chloroplast genome should provide a highly informative assay for the analysis of the genetic structure of plant populations.

Base Sequence↗

[Structural changes of 4V chromosome of Haynaldia villosa induced by gametocidal chromosome 3C of Aegilops triuncialis].

Chromosome 3C of Aegilops triuncialis was discovered with ability to be transferred preferentially in the case of its monosomic status in wheat background, whereas, those gametes without 3C would result in chromosome structural changes including deletions and translocations. In the present study, Triticum aestivum-Haynaldia villosa substitution line 4V(4D) developed in our laboratory, was crossed to T. aestivum c.v. Norin 26-Aegilops triuncialis 3C addition line, and the hybrids F1 were then backcrossed with common wheat in order to induce structural changes of 4V. Both chromosome C-banding and genomic in situ hybridization was applied to search such chromosome variations. In this case, total genomic DNA of Haynaldia villosa was labelled by Biotin-11-dUTP as probes and total genomic DNA of Chinese Spring as the block. Moreover, several chromosome changes within common wheat such as isochromosome 1BL.1BL(B39-2) and others were also revealed. The result indicated that two translocation lines T4VL.3AS(A47-10-3) and T4VS.4DL(A47-25-4), two telocentric chromosome lines A47-7-2(4VS) and A47-32-2(4VL), and two isochromosomes including 4VS.4VS(A47-23) and 4VL.4VL(A412-5-4) were identified from BC1F2 or BC1F3. This result indicated that gametocidal chromosome 3C of Aegilop triuncialis could effectively induce structural changes of both chromosome 4V of Haynaldia villosa and chromosomes of wheat.

Chromosome Aberrations↗

Cloning and characterization of the endogenous retroviral-tRNA(Glu) multigene family from human genomes of different racial backgrounds.

An 8.3-kb human endogenous retroviral-tRNA(Glu) (HERV-E)-encoding cDNA clone and a 1.5-kb genomic clone were isolated from a Chinese-derived cervical cancer cell line, CC7T, and their sequences determined. The former is a full-length endogenous retroviral cDNA containing corresponding u5-gag-pol-env-u3-r regions. The latter is a partial retroviral DNA segment, covering the gag and pol genes. Analysis of normal human DNA by Southern blot hybridization with three specific HERV-E molecular DNA probes revealed complex restriction-fragment length polymorphisms (RFLP), implying that the human genome contains diverse proviral structures and dispersed integration sites. The complex patterns were virtually identical between DNAs from African-Americans, Asians and Caucasians, with only a few minor variations. The data suggest that these proviral sequences were mostly incorporated into the human genome before racial divergence and, hence, may serve as markers for distinct chromosomal sites.

Amino Acid Sequence↗

Genome-wide intraspecific DNA-sequence variations in rice.

Genome-wide comparative analysis of the DNA sequences of two major cultivated rice subspecies, Oryza sativa L. ssp indica and Oryza sativa L. ssp japonica, have revealed their extensive microcolinearity in gene order and content. However, deviations from colinearity are frequent owing to insertions or deletions. Intraspecific sequence polymorphisms commonly occur in both coding and non-coding regions. These variations often affect gene structures and may contribute to intraspecific phenotypic adaptations.

Genetic Variation↗

Multi-locus allelic architecture underlying natural variation in leaf rolling in japonica rice.

Leaf rolling is a key component of rice canopy architecture that affects light interception, microclimate formation, and planting density. The contribution of naturally occurring allelic variation to quantitative variation in leaf rolling within cultivated rice remains poorly understood, while extreme leaf rolling caused by loss-of-function mutations often results in detrimental pleiotropic effects. Herein, we examined how multi-locus allelic variation contributes to natural variation in leaf rolling within japonica rice. Leaf rolling was quantified based on the leaf rolling index (LRI) using a panel of 201 japonica accessions. The phenotype was transformed using the Yeo-Johnson method to reduce strong right skewness and improve the distributional properties of the data, thereby facilitating subsequent regression modeling. Haplotype analyses were performed for previously reported leaf rolling-associated genes and genome-wide association study (GWAS) lead loci, leading to the identification of five loci exhibiting substantial haplotype-dependent phenotypic variation. Phenotypically defined allelic groups represented these loci were subsequently evaluated using multiple linear regression (MLR), with the first two principal components derived from genome-wide SNP data included as covariates to account for population structure. The final MLR model identified four loci (qALR1, OsYABBY1, OsSLL2, and OsSRL10) as the independent contributors to leaf rolling variation, collectively explaining 21% of the variance in the transformed phenotype after accounting for population structure. Model diagnostics and ten-fold cross-validation supported the statistical validity of the framework and indicated stable model performance across validation folds. Analysis of multi-locus allelic combinations showed 13 distinct configurations that clustered into three phenotypically differentiated groups. This reflected the cumulative dosage of high-leaf rolling alleles. Thus, the natural variation in leaf rolling in japonica rice is governed by the additive effects of multiple moderate-impact loci. The multi-locus allelic framework established here provides a statistically sound and biologically interpretable basis for dissecting polygenic canopy traits and practical guidance for developing genetic materials aimed at optimizing rice plant architecture.

cross-validation↗

Complete sequence of two tick-borne flaviviruses isolated from Siberia and the UK: analysis and significance of the 5' and 3'-UTRs.

The complete nucleotide sequence of two tick-transmitted flaviviruses, Vasilchenko (Vs) from Siberia and louping ill (LI) from the UK, have been determined. The genomes were respectively, 10928 and 10871 nucleotides (nt) in length. The coding strategy and functional protein sequence motifs of tick-borne flaviviruses are presented in both Vs and LI viruses. The phylogenies based on maximum likelihood, maximum parsimony and distance analysis of the polyproteins, identified Vs virus as a member of the tick-borne encephalitis virus subgroup within the tick-borne serocomplex, genus Flavivirus, family Flaviviridae. Comparative alignment of the 3'-untranslated regions revealed deletions of different lengths essentially at the same position downstream of the stop codon for all tick-borne viruses. Two direct 27 nucleotide repeats at the 3'-end were found only for Vs and LI virus. Immediately following the deletions a region of 332-334 nt with relatively conserved primary structure (67-94% identity) was observed at the 3'-non-coding end of the virus genome. Pairwise comparisons of the nucleotide sequence data revealed similar levels of variation between the coding region, and the 5' and 3'-termini of the genome, implying an equivalent strong selective control for translated and untranslated regions. Indeed the predicted folding of the 5' and 3'-untranslated regions revealed patterns of stem and loop structures conserved for all tick-borne flaviviruses suggesting a purifying selection for preservation of essential RNA secondary structures which could be involved in translational control and replication. The possible implications of these findings are discussed.

Animals↗

The PHYTOCHROME C photoreceptor gene mediates natural variation in flowering and growth responses of Arabidopsis thaliana.

Light has an important role in modulating seedling growth and flowering time. We show that allelic variation at the PHYTOCHROME C (PHYC) photoreceptor locus affects both traits in natural populations of A. thaliana. Two functionally distinct PHYC haplotype groups are distributed in a latitudinal cline dependent on FRIGIDA, a locus that together with FLOWERING LOCUS C explains a large portion of the variation in A. thaliana flowering time. In a genome-wide scan for association of 65 loci with latitude, there was an excess of significant P values, indicative of population structure. Nevertheless, PHYC was the most strongly associated locus across 163 strains, suggesting that PHYC alleles are under diversifying selection in A. thaliana. Our work, together with previous findings, suggests that photoreceptor genes are major agents of natural variation in plant flowering and growth response.

Arabidopsis↗

The biological improbability of a clone.

Empirical evidence for intraclonal genetic variation is described here for clonal systems using a variety of molecular techniques and implicating a diversity of mechanisms. However, clonal systems are still generally perceived as having strict genetic fidelity. As concepts of genetic variability move from primary sequence data to include epigenetic and structural influences on genetic expression, the ability to detect changes in the genome at short intervals allows precedence to be given to inherent biological variation that is often analytically ignored. Therefore, the advent of powerful molecular techniques, like genome mapping, mean that our concepts of genetic fidelity within eukaryotic clones and the whole philosophy of the 'clone' needs to be re-evaluated and redefined to replace old unproven dogma in this aspect of science.

Animals↗

Regional base composition variation along yeast chromosome III: evolution of chromosome primary structure.

The recent determination of the complete sequence of chromosome III from the yeast Saccharomyces cerevisiae allows, for the first time, the investigation of the long range primary structure of a eukaryotic chromosome. We have found that, against a background G+C level of about 35%, there are two regions (one in each chromosome arm) in which G+C values rise to over 50%. This effect is seen in silent sites within genes, but not in noncoding intergenic sequences. The variation in G+C content is not related to differential selection of synonymous codons, and probably reflects mutational biases. That the intergenic regions do not exhibit the same phenomenon is particularly interesting, and suggests that they are under substantial constraint. The yeast chromosome may be a model of the structure of the human genome, since there is evidence that it is also a mosaic of long regions of different base compositions, reflected in wide variation of G+C content at silent sites among genes. Two possible causes of this regional effect, replication timing, and recombination frequency, are discussed.

Animals↗

Guidelines and recommendations for content, structure, and deployment of mutation databases.

These Guidelines recognize the need for annotated online mutation databases documenting allelic variation (both pathogenic and phenotype modifying, and also neutral polymorphic); the databases will be both generalized (genomic) and specialized (locus specific), and a seamless integration of the two types is intended. Each requires a Document (its "biography"). Different mutation databases will have different content and structure, but a minimum core of content in a shared syntax is a necessity; the core includes: (1) a unique identifier of the allele; (2) the source/report of the data; (3) context of the allele; and (4) the allele itself (the description). The allele description should be validated. There is no single correct way to design a mutation database. The uses to which databases are put dictate the design. Software and deployment together recognize the different needs of specialized and generalized databases, while making them mutually compatible through shared content and the appropriate search facilities. A set of eight Recommendations completes these Guidelines for Content, Design, and Deployment of Mutation Databases.

Alleles↗

Whole-genome sequencing of 490,640 UK Biobank participants.

Whole-genome sequencing provides an unbiased and complete view of the human genome and enables the discovery of genetic variation without the technical limitations of other genotyping technologies. Here we report on whole-genome sequencing of 490,640 UK Biobank participants, building on previous genotyping effort1. This advance deepens our understanding of how genetics associates with disease biology and further enhances the value of this open resource for the study of human biology and health. Coupling this dataset with rich phenotypic data, we surveyed within- and cross-ancestry genomic associations and identified novel genetic and clinical insights. Although most associations with disease traits were primarily observed in individuals of European ancestries, strong or novel signals were also identified in individuals of African and Asian ancestries. With the improved ability to accurately genotype structural variants and exonic variation in both coding and UTR sequences, we strengthened and revealed novel insights relative to whole-exome sequencing2,3 analyses. This dataset, representing a large collection of whole-genome sequencing data that is available to the UK Biobank research community, will enable advances of our understanding of the human genome, facilitate the discovery of diagnostics and therapeutics with higher efficacy and improved safety profile, and enable precision medicine strategies with the potential to improve global health.

Humans↗

Predicting the secondary structures and tertiary interactions of 211 group I introns in IE subgroup.

The large number of currently available group I intron sequences in the public databases provides opportunity for studying this large family of structurally complex catalytic RNA by large-scale comparative sequence analysis. In this study, the detailed secondary structures of 211 group I introns in the IE subgroup were manually predicted. The secondary structure-favored alignments showed that IE introns contain 14 conserved stems. The P13 stem formed by long-range base-pairing between P2.1 and P9.1 is conserved among IE introns. Sequence variations in the conserved core divide IE introns into three distinct minor subgroups, namely IE1, IE2 and IE3. Co-variation of the peripheral structural motifs with core sequences supports that the peripheral elements function in assisting the core structure folding. Interestingly, host-specific structural motifs were found in IE2 introns inserted at S516 position. Competitive base-pairing is found to be conserved at the junctions of all long-range paired regions, suggesting a possible mechanism of establishing long-range base-pairing during large RNA folding. These findings extend our knowledge of IE introns, indicating that comparative analysis can be a very good complement for deepening our understanding of RNA structure and function in the genomic era.

Base Pairing↗

Variation and genetic control of surface antigen expression in mycoplasmas: the Vlp system of Mycoplasma hyorhinis.

Surface antigenic diversity in the swine pathogen Mycoplasma hyorhinis is generated by random combinatorial expression and high-frequency phase variation of multiple, size-variant membrane surface lipoproteins (Vlps) which represent the major coat proteins of this wall-less procaryote. The distinctive structural basis for Vlp variation was revealed in a family of several related but divergent vlp genes. These occur in one cluster as single chromosomal copies, each encoding a conserved domain for membrane insertion and lipoprotein processing, and a divergent external domain that changes size by deletion or insertion of repetitive intragenic coding sequences while retaining a distinctive charge motif. Lack of detectable changes in restriction fragment patterns or DNA sequence of vlp structural genes during phase transitions between ON and OFF expression states ruled out long range genomic rearrangements and frameshift mutations as a means of controlling Vlp phase variation. However, highly homologous vlp promoter regions contain a homopolymeric tract of contiguous adenine residues [poly(A)] upstream of the transcriptional start site which is subject to frequent mutations altering its length. These mutations are the only sequence changes detected during phase transitions, and are highly correlated with the expression state of each vlp gene. This suggests a mechanism of transcriptional control regulating Vlp phase variation by critical changes within the poly(A) region affecting the spacing between the -10 and -35 hexamers or a putative regulator binding site. The multiple levels of structural and antigenic diversity embodied in the vlp gene family may provide essential adaptive capabilities for this wall-less microbial pathogen.

Animals↗

Divergence in the chloroplast genome and nuclear rDNA of the rare western australian plant lambertia orbifolia gardner (Proteaceae)

The population genetic structure of the Australian plant Lambertia orbifolia was investigated for chloroplast DNA (cpDNA) and rDNA based on restriction fragment length polymorphism. Variation was assessed in 14-20 individuals from six populations with probes covering the majority of the chloroplast genome and the whole rRNA gene unit. For cpDNA, eight mutations were detected which were distributed over five haplotypes. Nucleotide diversity in the species was high and the majority of this diversity was distributed between populations with diversity within populations restricted to a single population. There was significant differentiation between the two regions in the species distribution with the Narrikup region being distinguished by a single haplotype that was characterized by six unique mutations. Variation in rDNA was detected with three gene length variants present in most individuals. However, the Narrikup region was characterized by homogenization of the gene unit to a single length variant in all individuals. The divergence of the Narrikup region suggests that the disjunction in the species distribution has been present for a long time and the two regions represent separate evolutionary lineages.

Journal Article↗