PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Divergence in the chloroplast genome and nuclear rDNA of the rare western australian plant lambertia orbifolia gardner (Proteaceae)

The population genetic structure of the Australian plant Lambertia orbifolia was investigated for chloroplast DNA (cpDNA) and rDNA based on restriction fragment length polymorphism. Variation was assessed in 14-20 individuals from six populations with probes covering the majority of the chloroplast genome and the whole rRNA gene unit. For cpDNA, eight mutations were detected which were distributed over five haplotypes. Nucleotide diversity in the species was high and the majority of this diversity was distributed between populations with diversity within populations restricted to a single population. There was significant differentiation between the two regions in the species distribution with the Narrikup region being distinguished by a single haplotype that was characterized by six unique mutations. Variation in rDNA was detected with three gene length variants present in most individuals. However, the Narrikup region was characterized by homogenization of the gene unit to a single length variant in all individuals. The divergence of the Narrikup region suggests that the disjunction in the species distribution has been present for a long time and the two regions represent separate evolutionary lineages.

Journal Article↗

Discovery of a new class of immunoglobulin heavy chain from fugu.

In teleosts, the genomic organization of the immunoglobulin (Ig) heavy (H)-chain locus was thought to follow a typical translocon-type multigene structure; however, recent studies have indicated a variation in the structure and this might be teleost specific. Isotypes of the Ig H-chain, namely IgM, IgD, IgZ and IgT, have been identified. In this study, we report the discovery of a new class of IgH from fugu. This isotype was first identified from the genomic sequence of the fugu IgH locus. This novel IgH gene is composed of two constant (C) domains, a hinge region, and two exons encoding membrane regions. Surprisingly, the new IgH gene is present between the variable (V)H and Cmu regions of the locus. The C domains of the new isotype do not show any significant similarity to mammalian or fish IgH genes. The cloned cDNA from the new isotype has typical Ig H-chain characteristics and is expressed as both secretory and membrane form. Transcript analyses suggest that the new IgH from fugu might only use the joining (J)H segments present in front of the new CH domains and that the usage of DH and JH segments is specific to the isotype expressed. The expression pattern of the gene has been confirmed by in situ hybridization and PCR studies.

Amino Acid Sequence↗

Human arylhydrocarbon receptor repressor (AHRR) gene: genomic structure and analysis of polymorphism in endometriosis.

The diversity of biological effects resulting from exposure to dioxin may reflect the ability of this environmental pollutant to alter gene expression by binding to the arylhydrocarbon receptor (AHR) gene and related genes. AHR function may be regulated by structural variations in AHR itself, in the AHR repressor (AHRR), in the AHR nuclear translocator (ARNT), or in AHR target molecules such as cytochrome P-4501A1 (CYP1A1) and glutathione S-transferase. Analysis of the genomic organization of AHRR revealed an open reading frame consisting of a 2094-bp mRNA encoded by ten exons. We found one novel polymorphism, a substitution of Ala by Pro at codon 185 (GCC to CCC), in exon 5 of the AHRR gene; among 108 healthy unrelated Japanese women, genotypes Ala/Ala, Ala/Pro, and Pro/Pro were represented, respectively, by 20 (18.5%), 49 (45.4%), and 39 (36.1%) individuals. We did not detect previously published polymorphisms of ARNT (D511N) or the CYP1A1 promoter (G-469A and C-459T) in our subjects, suggesting that these polymorphisms are rare in the Japanese population. No association was found between uterine endometriosis and any polymorphisms in the AHRR, AHR, ARNT, or CYP1A1 genes analyzed in the present study.

Aryl Hydrocarbon Receptor Nuclear Translocator↗

[Modern variations of human influenza group A viruses at the molecular level].

The authors own results on the variety of the genomic primary structures in human influenza A viruses participating in the epidemic process, including the atypical viruses. The comparative studies revealed new trends in the HA gene antigenic drift on the late stages and the PB1 gene shift. Modifications occurring in the primary structure of the influenza A viruses native genomes during laboratory treatment (adaptation to new hosts, vaccine preparation, egg passaging) have been analyzed. Sequencing of several types of "antigenic anachronisms" revealed the direct links between some of such viruses and the anthropogenic pollution of the biosphere by vaccine strains. Modifications in the HA genes of influenza A viruses during the persistent infection have also been studied.

Amino Acid Sequence↗

Molecular typing of West Nile Virus, Dengue, and St. Louis encephalitis using multiplex sequencing.

We report the development of an assay to simultaneously identify three of the clinically important flaviviruses (West Nile Virus, Dengue, and St. Louis encephalitis). This assay is based on the nucleotide sequence variations within a 266-bp region of the non-structural protein 5. Further, based on the nucleotide variations in the same region of the non-structural protein 5, four of the present Dengue serotypes were identified. To identify some of the subtypes of WNV we have developed a second assay using multiplex sequencing technology. The format of the result of this assay is an electropherogram of two genomic segments of the WNV genome: a 48-nucleotide sequence from the anchored core protein C and a 45-nucleotide sequence coding for the non-structural proteins (proteinase and putative helicase genes).

Base Sequence↗

Molecular epidemiology of African horse sickness virus based on analyses and comparisons of genome segments 7 and 10.

This paper describes a method to rapidly identify African horse sickness virus (AHSV), using a single tube reverse transcription polymerase chain reaction (PCR). This method was used to amplify cDNA copies of genome segments 7 and 10 from several different AHSV strains, of different serotypes, which were then analysed by sequencing and/or endonuclease digestion. AHSV VP7 (encoded by genome segment 7) is one of the two major capsid proteins of the inner capsid layer, forming the outer surface of the core particle. VP7 is highly conserved and is the major serogroup specific antigen common to all nine AHSV serotypes. Digestion of the 1179 bp cDNA with restriction enzymes, allowed differentiation of several strains of different serotypes and identified six distinct groups containing AHSV-1, 3, 6 and 8; AHSV-2; AHSV-4; AHSV-5; AHSV-7; and AHSV-9. Differences were detected between wild type viruses and vaccine strains that had been attenuated by multiple passage in suckling mouse brain or in tissue cultures. RFLP analysis was also used to study variation the 758 bp cDNA copies of AHSV genome segment 10, which encodes the two small non-structural membrane proteins NS3 and NS3a. In this way it was possible to distinguish each of the strains tested, except AHSV 4 (USDA) and AHSV 9 (USDA). However, these isolates could be distinguished by RFLP analysis of genome segment 7 cDNA. Using sequence analysis of genome segment 10 we were able to classify the virus isolates into three groups: AHSV-1, 2 and 8; AHSV-3 and 7; AHSV 4, 5, 6 and 9. These studies confirmed that the virus which first appeared in central Spain in July 1987, subsequently spread into northern Morocco in October 1989.

African Horse Sickness↗

Heterogeneity in codon usage in the flatworm Schistosoma mansoni.

Synonymous codon choices vary considerably among Schistosoma mansoni genes. Principal components analysis detects a single major trend among genes, which highly correlates with GC content in third codon positions and exons, but does not discriminate among putatively highly and lowly expressed genes. The effective number of codons used in each gene, and its distribution when plotted against GC3, suggests that codon usage is shaped mainly by mutational biases. The GC content of exons, GC3, 5', 3', and flanking (5' + 3' + introns) regions are all correlated among them, suggesting that variations in GC content may exist among different regions of the S. mansoni genome. We propose that this genome structure might be among the most important factors shaping codon usage in this species, although the action of selection on certain sequences cannot be excluded.

Animals↗

Population subdivision in Europe's great bustard inferred from mitochondrial and nuclear DNA sequence variation.

A continent-wide survey of sequence variation in mitochondrial (mt) and nuclear (n) DNA of the endangered great bustard (Otis tarda) was conducted to assess the extent of phylogeographic structure in a morphologically monotypic bird. DNA sequence variation in a combined 809 bp segment of the mtDNA genome from 66 individuals from the last six breeding regions showed relatively low levels of intraspecific sequence diversity (n = 0.32%) but significant differences in the regional distribution of 11 haplotypes (phiST = 0.49). Despite their exceptional potential for dispersal, a complete and long-term historical separation between the populations from the Iberian Peninsula (Spain) and mainland Europe (Hungary, Slovakia, Germany, and Russia) was demonstrated. Divergence between populations based on a 3-bp insertion-deletion polymorphism within the intron region of the nuclear CHD-Z gene was geographically concordant with the primary subdivision identified within the mtDNA sequences. Inferred aspects of phylogeography were used to formulate conservation recommendations for this endangered species.

Animals↗

The lipo-oligosaccharides of Haemophilus influenzae: an interesting array of characters.

The composition of the lipo-oligosaccharide (LOS) of Haemophilus influenzae is highly variable, especially in the oligosaccharide region. Many of the biosynthetic and transferase genes involved in LOS biosynthesis vary in seemingly random fashion by means of polymerase stuttering within redundant sequences in the 5'-portion of the genes. This results in a heterogeneous population of individual bacteria expressing literally thousands of LOS glycoforms. The simultaneous variation in the expression and structural context of a large number of individual carbohydrate and lipid structures within the LOS yields a diverse array of LOS glycoforms. The expression of glycoforms that mimic host structures may allow the organism to evade innate defenses and to manipulate host cell biology. We review how this randomly generated bacterial combinatorial chemistry results in the production of a large number of carbohydrate structures, in essentially any conceivable structural context, some of which allow the organism to utilize host cell receptors. By generating a diverse population of bacteria expressing different LOS glycoforms, discrete H. influenzae subpopulations may be adapted for survival of different environmental stresses within the airways. Thus, H. influenzae utilizes a simple and efficient "Monte Carlo" strategy for achieving maximal variation in cell surface structures, which allow the organism to adapt efficiently to environmental stresses with a small genome.

Antigens, Bacterial↗

A computer simulation analysis of the accuracy of partial genome sequencing and restriction fragment analysis in the reconstruction of phylogenetic relationships.

Partial genome sequencing (PGS) and restriction fragment analysis (RFA) are used frequently in molecular epidemiologic investigations. The relative accuracy of PGS and RFA in phylogenetic reconstruction has not been assessed. In this study, 32 model phylogenetic trees with 16 extant lineages were generated, for which DNA sequences were simulated under varying conditions of genome length, nucleotide substitution rate, and between-site substitution rate variation. Genotyping using PGS and RFA was simulated. The effect of tree structure (stemminess, imbalance, lineage variation) on the accuracy of phylogenetic reconstruction (topological and branch length similarity) was evaluated. Overall, PGS was more accurate than RFA. The accuracy of PGS increased with increasing sequence length. The accuracy of RFA increased with the number of restriction enzymes used. In fragment size comparison, the Dice and Nei-Li algorithms differed little, with both more accurate than the Fragment Size Distribution algorithm. For RFA, higher tree stemminess and longer genome length were associated with higher topological accuracy, whereas lower tree stemminess and lower substitution rates were associated with higher branch length accuracy. For PGS, lower tree imbalance was associated with higher topological accuracy, whereas lower tree stemminess, higher substitution rate, and lower between-site substitution rate variation were associated with higher branch length accuracy. RFA had higher topological accuracy than PGS only for the shortest sequence length (200 bps) at a low substitution rate, high tree stemminess, and long genome length. PGS had equal or higher accuracy in branch length reconstruction than RFA under all conditions investigated. Thus, partial genome sequencing is recommended over restriction fragment analysis for conditions within the parameter space examined.

Computational Biology↗

Diversification of Sulawesi macaque monkeys: decoupled evolution of mitochondrial and autosomal DNA.

In macaque monkeys, females are philopatric and males are obligate dispersers. This social system is expected to differently affect evolution of genetic elements depending on their mode of inheritance. Because of this, the geographic structure of molecular variation may differ considerably in mitochondrial DNA (mtDNA) and in autosomal DNA (aDNA) in the same individuals, even though these genomes are partially co-inherited. On the Indonesian island of Sulawesi, macaque monkeys underwent an explosive diversification as a result of range fragmentation. Today, barriers to dispersal have receded and fertile hybrid individuals can be found at contact zones between parapatric species. In this study, we examine the impact of range fragmentation on Sulawesi macaque mtDNA and aDNA by comparing evolution, phylogeography, and population subdivision of each genome. Our results suggest that mtDNA is paraphyletic in some species, and that mtDNA phylogeography is largely consistent with a pattern of isolation by distance. Autosomal DNA, however, is suggestive of fragmentation, in that interspecific differentiation across most contact zones is significant but intraspecific differentiation between contact zones is not. Furthermore, in mtDNA, most molecular variation is partitioned between populations within species but in aDNA most variation is partitioned within populations. That mtDNA has a different geographic structure than aDNA (and morphology) in these primates is a probable consequence of (1) a high level of ancestral polymorphism in mtDNA, (2) differences between patterns of ancestral dispersal of matrilines and contemporary dispersal of males, and (3) the fact that female philopatry impedes gene flow of macaque mtDNA.

Alleles↗

Significant variation in haplotype block structure but conservation in tagSNP patterns among global populations.

The initial belief that haplotype block boundaries and haplotypes were largely shared across populations was a foundation for constructing a haplotype map of the human genome using common SNP markers. The HapMap data document the generality of a block-like pattern of linkage disequilibrium (LD) with regions of low and high haplotype diversity but differences among the populations. Studies of many additional populations demonstrate that LD patterns can be highly variable among populations both across and within geographic regions. Because of this variation, emphasis has shifted to the generalizability of tagSNPs, those SNPs that capture the bulk of variation in a region. We have examined the LD and tagSNP patterns based upon over 2000 individual samples in 38 populations and 134 SNPs in 10 genetically independent loci for a total of 517 kb with an average density of 1 SNP/5 kb. Four different 'block' definitions and the pairwise LD tagSNP selection algorithm have been applied. Our results not only confirm large variation in block partition among populations from different regions (agreeing with previous studies including the HapMap) but also show that significant variation can occur among populations within geographic regions. None of the block-defining algorithms produces a consistent pattern within or across all geographic groups. In contrast, tagSNP transferability is much greater than the similarity of LD patterns and, although not perfect, some generalizations of transferability are possible. The analyses show an asymmetric pattern of tagSNP transferability coinciding with the subsetting of variation attributed to the spread of modern humans around the world.

Genetic Variation↗

Structure and sequence variation at the human leptin receptor gene in lean and obese Pima Indians.

The cloning of human and mouse cDNAs from brain that encode high affinity leptin receptors was recently reported. We have physically localized the human leptin receptor gene (LEPR) to a region at 1p31, between the anonymous microsatellite markers D1S515 and D1S198. The genomic structure of the human leptin receptor gene, corresponding to the published human brain cDNA sequence, spans over 70 kb and includes 20 exons. Since the leptin receptor gene is a candidate gene for obesity, and because of its proximity to D1S198, a marker previously linked to insulin secretion, the LEPR gene was sequenced in 20 non-diabetic Pima Indians chosen for extremes in percent body fat and in their acute insulin response to intravenous glucose. Seven polymorphic sites were identified. Two of these polymorphisms, Lys109Arg and Gln223Arg, are amino acid substitutions in the extracellular domain of the leptin receptor, one polymorphism is a silent substitution, and four occur in non-coding regions of the leptin receptor. Four of these sites are in linkage disequilibrium with one another. Nucleotides at three noncoding polymorphic sites were found exclusively in obese Pima Indians. This demonstrates an association between variation at the leptin receptor gene and obesity in humans.

Adipose Tissue↗

Haplotype structure and evidence for positive selection at the human IL13 locus.

Interleukin-13 (IL13) is believed to play an important role in the pathogenesis of atopy and allergic asthma. To better understand genetic variation at the IL13 locus, we resequenced a 5.1-kb genomic region spanning the entire locus and identified 26 single-nucleotide polymorphisms (SNPs) in 74 individuals from three major populations-Chinese, Caucasian, and African. Our survey suggests exceptionally high and significant geographic structure at the IL13 locus between African and outside Africa populations. This unusual pattern suggests that positive selection that acts in some local populations may have played a role on the IL13 locus. In support of this suggestion, we found a significant excess of high frequency-derived SNPs in the Chinese population and Caucasian population, respectively, as expected after a recent episode of positive selection. Further, the unusual haplotype structure indicates that different scenarios of the action of positive selection on the IL13 locus in different populations may exist. In the Caucasian population, the skewed haplotype distribution dominated by one common haplotype supports the hypothesis of simple directional selection. Whereas, in the Chinese population, the two-round hitchhiking hypothesis may explain the skewed haplotype structure with three dominant ones. These findings may provide insight into the likely relative roles of selection and population history in establishing present-day variation at the IL13 locus, and, motivate further studies of this locus as an important candidate in common diseases association studies.

Asian People↗

Constraint structure analysis of gene expression.

A microarray experiment gives a snapshot of the state of an organism in terms of the relative abundances of its mRNA transcripts, locating the organism at a point in a high dimensional state space where each axis represents the relative expression level of a single gene. Multiple experiments generate a cloud of points in this gene expression space. We present a geometric approach to analyzing the covariational properties of such a cloud and use a dataset from Saccharomyces cerevisiae as an illustration. In particular, we use singular value decomposition to identify significant linear sub-structures in the data and analyze the contributions of both individual genes and functional classes of genes to these major directions of variation. Analyzing the publicly available yeast expression data, we show that under all experimental conditions the variation in expression is limited to a small number of linear dimensions. Projections of individual gene axes onto the significant dimensions can order the contribution of individual genes to variation in expression within an experiment. We show that no particular groups of genes characterize particular experimental conditions. Instead, the particular structure of the coordinated expression of the entire genome characterizes a particular experiment.

Gene Expression↗

Complex structural variation, phylogeny, and disease associations of the mucin pangenome.

Mucins are large glycoproteins that provide hydration and barrier function to epithelial tissues. Although genetically heterogeneous, all mucins harbor a large exon composed of variable number tandem repeats (VNTRs). Short-read sequencing has limited our understanding of mucin VNTR diversity and makes disease association studies challenging. We leverage 296 long-read phased genome assemblies to characterize 14 mucin family members, achieving &#x2265;97% accuracy across 572 haplotypes. Phylogenetic haplogroup analysis reveals extraordinary structural heterozygosity, with MUC4 harboring the greatest allelic diversity (n=240 distinct lengths) and MUC12 the greatest size range (&#x394; = 55,233 bp; 23,080 amino acids). Ten mucins show significant population stratification (pFDR < 0.05). At the MUC4/MUC20 locus, we characterize higher-order structural variation, including a recurrent inversion, copy number variation, and interlocus gene conversion. Optimized genotyping achieves &#x2265;95% haplogroup concordance across 10 loci. We apply this to 4,637 deeply phenotyped cystic fibrosis patients and identify a significant association between short MUC1 VNTRs and severe disease (p=0.0056), demonstrating the pangenome's utility for complex locus genotyping and disease discovery.

Journal Article↗

Molecular cloning and sequence analysis of duck hepatitis B virus genomes of a new variant isolated from Shanghai ducks.

The genomes of duck hepatitis B virus (DHBV) from a brown duck (S5) and a white duck (S31) kept independently in Shanghai, China, were cloned and the complete nucleotide sequence of each virion DNA (DHBV-S5 and DHBV-S31) was determined. DHBV-S5 and DHBV-S31 were both 3027 bp in length and 6 bp longer than the other two DHBVs analyzed previously, DHBV16 and DHBV3. The genomes of DHBV-S5 and DHBV-S31 encoded three long overlapping open reading frames designated as P, S, and C. A possible new open reading frame was found in a complementary strand of each viral genome, as 336 bp for DHBV-S5 and 306 bp for DHBV-S31, respectively. A pair of 3-bp insertions were found in the overlapping region of pre-S2 and P and so two amino acids were inserted in this region in DHBV-S5 and DHBV-S31. The nucleotide sequence variation between DHBV-S5 and DHBV-S31 (4.9%) was similar to that between DHBV16 and DHBV3 (5.6%), and less than the variations between either of these Shanghai clones and DHBV16 or DHBV3 (9.5-10.4%). The amino acid sequence was also conserved in the two Shanghai clones but showed group difference from DHBV16 or DHBV3. Thus these two independent Shanghai clones of DHBV showed geographical characteristics of genomic structure.

Amino Acid Sequence↗

A worldwide survey of haplotype variation and linkage disequilibrium in the human genome.

Recent genomic surveys have produced high-resolution haplotype information, but only in a small number of human populations. We report haplotype structure across 12 Mb of DNA sequence in 927 individuals representing 52 populations. The geographic distribution of haplotypes reflects human history, with a loss of haplotype diversity as distance increases from Africa. Although the extent of linkage disequilibrium (LD) varies markedly across populations, considerable sharing of haplotype structure exists, and inferred recombination hotspot locations generally match across groups. The four samples in the International HapMap Project contain the majority of common haplotypes found in most populations: averaging across populations, 83% of common 20-kb haplotypes in a population are also common in the most similar HapMap sample. Consequently, although the portability of tag SNPs based on the HapMap is reduced in low-LD Africans, the HapMap will be helpful for the design of genome-wide association mapping studies in nearly all human populations.

Chromosome Mapping↗