PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

On the trail of a cereal killer: Exploring the biology of Magnaporthe grisea.

The blast fungus Magnaporthe grisea causes a serious disease on a wide variety of grasses including rice, wheat, and barley. Rice blast is the most serious disease of cultivated rice and therefore poses a threat to the world's most important food security crop. Here, I review recent progress toward understanding the molecular biology of plant infection by M. grisea, which involves development of a specialized cell, the appressorium. This dome-shaped cell generates enormous turgor pressure and physical force, allowing the fungus to breach the host cuticle and invade plant tissue. The review also considers the role of avirulence genes in M. grisea and the mechanisms by which resistant rice cultivars are able to perceive the fungus and defend themselves. Finally, the likely mechanisms that promote genetic diversity in M. grisea and our current understanding of the population structure of the blast fungus are evaluated.

Genetic Variation↗

Extensive sequence divergence and phylogenetic relationships between the fusogenic and nonfusogenic orthoreoviruses: a species proposal.

The orthoreoviruses can be divided into subgroups based on either their restricted host range or the unusual ability of certain members of this group of nonenveloped viruses to induce cell-cell fusion from within. Phylogenetic relationships cannot be inferred based on these biological properties because fusogenic reoviruses are present in both the avian and mammalian subgroups. To address this issue, the complete nucleotide sequences of the three S-class genome segments encoding the major sigma-class core, outer capsid, and nonstructural proteins of four fusogenic reoviruses were determined and used to establish the phylogeny of the orthoreoviruses. The viruses analysed included two strains of avian reovirus and the only known fusogenic mammalian reoviruses, Nelson Bay virus and baboon reovirus. Comparative sequence analysis of these fusogenic reoviruses and the prototypical nonfusogenic mammalian reoviruses indicated a highly diverged genus with both conserved and unique sequence-predicted structural motifs in the major sigma-class proteins. Phylogenetic analysis provided the basis for the first taxonomic subdivision of the orthoreoviruses into species classes based on inferred evolutionary relationships. It is proposed that the orthoreoviruses consist of at least four species that separate into three clades. The nonfusogenic mammalian reovirus species represent a single clade, and the fusogenic reoviruses separate into two distinct clades. The first clade of fusogenic reoviruses contains the avian reovirus- and Nelson Bay virus-type species, with the second clade being occupied by the single baboon reovirus isolate that represents a fourth orthoreovirus species.

Amino Acid Sequence↗

Extent and distribution of linkage disequilibrium in three genomic regions.

The positional cloning of genes underlying common complex diseases relies on the identification of linkage disequilibrium (LD) between genetic markers and disease. We have examined 127 polymorphisms in three genomic regions in a sample of 575 chromosomes from unrelated individuals of British ancestry. To establish phase, 800 individuals were genotyped in 160 families. The fine structure of LD was found to be highly irregular. Forty-five percent of the variation in disequilibrium measures could be explained by physical distance. Additional factors, such as allele frequency, type of polymorphism, and genomic location, explained <5% of the variation. Nevertheless, disequilibrium was occasionally detectable at 500 kb and was present for over one-half of marker pairs separated by <50 kb. Although these findings are encouraging for the prospects of a genomewide LD map, they suggest caution in interpreting localization due to allelic association.

Computer Simulation↗

A second-site suppressor strategy for chemical genetic analysis of diverse protein kinases.

Chemical genetic analysis of protein kinases involves engineering kinases to be uniquely sensitive to inhibitors and ATP analogs that are not recognized by wild-type kinases. Despite the successful application of this approach to over two dozen kinases, several kinases do not tolerate the necessary modification to the ATP binding pocket, as they lose catalytic activity or cellular function upon mutation of the 'gatekeeper' residue that governs inhibitor and nucleotide substrate specificity. Here we describe the identification of second-site suppressor mutations to rescue the activity of 'intolerant' kinases. A bacterial genetic selection for second-site suppressors using an aminoglycoside kinase APH(3')-IIIa revealed several suppressor hotspots in the kinase domain. Informed by results from this selection, we focused on the beta sheet in the N-terminal subdomain and generated a structure-based sequence alignment of protein kinases in this region. From this alignment, we identified second-site suppressors for several divergent kinases including Cdc5, MEKK1, GRK2 and Pto. The ability to identify second-site suppressors to rescue the activity of intolerant kinases should facilitate chemical genetic analysis of the majority of protein kinases in the genome.

Amino Acid Substitution↗

Variability and conservation in hepatitis B virus core protein.

BACKGROUND: Hepatitis B core protein (HBVc) has been extensively studied from both a structural and immunological point of view, but the evolutionary forces driving sequence variation within core are incompletely understood. RESULTS: In this study, the observed variation in HBVc protein sequence has been examined in a collection of a large number of HBVc protein sequences from public sequence repositories. An alignment of several hundred sequences was carried out, and used to analyse the distribution of polymorphisms along the HBVc. Polymorphisms were found at 44 out of 185 amino acid positions analysed and were clustered predominantly in those parts of HBVc forming the outer surface and spike on intact capsid. The relationship between HBVc diversity and HBV genotype was examined. The position of variable amino acids along the sequence was examined in terms of the structural constraints of capsid and envelope assembly, and also in terms of immunological recognition by T and B cells. CONCLUSION: Over three quarters of amino acids within the HBVc sequence are non-polymorphic, and variation is focused to a few amino acids. Phylogenetic analysis suggests that core protein specific forces constrain its diversity within the context of overall HBV genome evolution. As a consequence, core protein is not a reliable predictor of virus genotype. The structural requirements of capsid assembly are likely to play a major role in limiting diversity. The phylogenetic analysis further suggests that immunological selection does not play a major role in driving HBVc diversity.

Amino Acid Sequence↗

Influence of genome imprinting on gene expression, phenotypic variations and development.

Genome imprinting confers functional differences on parental chromosomes as a result of the differences in epigenetic inheritance from parental germlines. Repressed and derepressed chromatin structures probably constitute the initial germline-dependent 'imprints'. Any subsequent modifications, such as DNA methylation, will be influenced by these initial epigenetic modifications. Hence, epigenetic modifications of parental alleles probably occur progressively and this will affect their potential for expression. It appears that imprinting of some parental alleles is critical for their dosage, affecting embryonic growth, cell proliferation and differentiation. Genetic studies highlight the influence of subsets of imprinted genes and identify those which are crucial for development. Genomic imprinting also affects some transgene loci and dominant mutations with accompanying variable penetrance and expressivity. The response of transgenes can be influenced by modifier genes whose presence is most readily detected in different inbred backgrounds. The influence of modifier genes can in turn be affected by their parental origin, perhaps partly by the maternally inherited oocyte cytoplasmic factors, as well as by complex interactions between some parental alleles and oocyte cytoplasmic factors. The resulting epigenetic modifications of unlinked loci can result in substantial phenotypic variations.

Animals↗

Structure of rDNA in the mosquito Anopheles gambiae and rDNA sequence variation within and between species of the A. gambiae complex.

The structure of the rDNA repeating unit of Anopheles gambiae (Diptera: Culicidae) was determined by restriction endonuclease mapping and hybridization analyses on four independent clones obtained from a genomic library of a colony (G3) from the Gambia (West Africa). rDNA gene coding sequences are conserved, but much intragenomic and intraspecific (geographic) variation occurs in the intergenic spacer. Hybridization of subclones from spacer and coding sequences to genomic DNA that was isolated from single mosquitoes from laboratory colonies of four other A. gambiae complex species reveals conservation of coding sequences but concerted evolution in the intergenic spacers.

Africa, Western↗

In silico generation of synthetic cancer genomes using generative AI.

Understanding how genomic alterations drive cancer is key to advancing precision oncology. To detect these alterations, accurate algorithms are used; however, due to privacy concerns, few deeply sequenced cancer genomes can be shared, limiting benchmarking and representing a major obstacle to the improvement of analytic tools. To address this, we developed OncoGAN, a generative AI model combining adversarial networks and variational autoencoders to create realistic synthetic cancer genomes. Trained on large-scale genomic datasets, OncoGAN accurately reproduces somatic mutations, copy number alterations, and structural variants across cancer types while preserving donors' privacy. The synthetic genomes reflect tumor-specific mutational signatures and positional mutation patterns. Using DeepTumour, we validated the synthetic data's fidelity, showing high concordance between generated and predicted tumors. Moreover, augmenting the training data with synthetic genomes improved DeepTumour's accuracy, underscoring OncoGAN's potential to generate shareable datasets with known ground truths for benchmarking and enhancement of cancer genome analysis tools.

Humans↗

A De Novo 16p13.3 Triplication Underlying Early-Onset Complex Neurodegeneration.

BACKGROUND: Neurodegenerative disorders are clinically and genetically heterogeneous, characterized by progressive neuronal loss and multidomain functional decline. Despite a presumed genetic etiology, a substantial proportion of cases remain molecularly undiagnosed. OBJECTIVE: The aim was to identify the genetic cause of an early-onset neurodegenerative disorder presenting with ataxia and cognitive impairment. METHODS: Rare copy-number variants were detected via short-read whole-genome sequencing (WGS), with candidate structural models inferred using long-read WGS. We performed transcriptomic profiling of peripheral blood leukocytes by RNA sequencing, with validation using reverse transcription-quantitative polymerase chain reaction (RT-qPCR). RESULTS: We identified a de novo copy-number gain at 16p13.3. Combined copy-number profiling and long-read WGS suggested a candidate model comprising a triplicated segment in tandem with a proximal duplication, joined to a distal duplication via an inverted junction. Transcriptomic analysis demonstrated significant upregulation of ATP6V0C, AMDHD2, and PDPK1. CONCLUSIONS: These findings support a role for structural variation in early-onset neurodegeneration and highlight the value of combining short-read copy-number profiling with long-read WGS to detect and characterize complex genomic rearrangements. &#xa9; 2026 International Parkinson and Movement Disorder Society.

16p13.3↗

Signals of Natural Selection Across Regions of Low Recombination in Wild Populations of the Purple Sea Urchin, Strongylocentrotus purpuratus.

Structural variants (SVs) are increasingly recognized as important components of genetic architecture. Yet our understanding of the evolutionary forces maintaining SVs in natural populations is limited. Chromosomal inversions in particular can facilitate local adaptation in populations with high gene flow, including many marine species. The purple sea urchin (Strongylocentrotus purpuratus) is a powerful system to study these dynamics due to its high gene flow, lack of population structure, and broad latitudinal range. We analyzed whole genome sequence data from 137 individuals sampled across seven populations to identify regions of low recombination using scans for elevated linkage disequilibrium and genetic differentiation. Such regions may arise from structural variants, including chromosomal inversions. We identified nine regions showing signatures of reduced recombination, including three way genotype clustering, long range linkage, and hanging bridge patterns frequently associated with inversion polymorphisms. The regions were polymorphic within locations and along the species range with three loci showing concordant signatures of balancing and spatially heterogeneous selection based on enrichment of outliers and distinct patterns of allelic age. Additionally, these loci showed enrichment for genes associated with biomineralization and development. Our results provide the first evidence for regions of low recombination in the purple sea urchin genome, several of which display genomic signatures consistent with structural variants such as chromosomal inversions. These findings add to growing evidence that regions of reduced recombination constitute an important component of standing genetic variation in natural populations and may play a key role in adaptation to heterogeneous environments.

Strongylocentrotus purpuratus↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗

The Fire Ant Social Chromosome Exerts a Major Influence on Genome Regulation.

Supergenes underlying complex trait polymorphisms ensure that sets of coadapted alleles remain genetically linked. Despite their prevalence in nature, the mechanisms of supergene effects on genome regulation are poorly understood. In the fire ant Solenopsis invicta, a supergene containing over 500 individual genes influences trait variation in multiple castes to collectively underpin a colony level social polymorphism. Here, we present results of an integrative investigation of supergene effects on gene regulation. We present analyses of ATAC-seq data to investigate variation in chromatin accessibility by supergene genotype and STARR-seq data to characterize enhancer activity by supergene haplotype. Integration with gene co-expression analyses, newly mapped intact transposable elements (TEs), and previously identified copy number variants (CNVs) collectively reveals widespread effects of the supergene on chromatin structure, gene transcription, and regulatory element activity, with a genome-wide bias for open chromatin and increased expression in the presence of the derived supergene haplotype, particularly in regions that harbor intact TEs. Integrated consideration of CNVs and regulatory element divergence suggests each evolved in concert to shape the expression of supergene encoded factors, including several transcription factors that may directly contribute to the trans-regulatory footprint of a heteromorphic social chromosome. Overall, we show how genome structure in the form of a supergene has wide-reaching effects on gene regulation and gene expression.

Animals↗

Crystal structure of the bacterial YhcH protein indicates a role in sialic acid catabolism.

The yhcH gene is part of the nan operon in bacteria that encodes proteins involved in sialic acid catabolism. Determination of the crystal structure of YhcH from Haemophilus influenzae was undertaken as part of a structural genomics effort in order to assist with the functional assignment of the protein. The structure was determined at 2.2-A resolution by multiple-wavelength anomalous diffraction. The protein fold is a variation of the double-stranded beta-helix. Two antiparallel beta-sheets form a funnel opened at one side, where a putative active site contains a copper ion coordinated to the side chains of two histidine and two carboxylic acid residues. A comparison to other proteins with a similar fold and analysis of the genomic context suggested that YhcH may be a sugar isomerase involved in processing of exogenous sialic acid.

Amino Acid Sequence↗

Structure of the human gene for monoamine oxidase type A.

Monoamine oxidases, type A and type B, are principal enzymes for the degradation of biogenic amines, including catecholamines and serotonin. These isozymes have been implicated in neuropsychiatric disorders. Previously, cDNA clones for both MAO-A and MAO-B have been sequenced and the genes encoding them have been localized to human chromosome Xp11.23-Xp11.4. In this work, we isolated human genomic clones spanning almost all the MAOA gene from cosmid and phage libraries using a cDNA probe for MAO-A. Restriction mapping and sequencing show that the human MAOA gene extends over 70 kb and is composed of 15 exons. The exon structure of human MAOA is similar to that described by others for human MAOB. Exon 12 (bearing the codon for cysteine, which carries the covalently bound FAD cofactor) and exon 13 are highly conserved between human MAOA and MAOB genes (92% at the amino acid level). Earlier work revealed two species of MAO-A mRNA, 2.1 kb and 4.5-5.5 kb. We now report on further cDNA isolation and sequencing, which demonstrates that the longer message has an extension of 2.2 kb in the 3' noncoding region. This extended region is contained entirely within exon 15. The two messages therefore appear to be generated by the use of two alternative polyadenylation sites. Results from the present work should facilitate the mutational analysis of functional domains of MAO-A and MAO-B. Knowledge of the gene structure will also help in evaluating the role of genetic variations in MAO-A in human disease through the use of genomic DNA, which is more accessible than the RNA, as a template for PCR-amplification and sequencing.

Amino Acid Sequence↗

Search for characteristic structural features of mammalian mitochondrial tRNAs.

A number of mitochondrial (mt) tRNAs have strong structural deviations from the classical tRNA cloverleaf secondary structure and from the conventional L-shaped tertiary structure. As a consequence, there is a general trend to consider all mitochondrial tRNAs as "bizarre" tRNAs. Here, a large sequence comparison of the 22 tRNA genes within 31 fully sequenced mammalian mt genomes has been performed to define the structural characteristics of this specific group of tRNAs. Vertical alignments define the degree of conservation/variability of primary sequences and secondary structures and search for potential tertiary interactions within each of the 22 families. Further horizontal alignments ascertain that, with the exception of serine-specific tRNAs, mammalian mt tRNAs do fold into cloverleaf structures with mostly classical features. However, deviations exist and concern large variations in size of the D- and T-loops. The predominant absence of the conserved nucleotides G18G19 and T54T55C56, respectively in these loops, suggests that classical tertiary interactions between both domains do not take place. Classification of the tRNA sequences according to their genomic origin (G-rich or G-poor DNA strand) highlight specific features such as richness/poorness in mismatches or G-T pairs in stems and extremely low G-content or C-content in the D- and T-loops. The resulting 22 "typical" mammalian mitochondrial sequences built up a phylogenetic basis for experimental structural and functional investigations. Moreover, they are expected to help in the evaluation of the possible impacts of those point mutations detected in human mitochondrial tRNA genes and correlated with pathologies.

Acylation↗

Impact of plasmids and genetic change on the numerical classification of staphylococci.

Newly isolated bacterial strains often contain extrachromosomal DNA as plasmid DNA. These accessory components of the DNA gene pool confer additional phenotypic properties on their host but, despite this, little attention has been paid to the impact of plasmid-mediated characters on bacterial classification. In the present study, the effect of antibiotic resistance plasmids on the classification of representative staphylococci was determined using numerical phenetic techniques. Over sixty percent of the eighty-one test strains contained one or more plasmids which varied in molecular weight from 1.4 to 36 Mdal. Antibiotic resistance phenotypes were eliminated from strains of S. aureus, S. chromogenes, S. cohnii, S. hyicus and S. xylosus, and from a laboratory isolate, to give sixteen derivative strains. Fourteen had lost one or more plasmids and two had deleted plasmids. In addition three further derivative strains were isolated which showed no plasmid loss but exhibited gross phenotypic changes. The test and derivative strains were the subject of numerical phenetic analyses based on seventy-eight unit characters. Data were examined using the simple matching, Jaccard and pattern coefficients and clustering achieved using the unweighted pair group method with arithmetic averages algorithm. Cluster composition was not markedly affected by the statistics used or by test error, estimated at 1.02%. Numerically circumscribed clusters and subclusters were equated with the established species S. aureus, S. chromogenes, S. cohnii, S. hyicus, S. lentus, S. intermedius, S. sciuri and S. xylosus. The sixteen derivative strains with either lost or delected plasmids were recovered in the same cluster or subcluster as their corresponding parent indicating that the removal of plasmid-expressed characters had little effect on the structure of the numerical classification. In contrast, two of the three strains of S. xylosus with genomically-derived phenotypic variation formed a cluster that separated from their parent strain at the 70% similarity level in the SSM, UPGMA analysis.

Animals↗

Repetitive DNA and chromosome evolution in plants.

Most higher plant genomes contain a high proportion of repeated sequences. Thus repetitive DNA is a major contributor to plant chromosome structure. The variation in total DNA content between species is due mostly to variation in repeated DNA content. Some repeats of the same family are arranged in tandem arrays, at the sites of heterochromatin. Examples from the Secale genus are described. Arrays of the same sequence are often present at many chromosomal sites. Heterochromatin often contains arrays of several unrelated sequences. The evolution of such arrays in populations is discussed. Other repeats are dispersed at many locations in the chromosomes. Many are likely to be or have evolved from transposable elements. The structures of some plant transposable elements, in particular the sequences of the terminal inverted repeats, are described. Some elements in soybean, antirrhinum and maize have the same inverted terminal repeat sequences. Other elements of maize and wheat share terminal homology with elements from yeast, Drosophila, man and mouse. The evolution of transposable elements in plant populations is discussed. The amplification, deletion and transposition of different repeated DNA sequences and the spread of the mutations in populations produces a turnover of repetitive DNA during evolution. This turnover process and the molecular mechanisms involved are discussed and shown to be responsible for divergence of chromosome structure between species. Turnover of repeated genes also occurs. The molecular processes affecting repeats imply that the older a repetitive DNA family the more likely it is to exist in different forms and in many locations within a species. Examples to support this hypothesis are provided from the Secale genus.

Animals↗

Major structural differences and novel potential virulence mechanisms from the genomes of multiple campylobacter species.

Sequencing and comparative genome analysis of four strains of Campylobacter including C. lari RM2100, C. upsaliensis RM3195, and C. coli RM2228 has revealed major structural differences that are associated with the insertion of phage- and plasmid-like genomic islands, as well as major variations in the lipooligosaccharide complex. Poly G tracts are longer, are greater in number, and show greater variability in C. upsaliensis than in the other species. Many genes involved in host colonization, including racR/S, cadF, cdt, ciaB, and flagellin genes, are conserved across the species, but variations that appear to be species specific are evident for a lipooligosaccharide locus, a capsular (extracellular) polysaccharide locus, and a novel Campylobacter putative licABCD virulence locus. The strains also vary in their metabolic profiles, as well as their resistance profiles to a range of antibiotics. It is evident that the newly identified hypothetical and conserved hypothetical proteins, as well as uncharacterized two-component regulatory systems and membrane proteins, may hold additional significant information on the major differences in virulence among the species, as well as the specificity of the strains for particular hosts.

Animals↗