PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

S100 proteins in mouse and man: from evolution to function and pathology (including an update of the nomenclature).

The S100 protein family is the largest subgroup within the superfamily of proteins carrying the Ca2+-binding EF-hand motif. Despite their small molecular size and their conserved functional domain of two distinct EF-hands, S100 proteins developed a plethora of tissue-specific intra- and extracellular functions. Accordingly, various diseases such as cardiomyopathies, neurodegenerative and inflammatory disorders, and cancer are associated with altered S100 protein levels. Here, we review the different S100 protein functions and related diseases from an evolutionary point of view. We analyzed the structural variations, which are the basis of functional diversification, as well as the genomic organization of the S100 family in human and compared it with the S100 repertoires in mouse and rat. S100 genes and proteins are highly conserved between the different mammalian species. Moreover, we identified evolutionary related subgroups of S100 proteins within the three species, which share functional similarity and form subclusters on the genomic level. The available S100-specific mouse models are summarized and the consequences of our results are discussed with regard to the use of genetically engineered mice as human disease models. An update of the S100 nomenclature is included, because some of the recently identified S100 genes and pseudogenes had to be renamed.

Amino Acid Sequence↗

Prediction of functional sites in proteins using conserved functional group analysis.

A detailed knowledge of a protein's functional site is an absolute prerequisite for understanding its mode of action at the molecular level. However, the rapid pace at which sequence and structural information is being accumulated for proteins greatly exceeds our ability to determine their biochemical roles experimentally. As a result, computational methods are required which allow for the efficient processing of the evolutionary information contained in this wealth of data, in particular that related to the nature and location of functionally important sites and residues. The method presented here, referred to as conserved functional group (CFG) analysis, relies on a simplified representation of the chemical groups found in amino acid side-chains to identify functional sites from a single protein structure and a number of its sequence homologues. We show that CFG analysis can fully or partially predict the location of functional sites in approximately 96% of the 470 cases tested and that, unlike other methods available, it is able to tolerate wide variations in sequence identity. In addition, we discuss its potential in a structural genomics context, where automation, scalability and efficiency are critical, and an increasing number of protein structures are determined with no prior knowledge of function. This is exemplified by our analysis of the hypothetical protein Ydde_Ecoli, whose structure was recently solved by members of the North East Structural Genomics consortium. Although the proposed active site for this protein needs to be validated experimentally, this example illustrates the scope of CFG analysis as a general tool for the identification of residues likely to play an important role in a protein's biochemical function. Thus, our method offers a convenient solution to rapidly and automatically process the vast amounts of data that are beginning to emerge from structural genomics projects.

Algorithms↗

Defining and cataloging variants in pangenome graphs.

Structural variation causes some human haplotypes to align poorly with the linear reference genome, leading to 'reference bias'. A pangenome reference graph could ameliorate this bias by relating a sample to multiple reference assemblies. However, this approach requires a new definition of a 'genetic variant.' We introduce a definition of pangenome variants and a method, pantree, to identify them. Our approach involves a pangenome reference tree which includes all nodes (sequences) of the pangenome graph, but only a subset of its edges; non-reference edges are variant edges. Our variants are biallelic and have well-defined positions. Analyzing the Minigraph-Cactus draft human pangenome reference graph, we identified 29.6 million genetic variants. Most variants (99.2%) are small, and most small variants (73.9%) are SNPs. 3.5 million variants (11.7%) have a reference allele which is not on GRCh38; these variants are difficult to detect without a pangenome reference, or with existing pangenome-based approaches. They tend to be embedded within tangled, multiallelic regions. We analyze two medically relevant regions, around the HLA-A and RHD genes, identifying thousands of small variants embedded within several large insertions, deletions, and inversions. We release an open-source software tool together with a VCF variant catalogue.

Journal Article↗

Distribution, expression, and motif variability of ankyrin domain genes in Wolbachia pipientis.

The endosymbiotic bacterium Wolbachia pipientis infects a wide range of arthropods, in which it induces a variety of reproductive phenotypes, including cytoplasmic incompatibility (CI), parthenogenesis, male killing, and reversal of genetic sex determination. The recent sequencing and annotation of the first Wolbachia genome revealed an unusually high number of genes encoding ankyrin domain (ANK) repeats. These ANK genes are likely to be important in mediating the Wolbachia-host interaction. In this work we determined the distribution and expression of the different ANK genes found in the sequenced Wolbachia wMel genome in nine Wolbachia strains that induce different phenotypic effects in their hosts. A comparison of the ANK genes of wMel and the non-CI-inducing wAu Wolbachia strain revealed significant differences between the strains. This was reflected in sequence variability in shared genes that could result in alterations in the encoded proteins, such as motif deletions, amino acid insertions, and in some cases disruptions due to insertion of transposable elements and premature stops. In addition, one wMel ANK gene, which is part of an operon, was absent in the wAu genome. These variations are likely to affect the affinity, function, and cellular location of the predicted proteins encoded by these genes.

Ankyrins↗

Genetic and biochemical factors associated with variation in blood pressure in a genetic isolate.

We previously found an association between blood pressure and genetic variation of angiotensinogen in Canadian Hutterites. We hypothesized that variation in other candidate genes would also be associated with variation in blood pressure. We included genotypes of 12 candidate genes, along with clinical features and biochemical variables as covariates in an association analysis. We found that sex and body mass were significantly associated with variation in both systolic and diastolic blood pressures. We found that genotypes of APOB codon 4154 and AGT codon 174 were significantly associated with variation in systolic blood pressure. We found that genotypes of APOB codon 4154, AGT codon 174, and F7 codon 353 were significantly associated with variation in diastolic blood pressure. We found a significant association between age and variation in systolic but not diastolic blood pressure. We found a significant association between plasma apo B concentration and variation in diastolic but not systolic blood pressure. The association of genomic variation with resting blood pressure is consistent with the existence of important structural elements within or proximal to some genes in lipoprotein metabolism, the renin-angiotensin system, and the coagulation cascade. The association between plasma apo B concentration and diastolic blood pressure suggests that these traits may share some determinants.

Angiotensinogen↗

Pharmacogenomics and its potential impact on drug and formulation development.

Recent advances in genomic research have provided the basis for new insights into the importance of genetic and genomic markers during the different stages of drug development. A new field of research, pharmacogenomics, which studies the relationship between drug effects and the genome, has emerged. Structural pharmacogenomics maps the complete DNA sequences of whole genomes (genotypes) including individual variations, and functional pharmacogenomics assesses the expression levels of thousands of genes in one single experiment. Together, these two areas of pharmacogenomics have generated massive databases, which have become a challenge for the research field of informatics and have fostered a new branch of research, bioinformatics. If skillfully used, the databases generated by pharmacogenomics together with data mining on the Web promise to improve the drug development process in a variety of areas: identification of drug targets, evaluation of toxicity, classification of diseases, evaluation of formulations, assessment of drug response and treatment, post-marketing applications, and development of personalized medicines.

Animals↗

Characterization of phi 12, a bacteriophage related to phi 6: nucleotide sequence of the small and middle double-stranded RNA.

The isolation of additional bacteriophages containing segmented double-stranded RNA genomes has expanded the Cystoviridae family to nine members. Comparing the genomic sequences of these viruses has allowed evaluation of important genetic as well as structural motifs. These comparative studies are resulting in greater understanding of viral evolution and the role played by genetic and structural variation in the assembly mechanisms of the cystoviruses. In this regard, the small and middle double-stranded RNA genomic segments of bacteriophage phi 12 were copied as cDNA and their nucleotide sequences determined. This genome's organization is similar to that of the small and middle segments of bacteriophages phi 6, phi 8, and phi 13. Although there is little similarity in the nucleotide sequences, similarity exists in the amino acid sequence of the lysis cassette proteins to those of phi 6. The host cell attachment proteins are found to have marked similarity to the phi 13 attachment proteins.

Bacteriophage phi 6↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

Genomic structure of human anion exchanger 3 and its potential role in hereditary neurological disease.

Alterations in ion channel permeability or selectivity have been shown to cause neurological defects in humans. Anion exchanger isoform 3 (AE3) is prominently expressed in the brain and performs an electroneutral exchange of chloride and bicarbonate ions. In order to study the potential role of AE3 in human neurological disease, we characterized AE3 genomic structure and performed mutational analysis on patients with an episodic movement disorder that maps to the same genetic locus. AE3 genomic organization, including the nucleotide sequence of the 5'-untranslated region and intron/ exon boundaries, is highly conserved between humans and homologs from mouse and rat. Mutational analysis revealed no disease-causing defect in patients with familial paroxysmal dyskinesia, although several benign polymorphisms were identified. AE3 variation may prove useful for further genetic studies, such as finer resolution mapping. Characterization of genomic structure will facilitate mutational analysis of AE3 in studies of neurological diseases mapped to the same locus.

5' Untranslated Regions↗

Population genomics in natural microbial communities.

Little is known about the evolutionary processes that structure and maintain microbial diversity because, until recently, it was difficult to explore individual-level patterns of variation at the microbial scale. Now, community-genomic sequence data enable such variation to be assessed across large segments of microbial genomes. Here, we discuss how population-genomic analysis of these data can be used to determine how selection and genetic exchange shape the evolution of new microbial lineages. We show that once independent lineages have been identified, such analyses enable the identification of genome changes that drive niche differentiation and promote the coexistence of closely related lineages within the same environment. We suggest that understanding the evolutionary ecology of natural microbial populations through population-genomic analyses will enhance our understanding of genome evolution across all domains of life.

Archaea↗

[Characterization of 5S rRNA gene sequence and secondary structure in gymnosperms].

In higher plants the primary and the secondary structures of 5S ribosomal RNA gene are considered highly conservative. Little is known about the 5S rRNA gene structure, organization and variation in gyimnosperms. In this study we analyzed sequence and structure variation of 5S rRNA gene in Pinus through cloning and sequencing multiple copies of 5S rDNA repeats from individual trees of five pines, P. bungeana, P. tabulaeformis, P. yunnanensis, P. massoniana and P. densata. Pinus bungeana is from the subgenus Strobus while the other four are from the subgenus Pinus (diploxylon pines). Our results revealed variations in both primary and secondary structure among copies of 5S rDNA within individual genomes and between species. 5S rRNA gene in Pinus is 120 bp long in most of the 122 clones we sequenced except for one or two deletions in three clones. Among these clones 50 unique sequences were identified and they were shared by different pine species. Our sequences were compared to 13 sequences each representing a different gymnosperm species, and to six sequences representing both angiosperm monocots and dicots. Average sequence similarity was 97.1% among Pinus species and 94.3% between Pinus and other gymnosperms. Between gymnosperms and angiosperms the sequence similarity decreased to 88.1%. Similar to other molecular data, significant sequence divergence was found between the two Pinus subgenera. The 5S gene tree (neighbor-joining tree) grouped the four diploxylon pines together and separated them distinctly from P. bungeana. Comparison of sequence divergence within individuals and between species suggested that concerted evolution has been very weak especially after the divergence of the four diploxylon pines. The phylogenetic information contained in the 5S rRNA gene is limited due to its shorter length and the difficulties in identifying orthologous and paralogous copies of rDNA multigene family further complicate its phylogenetic application. Pinus densata is a diploid hybrid between P. tabulaeformis and P. yunnanensis. Its 5S rDNA composition is consistent with its hybrid origin. 5S rRNA of all gymnosperms published so far could be folded into a general secondary structure. Variation in this secondary structure was detected among species. About 55% of the 120 bp nucleotide positions was variable, in which 68% was on stem regions. Nevertheless, the positions at the end of the stems and those adjacent to loops are conserved. Their stability directly determines the size of the loops. Some mutations such as compensatory base-pair substitutions, and G-U pairing could be regarded as mechanisms for maintaining a stable secondary structure. The loops of the secondary structure are also relatively conserved. It seems that stable helices are necessary for the function of the gene. The conserved nucleotides in the loops are probably involved in the interaction with proteins and/or RNAs or with other nucleotide in the formation of the tertiary structure. However, unlike other reports, Loop E was found quite mutable among pines. These variations together with those on stems might be caused by the presence of pseudogenes among our clones. A preliminary evaluation indicates that only seven of 50 unique sequences are potentially functional genes.

Base Sequence↗

Optical genome mapping improves clinical interpretation of constitutional copy-number gains and reduces their VUS burden.

PURPOSE: Genomic structure of copy-number gains is critical for their clinical interpretation but cannot be determined by chromosomal microarray (CMA) analysis, which does not provide information about chromosomal location and orientation of multiplied regions. We thus hypothesized that in CMA testing gains have higher probability than losses to be classified as variants of uncertain significance (VUS) and that structural information from optical genome mapping (OGM) may improve their interpretation. METHODS: Using a χ2 test, we assessed the association between classification of copy-number variants as VUS and their type (gains vs losses) in a cohort of 4073 CMA cases. Thirty-three VUS gains involving disease-associated genes were characterized by OGM to evaluate if OGM data enable their more conclusive clinical interpretation. RESULTS: The proportion of variants reported as VUS compared with likely pathogenic/pathogenic was significantly higher for gains than losses, confirming their increased VUS burden. OGM successfully determined genomic structure for all 33 copy-number gains, showing that 26 of 33 were tandem duplications and 7 of 33 were complex rearrangements. Structural information facilitated clinical interpretation in majority of the cases; it supported benign nature for 27 of 33 gains and was inconclusive or supported pathogenic role for 6 of 33. An estimated 20% of reported VUS gains would not have been reportable if we had OGM data. CONCLUSION: We illustrate a specific advantage of OGM compared with CMA: in addition to detecting both copy-number variants and balanced rearrangements, OGM improves clinical interpretation of copy-number gains by providing structural information and is thus expected to significantly decrease their VUS burden.

Humans↗

Perspectives on human genetic variation from the HapMap Project.

The completion of the International HapMap Project marks the start of a new phase in human genetics. The aim of the project was to provide a resource that facilitates the design of efficient genome-wide association studies, through characterising patterns of genetic variation and linkage disequilibrium in a sample of 270 individuals across four geographical populations. In total, over one million SNPs have been typed across these genomes, providing an unprecedented view of human genetic diversity. In this review we focus on what the HapMap Project has taught us about the structure of human genetic variation and the fundamental molecular and evolutionary processes that shape it.

Alleles↗

Involvement of cross-genus phages in bacterial resistance to chlorine disinfection.

Chlorine disinfection resistance in pathogenic microorganisms poses severe environmental concerns and public health risks. While phages play critical roles in host adaptation to environmental stress, how poly-host phages contribute to bacterial resistance to chlorine disinfectants remains poorly understood. Here, we investigated shifts in the population dynamics, transcriptional profiles, and function potentials of cross-genus phage-bacterial communities under exposure to chlorine disinfectants in a continuously operated anaerobic-anoxic-oxic system over a 92-day period, using integrated metagenomic and metatranscriptomic approaches. In the presence and absence of chlorine disinfectants, the genomic abundance and diversity of phage and bacterial communities showed similar variation trends, and the community structures of both exhibited clear differences. A strong significant positive correlation was observed between phage and bacterial diversity under chlorine exposure (R&#x202f;=&#x202f;0.975, p&#x202f;=&#x202f;0.00,057), whereas no significant correlation was detected in the absence of chlorine disinfection (R&#x202f;=&#x202f;-0.314, p&#x202f;=&#x202f;0.613), suggesting that chlorine disinfectants may enhance phage-bacteria interactions. Host-associated phages exhibited high consistency with their corresponding putative hosts in terms of genomic abundance (M2&#x202f;=&#x202f;0.0945, p&#x202f;=&#x202f;0.001) and transcript abundance (M2&#x202f;=&#x202f;0.3668, p&#x202f;=&#x202f;0.001), and they were also significantly correlated with cross-genus phages in both genomic abundance (R&#x202f;=&#x202f;0.97, p&#x202f;<&#x202f;2.2e-16) and transcript abundance (R&#x202f;=&#x202f;0.83, p&#x202f;<&#x202f;2.2e-16), which collectively suggests the critical role of cross-genus phages in the resistance of microbial communities to chlorine disinfectants. Bipartite association network analysis shows that cross-genus phages carry highly homologous genes to their putative hosts and may be involved in the horizontal transfer of these genes among bacteria. These homologous genes are involved in DNA repair, redox balance regulation, environmental stress adaptation and efflux pump functions, suggesting a synergistic role between cross-genus phages and their putative hosts in chlorine resistance. Our findings reveal that cross-genus phages can contribute to the resistance of bacterial communities to chlorine disinfectants, providing the theoretical foundation for evaluating the role of poly-host phages in microbial communities.

Chlorine resistance↗

Capturing genomic signatures of DNA sequence variation using a standard anonymous microarray platform.

Comparative genomics, using the model organism approach, has provided powerful insights into the structure and evolution of whole genomes. Unfortunately, only a small fraction of Earth's biodiversity will have its genome sequenced in the foreseeable future. Most wild organisms have radically different life histories and evolutionary genomics than current model systems. A novel technique is needed to expand comparative genomics to a wider range of organisms. Here, we describe a novel approach using an anonymous DNA microarray platform that gathers genomic samples of sequence variation from any organism. Oligonucleotide probe sequences placed on a custom 44 K array were 25 bp long and designed using a simple set of criteria to maximize their complexity and dispersion in sequence probability space. Using whole genomic samples from three known genomes (mouse, rat and human) and one unknown (Gonystylus bancanus), we demonstrate and validate its power, reliability, transitivity and sensitivity. Using two separate statistical analyses, a large numbers of genomic 'indicator' probes were discovered. The construction of a genomic signature database based upon this technique would allow virtual comparisons and simple queries could generate optimal subsets of markers to be used in large-scale assays, using simple downstream techniques. Biologists from a wide range of fields, studying almost any organism, could efficiently perform genomic comparisons, at potentially any phylogenetic level after performing a small number of standardized DNA microarray hybridizations. Possibilities for refining and expanding the approach are discussed.

Animals↗

Genome size, quantitative genetics and the genomic basis for flower size evolution in Silene latifolia.

BACKGROUND AND AIMS: The overall goal of this paper is to construct an overview of the genetic basis for flower size evolution in Silene latifolia. It aims to examine the relationship between the molecular bases for flower size and the underlying assumption of quantitative genetics theory that quantitative variation is ultimately due to the impact of a number of structural genes. SCOPE: Previous work is reviewed on the quantitative genetics and potential for response to selection on flower size, and the relationship between flower size and nuclear DNA content in S. latifolia. These earlier findings provide a framework within which to consider more recent analyses of a joint quantitative trait loci (QTL) analysis of flower size and DNA content in this species. KEY RESULTS: Flower size is a character that fits the classical quantitative genetics model of inheritance very nicely. However, an earlier finding that flower size is correlated with nuclear DNA content suggested that quantitative aspects of genome composition rather than allelic substitution at structural loci might play a major role in the evolution of flower size. The present results reported here show that QTL for flower size are correlated with QTL for DNA content, further corroborating an earlier result and providing additional support for the conclusion that localized variations in DNA content underlie evolutionary changes in flower size. CONCLUSIONS: The search image for QTL should be broadened to include overall aspects of genome regulation. As we prepare to enter the much-heralded post-genomic era, we also need to revisit our overall models of the relationship between genotype and phenotype to encompass aspects of genome structure and composition beyond structural genes.

Biological Evolution↗