PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Characterization of phi 12, a bacteriophage related to phi 6: nucleotide sequence of the small and middle double-stranded RNA.

The isolation of additional bacteriophages containing segmented double-stranded RNA genomes has expanded the Cystoviridae family to nine members. Comparing the genomic sequences of these viruses has allowed evaluation of important genetic as well as structural motifs. These comparative studies are resulting in greater understanding of viral evolution and the role played by genetic and structural variation in the assembly mechanisms of the cystoviruses. In this regard, the small and middle double-stranded RNA genomic segments of bacteriophage phi 12 were copied as cDNA and their nucleotide sequences determined. This genome's organization is similar to that of the small and middle segments of bacteriophages phi 6, phi 8, and phi 13. Although there is little similarity in the nucleotide sequences, similarity exists in the amino acid sequence of the lysis cassette proteins to those of phi 6. The host cell attachment proteins are found to have marked similarity to the phi 13 attachment proteins.

Bacteriophage phi 6↗

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.

Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.

Proteins↗

Genomic structure of human anion exchanger 3 and its potential role in hereditary neurological disease.

Alterations in ion channel permeability or selectivity have been shown to cause neurological defects in humans. Anion exchanger isoform 3 (AE3) is prominently expressed in the brain and performs an electroneutral exchange of chloride and bicarbonate ions. In order to study the potential role of AE3 in human neurological disease, we characterized AE3 genomic structure and performed mutational analysis on patients with an episodic movement disorder that maps to the same genetic locus. AE3 genomic organization, including the nucleotide sequence of the 5'-untranslated region and intron/ exon boundaries, is highly conserved between humans and homologs from mouse and rat. Mutational analysis revealed no disease-causing defect in patients with familial paroxysmal dyskinesia, although several benign polymorphisms were identified. AE3 variation may prove useful for further genetic studies, such as finer resolution mapping. Characterization of genomic structure will facilitate mutational analysis of AE3 in studies of neurological diseases mapped to the same locus.

5' Untranslated Regions↗

[Characterization of 5S rRNA gene sequence and secondary structure in gymnosperms].

In higher plants the primary and the secondary structures of 5S ribosomal RNA gene are considered highly conservative. Little is known about the 5S rRNA gene structure, organization and variation in gyimnosperms. In this study we analyzed sequence and structure variation of 5S rRNA gene in Pinus through cloning and sequencing multiple copies of 5S rDNA repeats from individual trees of five pines, P. bungeana, P. tabulaeformis, P. yunnanensis, P. massoniana and P. densata. Pinus bungeana is from the subgenus Strobus while the other four are from the subgenus Pinus (diploxylon pines). Our results revealed variations in both primary and secondary structure among copies of 5S rDNA within individual genomes and between species. 5S rRNA gene in Pinus is 120 bp long in most of the 122 clones we sequenced except for one or two deletions in three clones. Among these clones 50 unique sequences were identified and they were shared by different pine species. Our sequences were compared to 13 sequences each representing a different gymnosperm species, and to six sequences representing both angiosperm monocots and dicots. Average sequence similarity was 97.1% among Pinus species and 94.3% between Pinus and other gymnosperms. Between gymnosperms and angiosperms the sequence similarity decreased to 88.1%. Similar to other molecular data, significant sequence divergence was found between the two Pinus subgenera. The 5S gene tree (neighbor-joining tree) grouped the four diploxylon pines together and separated them distinctly from P. bungeana. Comparison of sequence divergence within individuals and between species suggested that concerted evolution has been very weak especially after the divergence of the four diploxylon pines. The phylogenetic information contained in the 5S rRNA gene is limited due to its shorter length and the difficulties in identifying orthologous and paralogous copies of rDNA multigene family further complicate its phylogenetic application. Pinus densata is a diploid hybrid between P. tabulaeformis and P. yunnanensis. Its 5S rDNA composition is consistent with its hybrid origin. 5S rRNA of all gymnosperms published so far could be folded into a general secondary structure. Variation in this secondary structure was detected among species. About 55% of the 120 bp nucleotide positions was variable, in which 68% was on stem regions. Nevertheless, the positions at the end of the stems and those adjacent to loops are conserved. Their stability directly determines the size of the loops. Some mutations such as compensatory base-pair substitutions, and G-U pairing could be regarded as mechanisms for maintaining a stable secondary structure. The loops of the secondary structure are also relatively conserved. It seems that stable helices are necessary for the function of the gene. The conserved nucleotides in the loops are probably involved in the interaction with proteins and/or RNAs or with other nucleotide in the formation of the tertiary structure. However, unlike other reports, Loop E was found quite mutable among pines. These variations together with those on stems might be caused by the presence of pseudogenes among our clones. A preliminary evaluation indicates that only seven of 50 unique sequences are potentially functional genes.

Base Sequence↗

Optical genome mapping improves clinical interpretation of constitutional copy-number gains and reduces their VUS burden.

PURPOSE: Genomic structure of copy-number gains is critical for their clinical interpretation but cannot be determined by chromosomal microarray (CMA) analysis, which does not provide information about chromosomal location and orientation of multiplied regions. We thus hypothesized that in CMA testing gains have higher probability than losses to be classified as variants of uncertain significance (VUS) and that structural information from optical genome mapping (OGM) may improve their interpretation. METHODS: Using a χ2 test, we assessed the association between classification of copy-number variants as VUS and their type (gains vs losses) in a cohort of 4073 CMA cases. Thirty-three VUS gains involving disease-associated genes were characterized by OGM to evaluate if OGM data enable their more conclusive clinical interpretation. RESULTS: The proportion of variants reported as VUS compared with likely pathogenic/pathogenic was significantly higher for gains than losses, confirming their increased VUS burden. OGM successfully determined genomic structure for all 33 copy-number gains, showing that 26 of 33 were tandem duplications and 7 of 33 were complex rearrangements. Structural information facilitated clinical interpretation in majority of the cases; it supported benign nature for 27 of 33 gains and was inconclusive or supported pathogenic role for 6 of 33. An estimated 20% of reported VUS gains would not have been reportable if we had OGM data. CONCLUSION: We illustrate a specific advantage of OGM compared with CMA: in addition to detecting both copy-number variants and balanced rearrangements, OGM improves clinical interpretation of copy-number gains by providing structural information and is thus expected to significantly decrease their VUS burden.

Humans↗

Involvement of cross-genus phages in bacterial resistance to chlorine disinfection.

Chlorine disinfection resistance in pathogenic microorganisms poses severe environmental concerns and public health risks. While phages play critical roles in host adaptation to environmental stress, how poly-host phages contribute to bacterial resistance to chlorine disinfectants remains poorly understood. Here, we investigated shifts in the population dynamics, transcriptional profiles, and function potentials of cross-genus phage-bacterial communities under exposure to chlorine disinfectants in a continuously operated anaerobic-anoxic-oxic system over a 92-day period, using integrated metagenomic and metatranscriptomic approaches. In the presence and absence of chlorine disinfectants, the genomic abundance and diversity of phage and bacterial communities showed similar variation trends, and the community structures of both exhibited clear differences. A strong significant positive correlation was observed between phage and bacterial diversity under chlorine exposure (R&#x202f;=&#x202f;0.975, p&#x202f;=&#x202f;0.00,057), whereas no significant correlation was detected in the absence of chlorine disinfection (R&#x202f;=&#x202f;-0.314, p&#x202f;=&#x202f;0.613), suggesting that chlorine disinfectants may enhance phage-bacteria interactions. Host-associated phages exhibited high consistency with their corresponding putative hosts in terms of genomic abundance (M2&#x202f;=&#x202f;0.0945, p&#x202f;=&#x202f;0.001) and transcript abundance (M2&#x202f;=&#x202f;0.3668, p&#x202f;=&#x202f;0.001), and they were also significantly correlated with cross-genus phages in both genomic abundance (R&#x202f;=&#x202f;0.97, p&#x202f;<&#x202f;2.2e-16) and transcript abundance (R&#x202f;=&#x202f;0.83, p&#x202f;<&#x202f;2.2e-16), which collectively suggests the critical role of cross-genus phages in the resistance of microbial communities to chlorine disinfectants. Bipartite association network analysis shows that cross-genus phages carry highly homologous genes to their putative hosts and may be involved in the horizontal transfer of these genes among bacteria. These homologous genes are involved in DNA repair, redox balance regulation, environmental stress adaptation and efflux pump functions, suggesting a synergistic role between cross-genus phages and their putative hosts in chlorine resistance. Our findings reveal that cross-genus phages can contribute to the resistance of bacterial communities to chlorine disinfectants, providing the theoretical foundation for evaluating the role of poly-host phages in microbial communities.

Chlorine resistance↗

De novo genome assemblies of threatened Asian hornbills (Bucerotidae) reveal declining population trajectories during the late Pleistocene.

BACKGROUND: Asian hornbills are flagship species of the wet tropics that face significant threats from hunting, habitat loss, and fragmentation. Despite being conservation flagships, whole genome information is available for only two of the 32 Asian hornbill species. In this study, we provide the first de novo genome assemblies for four hornbill species (Bucerotidae) in Asia. METHODS: We used a combination of long-read and short-read sequencing data to assemble and annotate de novo hybrid genomes of four species of hornbills. We also assembled and compared mitochondrial genomes of these species. Using a comparative genomics approach, we performed orthology assignment and gene evolution analyses to identify unique gene families in Asian hornbills, gene families that showed significant expansion, their functions and structural variation. Furthermore, using the Pairwise Sequentially Markov Coalescent (PSMC) method, we reconstructed demographic histories of hornbill species to examine changes in their population trajectories in the past. RESULTS: We present hybrid genome assemblies for Great Hornbill (B. bicornis - GH), Rufous-necked Hornbill (A. nipalensis- RNH), Malabar Pied Hornbill (A. coronatus- MPH) and Wreathed Hornbill (R. undulatus- WH). The genome sizes of these hornbills range from 1.1 Gb to 1.3 Gb, with over 95.9% completeness and gene prediction BUSCO. We reported 10,525 orthogroups shared among four Asian hornbill species and identified significant expansion in gene families associated with structural keratin development in Asian hornbills compared to their ancestors. We also provide annotated mitogenomes for each of these species. Furthermore, we found that the WH, a more abundant, widely distributed, and migratory species, showed a higher Ne than the other three hornbill species. However, an overall decline in Ne for all species was recorded during the Pleistocene climatic fluctuations. CONCLUSIONS: We present the first-ever, high-quality reference genomes for the threatened hornbill species from Asia. Hornbills have shown significant expansion in genes involved in structural keratin development. Our results indicate that Pleistocene climatic fluctuations have led to dramatic population declines in all four species. We believe that this study provides robust genomic resources to support future comparative and conservation genomics efforts for hornbills.

Animals↗

Haplotype structure and population genetic inferences from nucleotide-sequence variation in human lipoprotein lipase.

Allelic variation in 9.7 kb of genomic DNA sequence from the human lipoprotein lipase gene (LPL) was scored in 71 healthy individuals (142 chromosomes) from three populations: African Americans (24) from Jackson, MS; Finns (24) from North Karelia, Finland; and non-Hispanic Whites (23) from Rochester, MN. The sequences had a total of 88 variable sites, with a nucleotide diversity (site-specific heterozygosity) of .002+/-.001 across this 9.7-kb region. The frequency spectrum of nucleotide variation exhibited a slight excess of heterozygosity, but, in general, the data fit expectations of the infinite-sites model of mutation and genetic drift. Allele-specific PCR helped resolve linkage phases, and a total of 88 distinct haplotypes were identified. For 1,410 (64%) of the 2,211 site pairs, all four possible gametes were present in these haplotypes, reflecting a rich history of past recombination. Despite the strong evidence for recombination, extensive linkage disequilibrium was observed. The number of haplotypes generally is much greater than the number expected under the infinite-sites model, but there was sufficient multisite linkage disequilibrium to reveal two major clades, which appear to be very old. Variation in this region of LPL may depart from the variation expected under a simple, neutral model, owing to complex historical patterns of population founding, drift, selection, and recombination. These data suggest that the design and interpretation of disease-association studies may not be as straightforward as often is assumed.

Animals↗

A 9.1-kb gap in the genome reference map is shown to be a stable deletion/insertion polymorphism of ancestral origin.

We show a mute 9.1-kb gap in the human genome reference map, unraveled by RDA studies, to be a worldwide deletion/insertion polymorphism of stable type. The molecular and population data presented suggest its origin from a unique ancestral transposition event in chromosomal region 22q11.2, overlapping the IglambdaV genes at about 450 kb from the cluster of the IglambdaJ-C genes. These findings are not meant to be just another report of a polymorphic marker suitable for population studies. Rather, we wish to stress that a large number of inborn mute gaps may be spread all over the genome and that the many RDA-detected microdeletions already available are efficient tools for the discovery of this otherwise hidden category of genetic variation. Apart from their possible impact on expression of structural genes, mute gaps must be filled for the reference map of our genome to be truly completed.

Chromosome Deletion↗

The study of variation in the human genome.

Regions of the genome showing high evolutionary stability are often conserved as a result of functional constraints. Conversely, more variable regions are likely to represent DNA with no functional or structural importance. However, as in the case of immunologically important regions, sequence divergence does not always indicate lack of functional importance. There is thus a wealth of information from both a functional and an evolutionary point of view that comes from studies of DNA sequence variation, a neglected aspect of the genome endeavor. Naturally, one cannot sequence hundreds of individuals in full, but a useful compromise is to use less expensive methods and to limit the more expensive types of analysis to an appropriately chosen sample of loci. The sample could be determined after careful consideration of categories of DNA segments with respect to individual variation. The study of such categories of DNA variation patterns can help in the understanding of the role of each gene and vice versa. One other important application requiring a study of DNA variation in different human populations is forensic DNA typing. This study requires a knowledge of allele frequencies in different human populations. Evidence of a match between two DNA samples is meaningless if the approximate population frequency of the DNA pattern is not known. It has been suggested (E. Lander) that one use the highest frequency for the most common allele as a baseline frequency estimate. Obviously, systems in which this is employed require an extensive analysis of population-specific allele frequencies. In general, the best way of studying interindividual variation when detecting or describing new polymorphisms is to include interethnic variation.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗

Heterogeneity in rates of recombination in the 6-Mb region telomeric to the human major histocompatibility complex.

Analysis of 784 informative meioses in the CEPH pedigrees revealed a total of 22 recombination events having occurred in the 6-Mb region between D6S265 (70 kb centromeric of HLA-A) and D6S276. These 22 breakpoints were localized with respect to anonymous polymorphic markers, leading to a detailed genetic map of the region telomeric to the human major histocompatibility complex. A nonrandom pattern of recombination was observed throughout this region: the low recombination rate of 0.19% within the 4-Mb interval centromeric to the HLA class I-like candidate gene for hemochromatosis indeed contrasts with the approximate 1% rate observed within the most telomeric two megabases. This reduced rate of recombination may be due to selective constraints depending on environmental factors related to immunity and iron status or to structural variations hampering proper meiotic pairing of homologous sequences. Population data from other human genome segments are now needed to determine whether linkage disequilibrium extending over 4 Mb is unique to this region.

Chromosome Mapping↗

Identification of two distinct subfamilies of alpha satellite DNA that are highly specific for human chromosome 15.

We report the isolation of two distinct subfamilies of alpha satellite DNA (pTRA-20 and -25) from human chromosome 15. In situ hybridization experiments indicated that both subfamilies are highly specific for this chromosome. Southern analysis of a somatic hybrid cell line carrying human chromosome 15 revealed a likely higher-order genomic band of 2.5 kb for pTRA-20. Similar analysis for pTRA-25 showed multiple higher-order bands of 3.5, 4.5, and 5 kb at moderately high hybridization stringency, but a predominance of the 4.5-kb species at very high stringency. Direct comparison with human genomic DNA confirmed the authenticity of these higher-order structures and demonstrated polymorphic variations using both probes. The origin of the different alphoid subfamilies on chromosome 15 is discussed. These sequences should be useful for the construction of centromere-based genetic linkage maps for human chromosome 15 and, in conjunction with the other alphoid sequences already reported for chromosomes 13, 14, 21, and 22, should allow a concerted analysis of the evolution and the possible etiological role of these DNAs in aberrations commonly seen in these chromosomes.

Blotting, Southern↗

Genotypic characteristics of bovine viral diarrhea virus 2 strains isolated in northern Italy.

Two strains of Bovine viral diarrhea virus 2 (BVDV-2) were isolated from calves in northern Italy. Variations in the 5'-untranslated region (UTR) of the genome were studied by primary structure alignment and neighbor-joining method based phylogenetic tree analyses and by palindromic nucleotide substitutions at the three variable loci in the 5'-UTR. Genetic analysis indicated their appurtenance to genovar BVDV-2a. Nucleotide sequence at the 5'-UTR of strain BS-95-II, one of the Italian isolates from healthy calves, showed 98% homology to that of the Japanese isolate OY89, a cytopathic strain derived from cattle with mucosal disease.

5' Untranslated Regions↗

Poliovirus type 3/Saukett: antigenic and structural correlates of sequence variation in the capsid proteins.

The Saukett/USA/50 strain is the type 3 component of the inactivated poliovirus vaccine. The capsid-coding region of genomic RNA of Saukett strains from five different sources was sequenced and the sequence differences were correlated with antigenic differences measurable with poliovirus type 3-specific neutralizing monoclonal antibodies. All strains appeared to have capsid protein genes identical in size to those of the entirely sequenced type 3 poliovirus strains. The nucleotide sequence identity between the strains was 91% on the average and the strains could be divided into three groups. Amino acid differences were seen in 30 positions located throughout the capsid region both within and outside the known antigenic sites. Substitutions at the known antigenic sites explained most of the observed antigenic differences. Use of the atomic coordinates of the crystal structure model of the Sabin 3 virus and prior data based on escape mutants and peptide scanning revealed that most of the exposed substitutions located outside the known antigenic sites are spatially associated with regions found to be antigenic by either or both of these methods.

Antigenic Variation↗

High guanine-cytosine content is not an adaptation to high temperature: a comparative analysis amongst prokaryotes.

The causes of the variation between genomes in their guanine (G) and cytosine (C) content is one of the central issues in evolutionary genomics. The thermal adaptation hypothesis conjectures that, as G:C pairs in DNA are more thermally stable than adenonine:thymine pairs, high GC content may he a selective response to high temperature. A compilation of data on genomic GC content and optimal growth temperature for numerous prokaryotes failed to demonstrate the predicted correlation. By contrast, the GC content of Structural RNAs is higher at high temperatures. The issue that we address here is whether more freely evolving sites in exons (i.e. codonic third positions) evolve in the same manner as genomic DNA as a whole, Showing no correlated response, or like structural RNAs showing a strong correlation. The latter pattern would provide strong support for the thermal adaptation hypothesis, as the variation in GC content between orthologous genes is typically most profoundly seen at codon third sites (GC3). Simple analysis of completely sequenced prokaryotic genomes shows that GC3, but not genomic GC, is higher on average in thermophilic species. This demonstrates, if nothing else, that the results from the two measures cannot be presumed to be the same. A proper analysis, however, requires phylogenetic control. Here, therefore, we report the results of a comparative analysis of GC composition and optimal growth temperature for over 100 prokaryotes. Comparative analysis fails to show, in either Archea or Eubacteria, any hint of connection between optimal growth temperature and GC content in the genome as a whole, in protein-coding regions or, more crucially at GC. Conversely, comparable analysis confirms that GC content of structural RNA is strongly correlated with optimal temperature. Against the expectations of the thermal adaptation hypothesis, within prokaryotes GC content in protein-coding genies, even at relatively freely evolving sites, cannot be considered an adaptation to the thermal environment.

Adaptation, Physiological↗

A single-nucleotide natural variation (U4 to C4) in an influenza A virus promoter exhibits a large structural change: implications for differential viral RNA synthesis by RNA-dependent RNA polymerase.

The influenza A virus promoter is recognized by the influenza A virus RNA-dependent RNA polymerase, and directs both transcription and replication of the viral RNA genome. Within the sequence of this promoter, flu strains exhibit a natural, unique variation, either a U or a C, at the fourth position from the 3' end. Promoters that contain a C residue (C4 promoter), which are invariably found in genome segments that encode the three RNA polymerase subunits (PB1, PB2 and PA), down-regulate transcription but activate genome replication. Here, we have determined the structure of the C4 promoter by NMR spectroscopy and compared it with the structure of the U4 promoter, which was determined previously. The structure of the internal loop in the C4 promoter is similar to that of the U4 promoter. However, the terminal stem of the C4 promoter is strikingly different from that of the U4 promoter. These structural data suggest that the internal loop is important for polymerase binding to the promoter, and the terminal stem is crucial for differential regulation of transcription and replication.

Base Sequence↗

Structural variation of the pseudoautosomal region between and within inbred mouse strains.

The pseudoautosomal region (PAR) is a segment of shared homology between the sex chromosomes. Here we report additional probes for this region of the mouse genome. Genetic and fluorescence in situ hybridization analyses indicate that one probe, PAR-4, hybridizes to the pseudoautosomal telomere and a minor locus at the telomere of chromosome 9 and that a PCR assay based on the PAR-4 sequence amplifies only the pseudoautosomal locus (DXYHgu1). The region detected by PAR-4 is structurally unstable; it shows polymorphism both between mouse strains and between animals of the same inbred strain, which implies an unusually high mutation rate. Variation occurs in the region adjacent to a (TTAGGG)n array. Two pseudoautosomal probes can also hybridize to the distal telomeres of chromosomes 9 and 13, and all three telomeres contain DXYMov15. The similarity between these telomeres may reflect ancestral telomere-telomere exchange.

Animals↗