PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

Haplotype and linkage disequilibrium architecture for human cancer-associated genes.

To facilitate association-based linkage studies we have studied the linkage disequilibrium (LD) and haplotype architecture around five genes of interest for cancer risk: ATM, BRCA1, BRCA2, RAD51, and TP53. Single nucleotide polymorphisms (SNPs) were identified and used to construct haplotypes that span 93-200 kb per locus with an average SNP density of 12 kb. These markers were genotyped in four ethnically defined populations that contained 48 each of African Americans, Asian Americans, Hispanic Americans, and European Americans. Haplotypes were inferred using an expectation maximization (EM) algorithm, and the data were analyzed using D', R(2), Fisher's exact P-values, and the four-gamete test for recombination. LD levels varied widely between loci from continuously high LD across 200 kb to a virtual absence of LD across a similar length of genome. LD structure also varied at each gene and between populations studied. This variation indicates that the success of linkage-based studies will require a precise description of LD at each locus and in each population to be studied. One striking consistency between genes was that at each locus a modest number of haplotypes present in each population accounted for a high fraction of the total number of chromosomes. We conclude that each locus has its own genomic profile with regard to LD, and despite this there is the widespread trend of relatively low haplotype diversity. As a result, a low marker density should be adequate to identify haplotypes that represent the common variation at a locus, thereby decreasing costs and increasing efficacy of association studies.

Alleles↗

Polymorphic variations in the ori sequences from the mitochondrial genomes of different wild-type yeast strains.

We determined the restriction maps and primary structures of two as yet poorly characterized regions of the mitochondrial genomes of different wild-type strains of Saccharomyces cerevisiae. These regions respectively comprised the ori1 sequence and the newly identified ori8 sequence. Ori1 and ori8, together with their flanking sequences, exhibit a large polymorphism, resulting from specific variations due to insertions or deletions of optional GC clusters at different locations. The mechanisms underlying such sequence rearrangements are discussed.

Base Sequence↗

Integrating Optical Genome Mapping into the Genetic Diagnostic Algorithm: Clinical Utility in Unresolved Autosomal Recessive Disorders from a Large Cohort.

INTRODUCTION: The identification of precise genetic etiologies is indispensable for the clinical management of monogenic disorders. However, conventional diagnostic methods and exome sequencing (ES) frequently fail to identify complex structural variations (SVs), leaving the genetic basis unexplained in approximately 30-60% of suspected cases. Optical genome mapping (OGM) emerges as a high-resolution technology capable of detecting cryptic SVs inaccessible to standard methodologies. METHODS: In this study, we evaluated the clinical utility of integrating OGM into the diagnostic algorithm for unresolved monogenic diseases. Following negative or inconclusive results from standard ES pipelines, OGM was applied to a targeted subset of patients (n = 7) selected from a comprehensive clinical cohort of 1,257 individuals with suspected genetic disorders. RESULTS: The integration of OGM identified candidate SVs that may represent the second allelic alteration in two distinct cases; however, confirmation through parental segregation analysis remains pending. Specifically, OGM identified an intronic insertion in the TTLL5 gene and a deletion in a putative regulatory region approximately 400 kb upstream of the NMNAT1 gene, both of which were missed by prior diagnostic testing. CONCLUSION: Our findings suggest that OGM has potential value in investigating the missing heritability of autosomal recessive disorders. By detecting candidate SVs invisible to conventional methods, OGM may warrant consideration as a complementary diagnostic approach following inconclusive ES; however, larger cohorts and confirmatory functional studies are needed to establish its clinical utility.

Autosomal recessive disorders↗

Cladistic structure within the human Lipoprotein lipase gene and its implications for phenotypic association studies.

Haplotype variation in 9.7 kb of genomic DNA sequence from the human lipoprotein lipase (LPL) gene was scored in three populations: African-Americans from Jackson, Mississippi (24 individuals), Finns from North Karelia, Finland (24), and non-Hispanic whites from Rochester, Minnesota (23). Earlier analyses had indicated that recombination was common but concentrated into a hotspot and that recurrent mutations at multiple sites may have occurred. We show that much evolutionary structure exists in the haplotype variation on either side of the recombinational hotspot. By peeling off significant recombination events from a tree estimated under the null hypothesis of no recombination, we also reveal some cladistic structure not disrupted by recombination during the time to coalescence of this variation. Additional cladistic structure is estimated to have emerged after recombination. Many apparent multiple mutational events at sites still remain after removing the effects of the detected recombination/gene conversion events. These apparent multiple events are found primarily at sites identified as highly mutable by previous studies, strengthening the conclusion that they are true multiple events. This analysis portrays the complexity of the interplay among many recombinational and mutational events that would be needed to explain the patterns of haplotype diversity in this gene. The cladistic structure in this region is used to identify four to six single-nucleotide polymorphisms (SNPs) that would provide disequilibrium coverage over much of this region. These sites may be useful in identifying phenotypic associations with variable sites in this gene. Evolutionary considerations also imply that the SNPs in the 3' region should have general utility in most human populations, but the 5' SNPs may be more population specific. Choosing SNPs at random would generally not provide adequate disequilibrium coverage of the sequenced region.

Black or African American↗

Four Arabidopsis RPP loci controlling resistance to the Noco2 isolate of Peronospora parasitica map to regions known to contain other RPP recognition specificities.

Interactions between Arabidopsis thaliana and the downy mildew fungus Peronospora parasitica provide a model system to study the genetic and molecular basis of plant-pathogen recognition. With the use of the Noco2 isolate of P. parasitica, the reaction phenotypes of 46 accessions of Arabidopsis were examined and 31 accessions exhibited resistance. Resistance phenotypes examined ranged from distinct necrotic pits or flecks to a weak necrosis accompanied by late and sparse fungal sporulation. Segregating populations generated from crosses between the susceptible accession Col-0 and the resistant accessions Ws-0, Pr-0, Oy-0, Po-1, Bch-1, Ge-1, Di-1, Ji-1, and Te-0 were also screened with Noco2. The genetic data were consistent with the presence of single resistance (RPP) loci in all of these accessions except Oy-0, in which resistance was inherited as a digenic trait. As a first step to molecular cloning, the map positions of four resistance loci were determined. These have been designated RPP14.1 from Ws-0, RPP14.2 from Pr-O, and RPP14.3 and RPP5.2 from Oy-0. RPP14.1 was mapped to a 3.2-cM interval on chromosome 3 that is linked to a region between the markers Gl-1 and m249 known to contain other P. parasitica resistance specificities. RPP14.2 from Pr-0 and RPP14.3 from Oy-0 were also positioned in this interval. Moreover, RPP14.1 and RPP14.2 showed linkage of < 0.05 cM, suggesting possible allelism. The second RPP locus from Oy-0, RPP5.2, was located on chromosome 4 and exhibited strong linkage (< 2 cM) to RRP5.1, a locus previously identified in the Arabidopsis accession Landsberg-erecta. The results reinforce evidence for RPP gene clustering in the Arabidopsis genome and provide new targets for cloning and examination of RPP gene structure, function, allelic variation, and organization within defined loci.

Alleles↗

Comparative genomics reveals population structure and functional differentiation in Limosilactobacillus fermentum.

Limosilactobacillus fermentum is a widely distributed lactic acid bacterium frequently detected in fermented foods and host-associated microbiota, yet its global genomic diversity and functional variability remain insufficiently characterized. Here, we performed a large-scale comparative genomic analysis of 336 high-quality L. fermentum genomes curated from public databases. Species identity was validated using average nucleotide identity (ANI), and population structure was examined using pairwise ANI comparisons together with Mash-based phylogenetic reconstruction. Clustering at &#x2265;&#x2009;99% ANI resolved the dataset into 15 genomic clusters, with four dominant lineages comprising the majority of genomes. Pangenome reconstruction identified 5,853 gene clusters, including 1,325 core genes (22.6%) and a large accessory component dominated by low-frequency genes. Heap's law modeling (&#x3bb;&#x2009;=&#x2009;0.19) indicated a weakly open pangenome, suggesting ongoing gene acquisition as additional genomes are sampled. Functional annotation revealed that core genes were primarily associated with essential cellular processes, whereas accessory genes were enriched in carbohydrate metabolism, membrane-associated functions, and defense-related systems. Variation in carbohydrate-active enzymes (CAZymes), transport systems, and stress-response genes was observed across lineages, indicating strain-level functional diversity. Although genomes from human and food sources were broadly distributed across phylogenetic lineages, multivariate analysis showed that gene-content variation was more strongly associated with genomic lineage than with isolation source. These results provide a population genomic framework for understanding genomic diversity and functional potential in L. fermentum.

Phylogeny↗

Variations of endogenous chicken proviruses: characterization of new loci of endogenous proviruses in the genome of Italian partridge chickens.

The composition and structure of endogenous proviruses present in the genome of Italian Partridge chickens were studied by the method of blot hybridization using RAV-2 [32P]DNA or LTR of RSV as hybridization probes. The genomes of 5 out of 39 chickens analyzed did not contain endogenous proviruses related to RAV-2. Different sets of five so far undescribed endogenous proviruses, differing in the structure and location, were detected in the DNA of other IP chickens. None of them is identical in its structure to the DNA of the endogenous chicken virus RAV-0, all five loci of endogenous proviruses of IP chickens were defective. The origin, the patterns of genetic variation and the function of endogenous proviruses are discussed.

Animals↗

A variation in the structure of the protein-coding region of the human p53 gene.

An extensive analysis of genomic DNA preparations from a number of normal and malignant tissues revealed BglII site polymorphism of the human p53 gene. Approximately 10% of p53 gene alleles were found to contain an additional BglII site localized in a region of intron I. This allelic form of p53 gene was also responsible for p53 protein having altered electrophoretic mobility. Molecular cloning and sequencing of both the alleles of p53 gene revealed a base-pair change in codon 72 causing arginine----proline substitution in the allele with the additional BglII site. Both variants of the p53 gene may occur in homozygous state and are therefore functional.

Amino Acid Sequence↗

Genetic relatedness of hepatitis B viral strains of diverse geographical origin and natural variations in the primary structure of the surface antigen.

A 681 nucleotide fragment of the hepatitis B virus (HBV) genome was sequenced that corresponded to the complete gene for hepatitis B surface antigen (HBsAg) in 80 HBsAg- and hepatitis B e antigen (HBeAg)-positive sera of diverse geographical origins. These and 42 previously published HBV sequences within the S gene were used for the construction of a dendrogram. In this comparison, each of the 122 HBsAg genes was found to be related to one or other of the six previously identified genomic groups of HBV, A to F. The HBV strains within each genomic group showed a characteristic geographical distribution. Group A genomes were represented by 23 strains mainly originating in northern Europe and sub-Saharan Africa. The group B and C genomes, represented by 17 and 28 strains respectively, were confined to populations with origins in eastern Asia and the Far East. The group D genomes, represented by 38 strains, were found worldwide, but were the predominant strains in the Mediterranean area, the Near and Middle East, and in south Asia. Group E genomes, represented by nine strains, were indigenous to western sub-Saharan Africa as far south as Angola. There were indications that the F group, made up of six strains, represented the genomic group of HBV among populations with origins in the New World. Thus, HBV has diverged into genomic groups according to the distribution of mankind in the different continents. As well as giving information on the genetic relationship of HBV strains of different geographical origin, this study also provides information on the primary structure of HBsAg in different regions of the world. Such data might prove valuable in explaining the reported failures to obtain protection with current HBV vaccines.

Amino Acid Sequence↗

Genotypic heterogeneity within Giardia lamblia isolates demonstrated by M13 DNA fingerprinting.

There has been considerable speculation regarding the possible relationship between the phenotypic and genotypic heterogeneity seen among human isolates of Giardia lamblia and the wide clinical spectrum of human giardiasis. Several workers have suggested that human giardiasis may be a mixed infection consisting of variant strains or subgroups which are present in the same infection and which are selectable, but it is not clear whether these apparent variant strains represent a truly heterogeneous infection or whether the genotypic heterogeneity observed is due to the susceptibility of the Giardia genome to a high rate of structural genetic rearrangement. We have therefore studied variation in Giardia intestinalis genotypes in 19 isolates in vitro and in vivo by using the technique of M13 DNA fingerprinting. Genotypes of isolates changed with time when cultured under standard conditions and when pressured with bile. Sequential isolates and their clones taken from a patient with chronic giardiasis both before and after several treatments with metronidazole had different genotypes. Finally, clones of isolate WB had different initial genotypes, which changed after 4 months in culture. These findings suggest that the apparent genotypic heterogeneity at least in these G. intestinalis isolates is more likely to be due to the plasticity of the Giardia genome than to the presence of a truly mixed population of strains within the same infection.

Animals↗

PangyPlot: multi-scale interactive visualization of pangenome variation graphs.

SUMMARY: Pangenome variation graphs integrate multiple samples into a unified representation, mitigating the reference bias inherent to linear genomes. However, these graphs can be large and structurally complex. Existing visualization tools are each confined to a fixed scale of resolution, requiring researchers to switch between multiple tools to examine variation at different levels of detail. PangyPlot is an interactive pangenome browser designed for multi-scale exploration of reference variation graphs from full chromosome to nucleotide-level sequence segments. PangyPlot anchors navigation to linear reference coordinates, organizes variation into hierarchical bubble structures, and uses a force-directed layout engine for automatic node arrangement. AVAILABILITY AND IMPLEMENTATION: An instance preloaded with data is available at https://pangyplot.research.sickkids.ca. Source code and documentation are openly available at https://github.com/strug-hub/pangyplot under the MIT License.

Software↗

Giant G+C% mosaic structures of the human genome found by arrangement of GenBank human DNA sequences according to genetic positions.

To determine the overall variation in the G+C% distribution over long ranges of the human genome, DNA sequences of human genes, which were closely linked genetically or physically, were surveyed from the GenBank Data Bank. A total of 72 sequences longer than 2 kb, which were mutually linked within 500 kb, were identified. The sequences belonged to 17 linkage groups and were ordered in each group according to their genetic positions. Analyses of the G+C% distribution along the ordered sequences showed that sequences within each group almost always had similar G+C% levels, but those belonging to different groups often had different levels. Similar analyses of more distantly linked sequences (e.g., greater than 10 Mb) showed mosaic structures of G+C% distribution. These findings are consistent with predictions made from the "isochore" structures found by CsCl equilibrium centrifugation, in that the structures having homogeneous base compositions stretched over at least several hundred kilobases. A possible boundary of the giant G+C% mosaic structures was identified between X-linked G6PD and F8C.

Base Composition↗

Long terminal repeat retrotransposons of Oryza sativa.

BACKGROUND: Long terminal repeat (LTR) retrotransposons constitute a major fraction of the genomes of higher plants. For example, retrotransposons comprise more than 50% of the maize genome and more than 90% of the wheat genome. LTR retrotransposons are believed to have contributed significantly to the evolution of genome structure and function. The genome sequencing of selected experimental and agriculturally important species is providing an unprecedented opportunity to view the patterns of variation existing among the entire complement of retrotransposons in complete genomes. RESULTS: Using a new data-mining program, LTR_STRUC, (LTR retrotransposon structure program), we have mined the GenBank rice (Oryza sativa) database as well as the more extensive (259 Mb) Monsanto rice dataset for LTR retrotransposons. Almost two-thirds (37) of the 59 families identified consist of copia-like elements, but gypsy-like elements outnumber copia-like elements by a ratio of approximately 2:1. At least 17% of the rice genome consists of LTR retrotransposons. In addition to the ubiquitous gypsy- and copia-like classes of LTR retrotransposons, the rice genome contains at least two novel families of unusually small, non-coding (non-autonomous) LTR retrotransposons. CONCLUSIONS: Each of the major clades of rice LTR retrotransposons is more closely related to elements present in other species than to the other clades of rice elements, suggesting that horizontal transfer may have occurred over the evolutionary history of rice LTR retrotransposons. Like LTR retrotransposons in other species with relatively small genomes, many rice LTR retrotransposons are relatively young, indicating a high rate of turnover.

Animals↗

Substantial non-homologous recombination and structural variation results from Brassica AABC and CCAB hybrid meiosis.

Meiotic crossovers contribute to genetic diversity and play a crucial role in homologous chromosome segregation. Non-homologous crossovers in Brassica, involving the exchange of genetic material between genomes, can be valuable for transferring novel traits or characteristics between Brassica species. However, there are a limited number of studies that specifically investigate crossover frequencies in populations of interspecific hybrids. We investigated the distribution and frequency of homologous crossover events, as well as non-homologous recombination and structural variation, in hybrids between B. juncea (AABB)&#x2009;&#xd7;&#x2009;B. napus (AACC) (resulting in AABC hybrids; 5 genotypes) and B. napus (AACC)&#x2009;&#xd7;&#x2009;B. carinata (BBCC) (resulting in CCAB hybrids; 4 genotypes). The analysis was performed on individuals derived from microspore culture of both unreduced and reduced gametes produced by the AABC and CCAB hybrids. All AABC and almost all CCAB unreduced gamete-derived individuals and most AABC and CCAB reduced gamete-derived individuals showed copy number variation indicative of non-homologous (A-C) recombination. Additionally, a higher frequency of homologous crossovers, also in centromeric and pericentromic regions, was observed in the diploid genomes of the AABC and CCAB hybrids. Overall, these hybrid types show high frequencies of A-C introgressions, which may be useful in B. juncea or B. carinata introgression breeding, and this increased recombination frequency may help break up existing linkage disequilibrium blocks in the Brassica A and C genomes.

Meiosis↗

Genome size variation in North American minnows (Cyprinidae). II. Variation among 20 species.

Genome sizes (nuclear DNA contents) from 200 individuals representing 20 species of North American cyprinid fishes (minnows) were examined spectrophotometrically. The distributions of DNA values of individuals within populations of the 20 species were essentially continuous and normal; the distribution of DNA values among species was continuous and overlapping. These observations suggest that changes in DNA quantity in cyprinids are small in amount, involve both gains and losses of DNA, and are cumulative and independent in effect. Significant heterogeneity in mean genome size occurs both between individuals within populations of species and among species. The former averages maximally around 6% of the cyprinid genome and is nearly the same as the amount of DNA theoretically needed for the entire cyprinid structural gene component. The majority of the DNA content variation among the 20 species is distributed above the level of individuals within populations. Comparisons of average genome size difference or distance between individuals drawn from different levels of taxonomic organization indicate that considerably greater divergence in genome size has occurred in the extremely speciose cyprinid genus Notropis as compared with other North American cyprinid genera. This may suggest that genome size change is concentrated in speciation episodes. Finally, no associations were found between interspecific variation in genome size and five life-history characters. This suggests that much of the variation in genome size within and among the 20 species may be phenotypically inconsequential.

Animals↗

Sequence conservation and antigenic variation of the structural proteins of equine rhinitis A virus.

The nucleotide and deduced amino acid sequences of the P1 region of the genomes of 10 independent equine rhinitis A virus (ERAV) isolates were determined and found to be very closely related. A panel of seven monoclonal antibodies to the prototype virus ERAV.393/76 that bound to nonneutralization epitopes conserved among all 10 isolates was raised. In serum neutralization assays, rabbit polyclonal sera and sera from naturally and experimentally infected horses reacted in a consistent and discriminating manner with the 10 isolates, which indicated the existence of variation in the neutralization epitopes of these viruses.

Amino Acid Sequence↗

Gene structure and promoter variation of expressed and nonexpressed variants of the KIR2DL5 gene.

Two variants of the novel KIR2DL5 gene (KIR2DL5.1 and.2) were identified in genomic DNA of a single donor. However, only the KIR2DL5.1 variant was transcribed in PBMC. In this study, analysis of seven additional donors reveals two new variants of the KIR2DL5 gene and indicates that transcription, or its lack, are consistently associated with particular variants of this gene. Comparison of the complete nucleotide sequences of the exons and introns of KIR2DL5.1 and KIR2DL5.2 reveals no structural abnormalities, but similar open reading frames for both variants. In contrast, the promoter region of KIR2DL5 shows a high degree of sequence polymorphism that is likely relevant for expression. Substitution within a putative binding site for the transcription factor acute myeloid leukemia gene 1 could determine the lack of expression for some KIR2DL5 variants.

Base Sequence↗