PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Structure of linkage disequilibrium and phenotypic associations in the maize genome.

Association studies based on linkage disequilibrium (LD) can provide high resolution for identifying genes that may contribute to phenotypic variation. We report patterns of local and genome-wide LD in 102 maize inbred lines representing much of the worldwide genetic diversity used in maize breeding, and address its implications for association studies in maize. In a survey of six genes, we found that intragenic LD generally declined rapidly with distance (r(2) < 0.1 within 1500 bp), but rates of decline were highly variable among genes. This rapid decline probably reflects large effective population sizes in maize during its evolution and high levels of recombination within genes. A set of 47 simple sequence repeat (SSR) loci showed stronger evidence of genome-wide LD than did single-nucleotide polymorphisms (SNPs) in candidate genes. LD was greatly reduced but not eliminated by grouping lines into three empirically determined subpopulations. SSR data also supplied evidence that divergent artificial selection on flowering time may have played a role in generating population structure. Provided the effects of population structure are effectively controlled, this research suggests that association studies show great promise for identifying the genetic basis of important traits in maize with very high resolution.

Chromosome Mapping↗

Histones of genetically active and inactive chromatin.

It has frequently been proposed that a variation in the relative content of lysine-rich, moderately lysine-rich, and arginine-rich histones might provide a mechanism by which specific portions of the genome may be genetically regulated. This possibility was investigated by comparing the electrophoretic pattern of these three fractions in cells differing markedly in their content of genetically active and genetically inactive chromatin. Three models were used: heterochromatin versus euchromatin; metaphase cells versus interphase cells, and mature lymphocytes versus phytohemagglutinin-stimulated lymphocytes. In no case was there a significant difference in the histone patterns of these contrasting models. It is concluded that, although histones may act as a generalized repressor and structural component of chromatin, factors other than a variation in histone pattern may be responsible for repression or derepression of specific segments of the genome.

Animals↗

The pattern of polymorphism in Arabidopsis thaliana.

We resequenced 876 short fragments in a sample of 96 individuals of Arabidopsis thaliana that included stock center accessions as well as a hierarchical sample from natural populations. Although A. thaliana is a selfing weed, the pattern of polymorphism in general agrees with what is expected for a widely distributed, sexually reproducing species. Linkage disequilibrium decays rapidly, within 50 kb. Variation is shared worldwide, although population structure and isolation by distance are evident. The data fail to fit standard neutral models in several ways. There is a genome-wide excess of rare alleles, at least partially due to selection. There is too much variation between genomic regions in the level of polymorphism. The local level of polymorphism is negatively correlated with gene density and positively correlated with segmental duplications. Because the data do not fit theoretical null distributions, attempts to infer natural selection from polymorphism data will require genome-wide surveys of polymorphism in order to identify anomalous regions. Despite this, our data support the utility of A. thaliana as a model for evolutionary functional genomics.

Arabidopsis↗

Chromosome-scale assembly with improved annotation provides insights into breed-wide genomic structure and diversity in domestic cats.

INTRODUCTION: Comprehensive genomic resources offer insights into biological features, including traits/disease-related genetic loci. The current reference genome assembly for the domestic cat (Felis catus), Felis_Catus_9.0 (felCat9), derived from sequences of the Abyssinian cat, may inadequately represent the general cat population, limiting the extent of deducible genetic variations. OBJECTIVES: The goal was to develop Anicom American Shorthair 1.0 (AnAms1.0), a reference-grade chromosome-scale cat genome assembly. METHODS: In contrast to prior assemblies relying on Abyssinian cat sequences, AnAms1.0 was constructed from the sequences of more popular American Shorthair breed, which is related to more breeds than the Abyssinian cat. By combining advanced genomics technologies, including PacBio long-read sequencing and Hi-C- and optical mapping data-based sequence scaffolding, we compared AnAms1.0 to existing Felidae genome assemblies (20 scaffolds, scaffolds N50&#xa0;>&#xa0;150 Mbp). Homology-based and ab initio gene annotation through Iso-Seq and RNA-Seq was used to identify new coding genes and splice variants. RESULTS: AnAms1.0 demonstrated superior contiguity and accuracy than existing Felidae genome assemblies. Using AnAms1.0, we identified over 1.5 thousand structural variants and 29 million repetitions compared to felCat9. Additionally, we identified > 1,600 novel protein-coding genes. Notably, olfactory receptor structural variants and cardiomyopathy-related variants were identified. CONCLUSION: AnAms1.0 facilitates the discovery of novel genes related to normal and disease phenotypes in domestic cats. The analyzed data are publicly accessible on Cats-I (https://cat.annotation.jp/), which we established as a platform for accumulating and sharing genomic resources to discover novel genetic traits and advance veterinary medicine.

Animals↗

Genetic structuring and estimation of reproductive adults in Onchocerca volvulus: A genome-wide analysis across hosts and regions.

Genomic analysis of parasites can deepen our understanding of their transmission, population structure, and important biological characteristics. Onchocerciasis (river blindness), caused by the parasitic nematode Onchocerca volvulus, involves adult worms residing in subcutaneous nodules that produce larval-stage microfilariae (mf), which are routinely detected in the skin for diagnosis. Whole-genome studies of mf are limited; most analyses have focused on the mitochondrial genome. We conducted a genome-wide analysis with 94% median nuclear genome coverage, analyzing 171, 37, and 98 mf from 16, 3, and 5 individuals from Ghana, Liberia, and the Democratic Republic of Congo, respectively. These data were used to investigate population differentiation, estimate the number of reproductive adult worms, and analyze genetic variation across chromosomes. Population genetic analyses across hosts and countries showed that nuclear genome diversity can reveal fine-scale genetic structure, even between geographically close countries, providing more resolution than mitochondrial haplotype data. By reconstructing maternal and paternal sibships, we estimated the number of reproductively active adult filariae. Comparisons between adult worm estimates from genetic data and nodule observations showed that genetics-based estimates were higher or equal to observed worm counts in 8 out of 9 hosts for female worms and 7 out of 9 hosts for male worms. Our analysis also revealed lower-than-expected X chromosome diversity, consistent with neo-X chromosome fusions in filarial species. This study represents an important step in using nuclear genome data from mf to support onchocerciasis elimination efforts and in developing genetic tools that could inform mass drug administration programs.

Onchocerca volvulus↗

Whole-genome analysis: annotations and updates.

The most important advances in the field of genome annotation over the past two years involve the use of cDNA sequences, protein structures and gene expression data to predict genes. These types of information not only improve gene identification, but they also give insights into variation in gene structure and function.

Alternative Splicing↗

Molecular systematics of Anopheles: from subgenera to subpopulations.

The century-old discovery of the role of Anopheles in human malaria transmission precipitated intense study of this genus at the alpha taxonomy level, but until recently little attention was focused on the systematics of this group. The application of molecular approaches to systematic problems ranging from subgeneric relationships to relationships at and below the species level is helping to address questions such as anopheline phylogenetics and biogeography, the nature of species boundaries, and the forces that have structured genetic variation within species. Current knowledge in these areas is reviewed, with an emphasis on the Anopheles gambiae model. The recent publication of the genome of this anopheline mosquito will have a profound impact on inquiries at all taxonomic levels, supplying better tools for estimating phylogeny and population structure in the short term, and ultimately allowing the identification of genes and/or regulatory networks underlying ecological differentiation, speciation, and vectorial capacity.

Animals↗

Isochore structures in the mouse genome.

The distribution of the G+C content in the mouse genome has been studied using a windowless technique. We have found that: (i). Abrupt variations of the G+C content from a GC-rich region to a GC-poor region, and vice versa, occur frequently at some sites along the sequence of the mouse genome. (ii). Long domains with relatively homogeneous G+C content (isochores) exist, which usually have sharp boundaries. Consequently, 28 isochores longer than 1 Mb have been identified in the mouse genome. A homogeneity index was used to quantify the variations of the G+C content within isochores. The precise boundaries, sizes, and G+C contents of these isochores have been determined. The windowless technique for the G+C content computation was also used to analyze the DNA sequence containing the mouse MHC region, which has a GC-poor isochore. This isochore is located at the central part of the sequence with boundaries at 468459 and 812716 bp, where the sequence is extended from the centromeric end to the telomeric end. In addition, the analysis of a segment of the rat genome shows that the rat genome also has clear isochore structures.

Algorithms↗

Genomic and Structural Analysis of Gamete Recognition Proteins in a Broadcast Spawning Echinoderm Mesocentrotus franciscanus.

Gamete recognition proteins are expressed on the surfaces of sperm and eggs, where they mediate interactions between gametes. The genetic basis for gamete recognition proteins, as well as their structure and interactions, have yet to be fully resolved. Using a new high-quality de novo genome assembly for the sea urchin Mesocentrotus franciscanus, we investigated the genomic structure, expression, and protein forms of several gamete recognition proteins: sperm bindin, egg receptor for sperm (HSP110), and egg bindin receptor (EBR1), as well as the receptor for egg jelly (REJ) and its paralogs. To inform future population genetic and evolutionary studies, we resolve the genomic structure of the large EBR1 protein, identifying fewer tandem CUB-TSP1 repeats in EBR1 compared to the initial characterization of this protein. As expected for an egg receptor for sperm, EBR1 is highly expressed in female reproductive tissues (eggs and female gonad), compared to other tissues. In contrast, HSP110 shows similar levels of expression across male and female reproductive tissues, as well as across non-reproductive tissues and development stages. HSP110 might be a pleiotropic gene that in part influences fertilization. Using protein structural modeling and functional domain predictions, we propose hypotheses about potential interactions among EBR1, bindin, and HSP110 proteins that may provide insight into sperm-egg interactions in sea urchins. Resolving the genomic structure of genes encoding gamete recognition proteins, in combination with functional annotations and protein structural modeling, enables deeper investigation into the consequences of variation in gamete recognition proteins and the evolution of reproductive isolation.

Mesocentrotus franciscanus↗

Evolution of the P-type II ATPase gene family in the fungi and presence of structural genomic changes among isolates of Glomus intraradices.

BACKGROUND: The P-type II ATPase gene family encodes proteins with an important role in adaptation of the cell to variation in external K+, Ca2+ and Na2+ concentrations. The presence of P-type II gene subfamilies that are specific for certain kingdoms has been reported but was sometimes contradicted by discovery of previously unknown homologous sequences in newly sequenced genomes. Members of this gene family have been sampled in all of the fungal phyla except the arbuscular mycorrhizal fungi (AMF; phylum Glomeromycota), which are known to play a key-role in terrestrial ecosystems and to be genetically highly variable within populations. Here we used highly degenerate primers on AMF genomic DNA to increase the sampling of fungal P-Type II ATPases and to test previous predictions about their evolution. In parallel, homologous sequences of the P-type II ATPases have been used to determine the nature and amount of polymorphism that is present at these loci among isolates of Glomus intraradices harvested from the same field. RESULTS: In this study, four P-type II ATPase sub-families have been isolated from three AMF species. We show that, contrary to previous predictions, P-type IIC ATPases are present in all basal fungal taxa. Additionally, P-Type IIE ATPases should no longer be considered as exclusive to the Ascomycota and the Basidiomycota, since we also demonstrate their presence in the Zygomycota. Finally, a comparison of homologous sequences encoding P-type IID ATPases showed unexpectedly that indel mutations among coding regions, as well as specific gene duplications occur among AMF individuals within the same field. CONCLUSION: On the basis of these results we suggest that the diversification of P-Type IIC and E ATPases followed the diversification of the extant fungal phyla with independent events of gene gains and losses. Consistent with recent findings on the human genome, but at a much smaller geographic scale, we provided evidence that structural genomic changes, such as exonic indel mutations and gene duplications are less rare than previously thought and that these also occur within fungal populations.

Base Sequence↗

[Epigenetic and synergistic types of inheritance of the reproductive characters in angiosperms].

The author considers three types of hereditary memory (structural, cell and signal), that are realized on different levels of biological organization. These three types of hereditary memory correspond to three types of reproduction: self-replication, cell division and reproduction s. str. Reproductive characters are exemplified with three essential characters in angiosperm plants: dimorphism in population by flower sex; mono-, di- and trystyly of flowers; uni- and biparental mode of seed reproduction. All these characters are considered as "supercharacters" that are controlled by gene ensembles. The correspondence between three types of reproduction and three types of hereditary memory are discussed. The authors reviews also the role of polyploidy (auto- and endoploidy) in the inheritance of reproductive condition. From the information theory point the increase in cell ploidy causes the growth of uncertainty in expression of genes and gene ensembles thus creating new type of variability--epigenetic variation. The change of reproductive strategy in plants is regulated by state of gene and gene ensembles and does not demand structural changes in genome. The reproductive characters of plants in spite its complex structure are inherited in number of generations as a discrete Mendel characters by mono-, di-, ot trihybrid schemas.

Biological Evolution↗

Exploring alternative transcript structure in the human genome using blocks and InterPro.

Understanding how alternative splicing affects gene function is an important challenge facing modern-day molecular biology. Using homology-based, protein sequence analysis methods, it should be possible to investigate how transcript diversity impacts protein function. To test this, high-quality exon-intron structures were deduced for over 8000 human genes, including over 1300 (17 percent) that produce multiple transcript variants. A data mining technique (DiffMotif) was developed to identify genes in which transcript variation coincides with changes in conserved motifs between variants. Applying this method, we found that 30 percent of the multi-variant genes in our test set exhibited a differential profile of conserved InterPro and/or BLOCKS motifs across different mRNA variants. To investigate these, a visualization tool (ProtAnnot) that displays amino acid motifs in the context of genomic sequence was developed. Using this tool, genes revealed by the DiffMotif method were analyzed, and when possible, hypotheses regarding the potential role of alternative transcript structure in modulating gene function were developed. Examples of these, including: MEOX1, a homeobox-containing protein; AIRE, involved in auto-immune disease; PLAT, tissue type plasminogen activator; and CD79b, a component of the B-cell receptor complex, are presented. These results demonstrate that amino acid motif databases like BLOCKS and InterPro are useful tools for investigating how alternative transcript structure affects gene function.

Algorithms↗

Characterization of iap gene in Listeria monocytogenes strains isolated in Japan.

Variation of the iap gene region (407bp) encoding an invasion-associated protein p60 was studied on 12 strains of Listeria monocytogenes of different origin in Japan. These 12 strains are known to have 2 types of serotype (1/2a and 4b) and have a diversity among the strains (Saito et al., 1998). The dye-primer cycle sequencing method was employed to determine the genomic structure, and the nucleotide sequences obtained were compared with those of reference strain SV 1/2a EGD. Differences found in the nucleotides were as follows; point mutations of 33 variations in 32 places; an insertion and 3 deletions of 3 bases; AAT position (po.) 1282-1283, and GCA po. 1307-1309, ACA po. 1412-1414, AAT po. 1439-1444, respectively. Different repeating numbers by 6 base unit, ACA AAT, were also found in the tandem repeat region (po. 1394-1423). Classification of 12 strains was attempted, then 8, 4 and 5 types were obtained from the point mutations, the insertions and deletions, and the repeating numbers, respectively. Consequently, 8 patterns were profiled regardless of each serotype. From these results, genomic structures were partially clarified in the iap gene 407bp of L. monocytogenes isolated in Japan. Then, the possibility of detailed epidemiology for L. monocytogenes infection using a combination of serotype and genome structure was suggested because of the previous polymorphism thought to be due to the nucleotide differences in the region.

Bacterial Proteins↗

A genomic perspective on the chromodomain-containing retrotransposons: Chromoviruses.

Chromoviruses, chromodomain-containing retrotransposons, are the only Metaviridae (Ty3/gypsy group of retrotransposons) clade with a Eukaryota-wide distribution. They have a common evolutionary origin and are the most prolific and diverse Metaviridae clade. The fusion of a retrotransposon and a chromodomain, was most probably responsible for their extreme evolutionary success in Eukaryota. Analysis of the massive amount of genome sequence data for different eukaryotic lineages has provided an in depth insight into the diversity, evolution, neofunctionalization, high rate of genomic turnover and origin of chromoviruses in Eukaryota. This review attempts to summarise the unique aspects of chromoviruses from a genomic perspective.

Amino Acid Sequence↗

The out of Africa model of varicella-zoster virus evolution: single nucleotide polymorphisms and private alleles distinguish Asian clades from European/North American clades.

Until 1998, varicella-zoster virus (VZV) was generally considered sufficiently stable to allow the use of a single sequenced virus (VZV-Dumas) as a consensual representation of the world VZV genotype. But recent investigations have uncovered a gE mutant virus called VZV-MSP with a second genotype and a distinguishable accelerated cell spread phenotype. A subsequent study suggested that single nucleotide polymorphisms (SNPs) could be applied toward the genetic analysis of the VZV genome. To further assess the scope of genetic variation in the VZV genome on a worldwide basis, we carried out an extensive SNP analysis of structural glycoprotein genes gB, gE, gH, gI, gL, as well as the IE62 regulatory gene in viruses collected from Western Europe, North America and Asia, including the VZV vaccine strain. The SNP data showed segregation of viral isolates of Asian origin from those of Western ancestry into distinct phylogenetic clades. Unexpectedly, however, VZV from Thailand segregated with VZV from Iceland and the United States, i.e. it was more Western than Asian in nature. Further, SNP analysis disclosed strikingly unusual genotypes, e.g. gH genes with up to five missense mutations and gL genes with insertions of an in-frame methionine codon. In summary, these VZV genomic analyses have shown that individual VZV strains, like closely related human beings, have distinctive SNP profiles containing private alleles within just five VZV genes (gB, gH, gE, gL and IE62) that provide a fingerprint to localize ancestry of the viral strain.

Africa↗

An integrated view of protein evolution.

Why do proteins evolve at different rates? Advances in systems biology and genomics have facilitated a move from studying individual proteins to characterizing global cellular factors. Systematic surveys indicate that protein evolution is not determined exclusively by selection on protein structure and function, but is also affected by the genomic position of the encoding genes, their expression patterns, their position in biological networks and possibly their robustness to mistranslation. Recent work has allowed insights into the relative importance of these factors. We discuss the status of a much-needed coherent view that integrates studies on protein evolution with biochemistry and functional and structural genomics.

Animals↗

Computational analysis suggests that alternative first exons are involved in tissue-specific transcription in rice (Oryza sativa).

MOTIVATION: Transcription start site selection and alternative splicing greatly contribute to diversifying gene expression. Recent studies have revealed the existence of alternative first exons, but most have involved mammalian genes, and as yet the regulation of usage of alternative first exons has not been clarified, especially in plants. RESULTS: We systematically identified putative alternative first exon transcripts in rice, verified the candidates using RT-PCR, and searched for the promoter elements that might regulate the alternative first exons. As a result, we detected a number of unreported alternative first exons, some of which are regulated in a tissue-specific manner. SUPPLEMENTARY INFORMATION: http://www.bioinfo.sfc.keio.ac.jp/research/intron.

Alternative Splicing↗

Microsatellite variation in North American populations of Drosophila melanogaster.

Computer database searching for microsatellites can be particularly effective for organisms like Drosophila melanogaster for which there are extensive sequence data. Here we demonstrate that 17 out of 18 such microsatellites are also highly polymorphic in natural populations of Drosophila, and that this variation is easily scorable with PCR followed by electrophoresis on high-resolution agarose. This form of variation is likely to be of great value in studies of the genomic distribution of polymorphism, population structure, the relation between intraspecific polymorphism and interspecific divergence and the mutation rate and pattern of mutations of microsatellites. In this preliminary survey of 15 lines, we find that the variance in repeat count is most strongly correlated with the maximum count, that perfect repeats are significantly more variable than imperfect repeats and that repeats which are split by an imperfection have unexpectedly low variance given the size of the perfectly repeated portion.

Alleles↗