PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “genomic tandem array”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Extensive amplification and transposition of a novel repetitive element, xstir, together with its terminal inverted repeat in the evolution of Xenopus.

A DNA fragment containing short tandem repeat sequences (approximately 86-bp repeat) was isolated from a Xenopus laevis cDNA library. Southern blot and in situ hybridization analyses revealed that the repeat was highly dispersed in the genome and was present at approximately 1 million copies per haploid genome. We named this element Xstir (Xenopus short tandemly and invertedly repeating element) after its arrangement in the genome. The majority of the genomic Xstir sequences were digested to monomer and dimer sizes with several restriction enzymes. Their sequences were found to be highly homogeneous and organized into tandem arrays in the genome. Alignment analyses of several known sequences showed that some of the Xstir-like sequences were also organized into interspersed inverted repeats. The inverted repeats consisted of an inverted pair of two differently modified Xstirs separated by a short insert. In addition, these were framed by another novel inverted repeat (Xstir-TIR). The Xstir-TIR sequence was also found at the ends of tandem Xstir arrays. Furthermore, we found that Xstir-TIR was linked to a motif characterizing the T2 family which belonged to a vertebrate MITE (miniature inverted-repeat transposable element) family, suggesting the importance of Xstir-TIR for their amplification and transposition. The present study of 11 anuran and 2 urodele species revealed that Xstir or Xstir-like sequences were extensively amplified in the three Xenopus species. Genomic Xstir populations of X. borealis and X. laevis were mutually indistinguishable but significantly different from that of X. tropicalis.

Animals↗

Molecular epidemiology of African swine fever virus studied by analysis of four variable genome regions.

Variable regions of the African swine fever virus genome, which contain arrays of tandem repeats, were compared in the genomes of isolates obtained over a 40-year period. Comparison of the size of products generated by polymerase chain reaction (PCR) from four different genome regions, within the B602L and KP86R genes and intergenic regions J286L and BtSj, placed 43 closely related isolated from Europe, the Caribbean, West and Central Africa into 17 different virus sub-groups. Sequence analysis of the most variable fragment, within the B602L gene, from 81 different isolates distinguished 31 sub-groups of virus isolates which varied in sequence and number of a tandem repeat encoding 4 amino acids. Thus, each of these analysis methods enabled isolates, which were previously grouped together by sequencing of a more conserved genome region, to be separated into multiple sub-groups. This provided additional information about strains of viruses circulating in different countries. The methods could be used in future to study the epidemiology and evolution of virus isolates and to trace the sources of disease outbreaks.

Africa↗

Organization and complete nucleotide sequence of the core-histone-gene cluster of the annelid Platynereis dumerilii.

The arrangement of the core-histone genes, their transcriptional polarity and their nucleotide sequences have been determined for the polychaete annelid Platynereis dumerili. A clone containing the core-histone genes was isolated from a annelid genomic library constructed in the EMBL-4 phage vector, using a trout H3 genomic probe. This clone was found to contain two and a half repeats of a 6-kbp EcoRV fragment that contained one copy of each of the core-histone genes. The clusters are tandemly arrayed in the genome and the gene order within the core-histone cluster does not vary. Absolutely no differences were found in the nucleotide sequences comprising the same part of two adjacent clusters (bases -225 to 2776 and bases 5821 to 8825). The number of copies of the cluster appeared to be high: approximately 660 copies/diploid cell, as also observed in sea urchins and amphibians. There are also some additional subtypes of histone gene organization: multimers of tandemly arrayed genes and isolated genes; these are present at a much lower copy number (an average of 40-50 copies/diploid genome). Two mRNAs (for H2B and H3) are transcribed from one DNA strand and the two other histone mRNAs (for H2A and H4) from the other strand as is the case for some insects and certain vertebrates. No H1-coding sequence has been found in the completely sequenced four-membered cluster. The organization of histone genes in P. dumerilii is similar to the clustering found in Caenorhabditis elegans but in this nematode worm several different types of organization are observed with a low copy number for each.

Amino Acid Sequence↗

AniAnn's: alignment-free annotation of tandem repeat arrays using fast average nucleotide identity estimates.

MOTIVATION: Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of tandem repeat arrays, are limited in their ability to annotate divergent or novel sequences. RESULTS: In this work, we introduce AniAnn's, an algorithm for annotating large blocks of tandemly repeating DNAs. AniAnn's exploits the high Average Nucleotide Identity (ANI) shared between repeat units of the same array to quickly and accurately infer the boundaries of such arrays. We show that AniAnn's improves the annotation of satellites and other tandem repeats within a variety of plant and animal genomes, while requiring only a fraction of the runtime compared to previous approaches. We conclude by exploring several use cases of AniAnn's as a lightweight method for masking repeats prior to whole-genome alignment as well as the de novo annotation and classification of satellite repeats. AVAILABILITY: AniAnn's is open source software and available at github.com/marbl/anianns.

Algorithms↗

Structure and expression of the woodchuck herpesvirus genome.

This report describes the genome structure and location from which immediate-early transcription originates in the recently characterized woodchuck herpesvirus (herpesvirus marmota: HVM). Cross-hybridization of restriction fragments indicates that the HVM genome contains a tandem array of 1.5-kb repeat units. Additionally, terminal labeling and exonuclease experiments demonstrate that the repeated sequences lie at the termini of the genome. Hybridization of probes representing immediate-early transcription indicates that only a single predominant species of immediate-early RNA originates from a region near one end of unique sequences in the HVM genome. These results show remarkable similarity with group 2 of the gammaherpesvirinae. However, no homology was detected by conventional Southern blot hybridization between HVM and the gamma-2 prototype, herpesvirus saimiri. Therefore, we propose HVM to be a new member of the gammaherpesvirinae subfamily of herpesviruses.

Animals↗

New methods for detection of low levels of DNA damage in human populations.

The use of a postlabeling method to characterize and to detect infrequent base modifications in DNA is outlined. This method has the advantage that low levels of DNA modifications, approximately 1 modified base per 10(5) nucleotides, can be detected. Moreover, a broad spectrum of modification can be identified by using this methodology. The basis for the method involves transfer of a radioactive phosphate from the gamma position of ATP to the 5'-hydroxyl terminus of 3'-phosphoryl nucleotides that are derived from modified DNA by appropriate nuclease digestion. The second method involves use of a defined DNA sequence within human cells. The alpha sequence is used as a probe for DNA damage to specific nucleotides. The alpha DNA sequence is reiterated approximately 300,000 times in the human genome and exists in tandem arrays. It comprises approximately 1% of the entire genome. The reiterated sequence is sufficiently homogeneous to permit its use as a probe for a site specific in DNA damage. Examples of the application of both of these methodologies to DNA damage inflicted in human cells by chemicals and ultraviolet light are provided.

Alkylating Agents↗

A gene family encoding heterogeneous histone H1 proteins in Trypanosoma cruzi.

A gene family encoding a set of histone H1 proteins in Trypanosoma cruzi is described. The sequence of 3 genomic and 4 cDNA clones revealed the presence of several motifs characteristic of histone H1, although heterogeneity at the polypeptide level was evident. The clones encode histone H1 proteins of an unusually small size (74-97 amino acids), which lack the globular domain found in histone H1 of higher eukaryotes. All histone H1 mRNAs from T. cruzi are polyadenylated, although no typical polyadenylation signal was found. Furthermore, the genes encoding the histone H1 proteins in T. cruzi are found in a tandem array containing 15-20 gene copies per haploid genome. This tandem array is located on a large chromosome of 2.2 Mb.

Amino Acid Sequence↗

The SRS superfamily of Toxoplasma surface proteins.

The surface of the protozoan parasite Toxoplasma gondii is coated with developmentally expressed, glycosylphosphatidylinositol-linked proteins structurally related to the highly immunogenic surface antigen SAG1. Collectively, these surface antigens are known as the SRS (SAG1-related sequences) superfamily of proteins. SRS proteins are thought to mediate attachment to host cells and activate host immunity to regulate the parasite's virulence. To better understand the number, evolution and developmental expression of SRS genes, this study has bioinformatically identified 161 unique SRS DNA sequences present in the T. gondii type II Me49 genome. The SRS superfamily of sequences phylogenetically bifurcates into two subfamilies, the prototypic members being SAG1 and SAG2A, respectively. Paralogous SRS sequences are 24-99% identical, are tandemly arrayed throughout the genome, and are present on most, if not all, chromosomes. All 11 SRS sequences on chromosomes Ia and Ib are clustered at sub-telomeric expression sites. Messenger RNA expression in the majority of SRS sequences for which multiple Expressed Sequence Tags exist is developmentally regulated. A consensus nucleotide sequence surrounding both the splice acceptor and donor sites was identified in those SRS sequences possessing an intron. Genotypic differences among SRS sequences are present at several loci (e.g. the absence of SAG5B, the truncation of SAG2D in Me49 compared with RH) indicating that different genotypes possess distinct sets of SRS sequences. Orthologous genes are restricted to tissue-dwelling coccidia (Neospora, Sarcocystis) with no related sequences present in other more distant apicomplexa such as Eimeria, Cryptosporidia, and Plasmodium spp.

Animals↗

Genomic analysis of sequence variation in tandemly repeated DNA. Evidence for localized homogeneous sequence domains within arrays of alpha-satellite DNA.

As a model to examine the local distribution of sequence variation within large arrays of tandemly repeated DNA in complex genomes, the long-range organization of alpha-satellite DNA from human chromosome 17 was investigated. Three individual chromosomes, representing different alpha-satellite haplotypes, were segregated into mouse and human somatic cell hybrids and their arrays sized by pulse-field gel electrophoresis. An inventory of the higher-order repeat units found in multiple separate regions of these megabase arrays was obtained using cosmid mapping and two-dimensional gel electrophoresis, a technique that combines the large-scale resolution of pulsed-field gel electrophoresis with the small-scale resolution of conventional gel electrophoresis. These analyses show that alpha-satellite arrays are characterized by the presence of localized homogeneous domains containing only one distinct type of repeat unit. These domains, which consist of sequence variants and/or higher-order repeat length variants, can be up to at least several hundred thousands of bases in length. Both abundant and rare variant repeat units can be localized in these distinct domains, which may correspond to transition states in the evolution of tandem multicopy DNA families. This description of the organization of large arrays of tandem repeats provides insight into mechanisms involved in their homogenization.

Amino Acid Sequence↗

Variability and evolution of highly repeated DNA sequences in the genus Beta.

Satellite DNA from wild beet species was separated from restriction endonuclease digested genomic DNA by polyacrylamide gel electrophoresis. Two nonhomologous HaeIII satellite DNA repeats were cloned from the wild beet Beta trigyna. The type I repeat is 140-149 bp long and AT rich, while the type II is 162 bp in size and GC rich. A third repetitive HaeIII element cloned from the related wild beet B. corolliflora was shown to be organized as a HinfI satellite DNA family in the cultivated beet B. vulgaris ssp. vulgaris and the wild beet B. vulgaris ssp. maritima. This type III satellite monomer is 149 bp long and contains a high number of short direct subrepeats. The monomer was found in different genomic organizations and copy numbers in all sections of the genus Beta indicating an amplification early in the phylogeny. The HaeIII repeats from B. trigyna are characterized by a lower variability and form long tandem arrays in the genomes of Corollinae species. The investigation of the distribution of all three sequence families provided data that may contribute to the solution of taxonomic problems of the genus Beta and be useful in the characterization of hybrids and derived lines with alien wild beet chromosomes.

Base Sequence↗

Evolution of tubulin gene arrays in Trypanosomatid parasites: genomic restructuring in Leishmania.

BACKGROUND: alpha- and beta-tubulin are fundamental components of the eukaryotic cytoskeleton and cell division machinery. While overall tubulin expression is carefully controlled, most eukaryotes express multiple tubulin genes in specific regulatory or developmental contexts. The genomes of the human parasites Trypanosoma brucei and Leishmania major reveal that these unicellular kinetoplastids possess arrays of tandem-duplicated tubulin genes, but with differences in organisation. While L. major possesses monotypic alpha and beta arrays in trans, an array of alternating alpha- and beta tubulin genes occurs in T. brucei. Polycistronic transcription in these organisms makes the chromosomal arrangement of tubulin genes important with respect to gene expression. RESULTS: We investigated the genomic architecture of tubulin tandem arrays among these parasites, establishing which character state is derived, and the timing of character transition. Tubulin loci in T. brucei and L. major were compared to examine the relationship between the two character states. Intergenic regions between tubulin genes were sequenced from several trypanosomatids and related, non-parasitic bodonids to identify the ancestral state. Evidence of alternating arrays was found among non-parasitic kinetoplastids and all Trypanosoma spp.; monotypic arrays were confirmed in all Leishmania spp. and close relatives. CONCLUSION: Alternating and monotypic tubulin arrays were found to be mutually exclusive through comparison of genome sequences. The presence of alternating gene arrays in non-parasitic kinetoplastids confirmed that separate, monotypic arrays are the derived state and evolved through genomic restructuring in the lineage leading to Leishmania. This fundamental reorganisation accounted for the dissimilar genomic architectures of T. brucei and L. major tubulin repertoires.

Animals↗

Introducing a null mutation in the mouse K6alpha and K6beta genes reveals their essential structural role in the oral mucosa.

Mammalian genomes feature multiple genes encoding highly related keratin 6 (K6) isoforms. These type II keratins show a complex regulation with constitutive and inducible components in several stratified epithelia, including the oral mucosa and skin. Two functional genes, K6alpha and K6beta, exist in a head-to-tail tandem array in mouse genomes. We inactivated these two genes simultaneously via targeting and homologous recombination. K6 null mice are viable and initially indistinguishable from their littermates. Starting at two to three days after birth, they show a growth delay associated with reduced milk intake and the presence of white plaques in the posterior region of dorsal tongue and upper palate. These regions are subjected to greater mechanical stress during suckling. Morphological analyses implicate the filiform papillae as being particularly sensitive to trauma in K6alpha/K6beta null mice, and establish the complete absence of keratin filaments in their anterior compartment. All null mice die about a week after birth. These studies demonstrate an essential structural role for K6 isoforms in the oral mucosa, and implicate filiform papillae as being the major stress bearing structures in dorsal tongue epithelium.

Animals↗

Characterization and chromosomal distribution of novel satellite DNA sequences of the lesser rhea (Pterocnemia pennata) and the greater rhea (Rhea americana).

Two different types of novel satellite DNA (stDNA) sequences were cloned from the lesser rhea (Ptercnemia pennata) and the greater rhea (Rhea americana) after digestion of genomic DNAs with a restriction endonuclease Pvu II, and characterized by filter hybridization and in-situ hybridization to metaphase chromosomes. These nucleotide sequences consisted of GC-rich 288-bp and 332-bp repeated elements in P. pennata and 288-bp and 336-bp repeated elements in R. americana, all of which were organized in tandem arrays in the genome. The 288-bp and 332-bp elements of P. pennata displayed strong sequence similarity with the 288-bp and 336-bp elements of R. americana, respectively. The 332-bp and 336-bp elements were located on almost all the microchromosomes in both the species. The other type of repeated elements, the 288-bp element, was located on four and nine pairs of microchromosomes in P. pennata and R. americana, respectively. All the stDNA sequences were not crosshybridized to genomic DNAs of another three ratite species, ostrich (Struthio camelus), cassowary (Casuarius casuarius) and emu (Dromaius novaehollandiae), suggesting that these stDNA sequences are conserved in the same family but fairly divergent among the different families of Struthioniformes.

Animals↗

The in vivo conformation of the plastid DNA of Toxoplasma gondii: implications for replication.

The Phylum Apicomplexa comprises thousands of obligate intracellular parasites, some of which cause serious disease in man and other animals. Though not photosynthetic, some of them, including the malaria parasites (Plasmodium spp.) and the causative organism of Toxoplasmosis, Toxoplasma gondii, possess a remnant plastid partially determined by a highly derived residual genome encoded in 35 kb DNA. The genetic maps of the plastid genomes of these two organisms are extremely similar in nucleotide sequence, gene function and gene order. However, a study using pulsed field gel electrophoresis and electron microscopy has shown that in contrast to the malarial version, only a minority of the plastid DNA of Toxoplasma occurs as circular 35 kb molecules. The majority consists of a precise oligomeric series of linear tandem arrays of the genome, each oligomer terminating at the same site in the genetic map, i.e. in the centre of a large inverted repeat (IR) which encodes duplicated tRNA and rRNA genes. This overall topology strongly suggests that replication occurs by a rolling circle mechanism initiating at the centre of the IR, which is also the site at which the linear tails of the rolling circles are processed to yield the oligomers. A model is proposed which accounts for the quantitative structure of the molecular population. It is relevant that a somewhat similar structure has been reported for at least three land plant chloroplast genomes.

Animals↗

Loss of polyoma virus infectivity as a result of a single amino acid change in a region of polyoma virus large T-antigen which has extensive amino acid homology with simian virus 40 large T-antigen.

The polyoma virus (Py) transformed cell line 7axB, selected by in vivo passage of an in vitro transformed cell, contains an integrated tandem array of 2.4 genomes and produces the large, middle, and small Py T-antigen species, with molecular weights of 100,000, 55,000, and 22,000, respectively (Hayday et al., J. Virol. 44:67-77, 1982; Lania et al., Cold Spring Harbor Symp. Quant. Biol. 44:597-603, 1980). The integrated viral and adjacent host DNA sequences have been molecularly cloned as three EcoRI fragments (Hayday et al.). One of these fragments (7B-M), derived from within the tandem viral sequences, is equivalent to an EcoRI viral linear molecule. Fragment 7B-M has been found to be transformation competent but incapable of producing infectious virus after DNA transfection (Hayday et al.). By constructing chimerae between 7B-M and Py DNA and by direct DNA sequencing, the mutation responsible for the loss of infectivity has been located to a single base change (adenine to guanine) at nucleotide 2503. This results in a conversion of an aspartic acid to a glycine in the C-terminal region of the Py large T-antigen but does not appear to affect the binding of the Py large T-antigen to Py DNA at the putative DNA replication and autoregulation binding sites. The mutation is located within a 21-amino acid homology region shared by the simian virus 40 large T-antigen (Friedmann et al., Cell 17:715-724, 1979). These results suggest that the mutation in the 7axB large T-antigen may be involved in the active site of the protein for DNA replication.

Amino Acid Sequence↗

Complete sequence of Euglena gracilis chloroplast DNA.

We report the complete DNA sequence of the Euglena gracilis, Pringsheim strain Z chloroplast genome. This circular DNA is 143,170 bp, counting only one copy of a 54 bp tandem repeat sequence that is present in variable copy number within a single culture. The overall organization of the genome involves a tandem array of three complete and one partial ribosomal RNA operons, and a large single copy region. There are genes for the 16S, 5S, and 23S rRNAs of the 70S chloroplast ribosomes, 27 different tRNA species, 21 ribosomal proteins plus the gene for elongation factor EF-Tu, three RNA polymerase subunits, and 27 known photosynthesis-related polypeptides. Several putative genes of unknown function have also been identified, including five within large introns, and five with amino acid sequence similarity to genes in other organisms. This genome contains at least 149 introns. There are 72 individual group II introns, 46 individual group III introns, 10 group II introns and 18 group III introns that are components of twintrons (introns-within-introns), and three additional introns suspected to be twintrons composed of multiple group II and/or group III introns, but not yet characterized. At least 54,804 bp, or 38.3% of the total DNA content is represented by introns.

Animals↗

Comparison of the Z and W sex chromosomal architectures in elegant crested tinamou (Eudromia elegans) and ostrich (Struthio camelus) and the process of sex chromosome differentiation in palaeognathous birds.

To clarify the process of avian sex chromosome differentiation in palaeognathous birds, we performed molecular and cytogenetic characterization of W chromosome-specific repetitive DNA sequences for elegant crested tinamou (Eudromia elegans, Tinamiformes) and constructed comparative cytogenetic maps of the Z and W chromosomes with nine chicken Z-linked gene homologues for E. elegans and ostrich (Struthio camelus, Struthioniformes). A novel family of W-specific repetitive sequences isolated from E. elegans was found to be composed of guanine- and cytosine-rich 293-bp elements that were tandemly arrayed in the genome as satellite DNA. No nucleotide sequence homologies were found for the Struthioniformes and neognathous birds. The comparative cytogenetic maps of the Z and W chromosomes of E. elegans and S. camelus revealed that there are partial deletions in the proximal regions of the W chromosomes in the two species, and the W chromosome is more differentiated in E. elegans than in S. camelus. These results suggest that a deletion firstly occurred in the proximal region close to the centromere of the acrocentric proto-W chromosome and advanced toward the distal region. In E. elegans, the W-specific repeated sequence elements were amplified site-specifically after deletion of a large part of the W chromosome occurred.

Animals↗

Molecular characterization of the histone gene family of Caenorhabditis elegans.

The core histone genes (H2A, H2B, H3 and H4) of Caenorhabditis elegans are arranged in approximately 11 dispersed clusters and are not tandemly arrayed in the genome. Three well-characterized genomic clones, which contain histone genes, have one copy of each core histone gene per cluster. One of the clones (lambda Ceh-1) carries one histone cluster surrounded by several thousand base-pairs of non-histone DNA, and another clone (lambda Ceh-3) contains a histone cluster duplication surrounded by non-histone DNA. A third clone (lambda Ceh-2) carries a cluster of core histone genes flanked on one side (12,000 base-pairs away) by a single H2B gene and on the other by non-histone DNA. A fourth cluster (clone BE9) has one copy each of H3 and H4 and two copies each of H2A and H2B. This cluster is also flanked by non-histone DNA. Analysis of cosmid clones which overlap three of the clusters shows that no other histone clusters are closer than 8000 to 60,000 base-pairs, although unidentified non-histone transcription units are present on the flanking regions. Gene order within the histone clusters varies, and histone mRNAs are transcribed from both DNA strands. No H1 sequences are found on these core histone clones. Restriction fragment length polymorphisms between two related nematode strains (Bristol and Bergerac) were used as phenotypic markers in genetic crosses to map one histone cluster to linkage group V and another to linkage group IV. Hybridization of gene-specific probes from sea urchin to C. elegans RNA identifies C. elegans core histone messenger RNAs of sizes similar to sea urchin early stage histone mRNAs (H2A, H2B, H3 and H4). The organization of histone genes in C. elegans resembles the clustering found in most vertebrate organisms and does not resemble the tandem patterns of the early stage histone gene family of sea urchins or the major histone locus of Drosophila.

Animals↗