PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Simple repeated sequences”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

On evolutionarily conserved simple repetitive DNA sequences: do "sex-specific" satellite components serve any sequence dependent function?

The nuclear genomes of eukaryotes contain DNA of varying degrees of repetition. Highly repetitious DNA and simple repetitive sequences as a fraction thereof appear to be distributed in a non-random fashion in the genome. There are arguments for and against functional roles of simple repetitive sequences, and the reasons for their evolutionary conservation are not at all clear. In order to learn more about the biologic role of simple repetitive sequences in the context of their evolutionary history, we report here the following results from studies of sex-specific snake satellite DNA: 1) The snake simple repeat sequence is 5'-GATAGACA-3' and it is strictly conserved throughout vertebrate evolution. 2) The simple repeat sequence is intimately interspersed with single-copy DNA throughout the mouse genome. 3) The simple repeat is transcribed into RNA in several animal systems and it is translatable in bacterial test systems. 4) The simple repeat sequence is sex-specifically arranged in vertebrates. 5) In snake DNA, the simple repeat is adjacent to a single-copy sequence which singles out a male-specific putative mRNA in mouse polysomal poly (A)+ RNA. Thus even if this snake simple repetitive sequence is not involved in a basic cellular function such as sex-determination, it is nevertheless a valuable tool to approach those problems.

Animals↗

Instability of simple sequence DNA in Saccharomyces cerevisiae.

All eukaryotic genomes thus far examined contain simple sequence repeats. A particularly common simple sequence in many organisms (including humans) consists of tracts of alternating GT residues on one strand. Allelic poly(GT) tracts are often of different lengths in different individuals, indicating that they are likely to be unstable. We examined the instability of poly(GT) and poly(G) tracts in the yeast Saccharomyces cerevisiae. We found that these tracts were dramatically unstable, altering length at a minimal rate of 10(-4) events per division. Most of the changes involved one or two repeat unit additions or deletions, although one alteration involved an interaction with the yeast telomeres.

Base Sequence↗

Frameshift mutations in TGFbetaRII, IGFIIR, BAX, hMSH3 and hMSH6 are absent in lung cancers.

A genome-wide instability at simple repeat sequences characterizes gastrointestinal and endometrial cancers of the microsatellite mutator phenotype (MMP). The genes encoding transforming growth factor-beta receptor type II (TGFbetaRII), insulin-like growth factor II receptor (IGFIIR), Bcl-2 associated X protein (BAX), hMSH3 and hMSH6 have simple repeat sequences in their coding regions. Consequently, mutations in the single repeat sequences in these genes provide one major route for carcinogenesis in these cancers. We examined 43 non-small cell lung carcinomas and 16 small cell carcinomas for frameshift mutations in simple repeat sequences of TGFbetaRII, IGFIIR, BAX, hMSH3 and hMSH6. In addition, MMP was assessed using a primer set for BAT-26. None of 59 lung cancers exhibited frameshift mutations or MMP. It is concluded that somatic frameshift mutations in these genes and MMP do not constitute important mechanisms in lung carcinogenesis. The possibility of some sort of genetic instability undetectable as a form of MMP cannot be precluded.

Base Sequence↗

DNA polymorphism in cytokine genes based on length variation in simple-sequence tandem repeats.

The possibility of the involvement of cytokines in the genetic predisposition to various diseases has been suggested by a large variety of studies. However, the study of potential disease linkage of cytokine genes has been hampered by a lack of sufficiently polymorphic markers at the restriction fragment length polymorphism (RFLP) level. We have investigated the distribution, the length polymorphism, the informativeness, and the efficiency of analysis, of simple-sequence tandem repeats in the mouse cytokine genes. Highly polymorphic sequences have been identified in the IL-1 beta, IL-1ra, IL-2, IL-4, IL-5, IL-6, IL-7, and IFN-gamma genes. The utility and the value of these sequences as gene markers is exemplified by mapping the IL-7 gene to mouse chromosome 3 close to pgk-1ps3 and Car-2 loci and the IFN-gamma gene to chromosome 10 near the pg locus. Advantages of short tandemly repeated sequences as genetic markers are discussed in comparison with RFLPs.

Animals↗

Fragile X syndrome unstable element, p(CCG)n, and other simple tandem repeat sequences are binding sites for specific nuclear proteins.

The trinucleotide repeat sequences which become unstable in fragile X syndrome and myotonic dystrophy are located in the untranslated regions of their respective genes, FMR1 and DM1. This implies that a functional constraint other than coding capacity maintains the presence of the repeats. In the case of fragile X syndrome, sequences adjacent to the repeat are methylated in affected individuals and the FMR1 gene is transcriptionally inactive. We demonstrate that the fragile X p(CCG)n repeat itself is methylated in vivo and that methylation of this repeat is able to inhibit in vitro binding of a novel, specific nuclear p(CCG)n binding protein (CCG-BP1)--one of at least 10 distinct simple tandem repeat sequence binding proteins (STR-BPs). We describe additional, apparently distinct, binding activities both for the methylated form of the p(CCG)n repeat and for each of the single strands of the repeat.

Base Sequence↗

Molecular structure of the chicken vitamin D-induced calbindin-D28K gene reveals eleven exons, six Ca2+-binding domains, and numerous promoter regulatory elements.

The seco-steroid hormone 1,25-dihydroxyvitamin D3 is known to induce the expression of a calcium binding protein termed calbindin-D28K in a variety of target tissues. In order to comprehend the mechanism of induction we have cloned and sequenced the chicken calbindin-D28K gene. The gene spans some 18.5 kilobases (kb) of chromosomal DNA from the putative Cap site to the polyadenylation site of the 2.8 kb mRNA. It is split into 11 coding exons by 10 intervening sequences. The promoter region of this gene is markedly G + C-rich (60-80%) extending from -225 to +400. Within this region we find 70 CpG dinucleotides, four G-C boxes, and numerous known promoter regulatory signals. These putative regulatory signals include a TATA box (ATAAATA) at -30 and a CAT box (CCAAT) at -326. Ten additional variant CAT boxes are found in the upstream promoter region (-218 to -770) of this gene. Furthermore we have identified a glucocorticoid-like responsive element at -410 (TCTACACACTGTTCC) and this element overlaps a metal responsive element (TGCACTC) and a variant CAT box (CCAAAT) and juxtaposes an enhancer-like core element (AAATGGT) on its 3'-side. In addition, the calbindin-D28K promoter is composed of a variety of simple repeated sequences, some of which are components of putative regulatory signals. All splice junctions were found to conform to the GT-AG rule. A consensus sequence of the 5'-splice junction reads AG/GTAAG-TTATA. A consensus sequence of the 3'-splice site consists of two elements: a pyrimidine track (mainly T) followed by ACAG/G-T. A two-dimensional model of calbindin-D28K was constructed which projects the existence of 6 alpha-helix-loop-alpha-helix regions characteristic of calcium binding domains. The 3'-end of the gene consists of a single large (2039 base pair) uninterrupted exon, an organizational feature common to other members of the calcium binding protein gene family which include calmodulin, parvalbumin, Spec I, myosin light chains, etc. Another feature common to the gene family is the presence of the repeated sequence ATTT or TTTA located in the 3'-untranslated exons. These simple repeat sequences could be involved in regulating mRNA degradation by serving as a ribonuclease recognition signal.

Amino Acid Sequence↗

The sequence of the gorilla fetal globin genes: evidence for multiple gene conversions in human evolution.

Two fetal globin genes (G gamma and A gamma) from one chromosome of a lowland gorilla (Gorilla gorilla gorilla) have been sequenced and compared to three human loci (a G gamma-gene and two A gamma-alleles). A comparison of regions of local homology among these five sequences indicates that long after the duplication that produced the two nonallelic gamma-globin loci of catarrhine primates, about 35 million years (Myr) ago, at least one gene conversion event occurred between these loci. This conversion occurred not long before the ancestral divergence (about 6 Myr ago) of Homo and Gorilla. After this ancestral divergence, a minimum of three more gene conversion events occurred in the human lineage. Each human A gamma-allele shares specific sequence features with the gorilla A gamma-gene; one such distinctive allelic feature involves the simple repeated sequence in IVS 2. This suggests that early in the human lineage the A gamma-genes may have undergone a crossing-over event mediated by this simple repeated sequence. The DNA sequences from coding regions of both G gamma- and A gamma-loci, a comparison of 292 codons in the corresponding gorilla and human genes, show an unusually low evolutionary rate, with only two nonsilent differences and, surprisingly, not even one silent substitution. The two nonsynonymous substitutions observed predict a glycine at codon 73 and an arginine at codon 104 in the gorilla A gamma-sequence rather than aspartic acid and lysine, respectively, in human A gamma. Because only arginine has been found at position 104 in gamma-chains of Old World monkeys, it may represent the ancestral residue lost in gorilla and human G gamma-chains and in the human A gamma-chain. Possibly the arginine codon (AGG) was replaced by the lysine codon (AAG) in the G gamma-gene of a common ancestor of Homo and Gorilla and then was transferred to the A gamma-gene by subsequent conversions in the human lineage. DNA sequence conversions, similar to that attributed to the fetal gamma-globin genes, appear to be relatively frequent phenomena and, if widespread throughout the genome, may have profound evolutionary consequences.

Alleles↗

Identification of novel simple sequence length polymorphisms (SSLPs) in mouse by interspersed repetitive element (IRE)-PCR.

Interspersed repetitive element (IRE)-PCR is a useful method for identification of novel human or mouse sequence tagged sites (STSs) from contigs of genomic clones. We describe the use of IRE-PCR with mouse B1 repetitive element primers to generate novel, PCR amplifiable, simple sequence length polymorphisms (SSLPs) from yeast artificial chromosome (YAC) clones containing regions of mouse chromosomes 13 and 14. Forty-two IRE-PCR products were cloned and sequenced from eight YACs. Of these, 29 clones contained multiple simple sequence repeat units. PCR analysis with primers derived from unique sequences flanking the simple sequence repeat units in seven clones showed all to be polymorphic between various mouse strains. This novel approach to SSLP identification represents an efficient method for saturating a genomic interval with polymorphic genetic markers that may expedite the positional cloning of genes for traits and diseases.

Animals↗

Mapping of the interleukin 5 receptor gene to human chromosome 3 p25-p26 and to mouse chromosome 6 close to the Raf-1 locus with polymorphic tandem repeat sequences.

Simple-sequence tandem repeat sequences in the 3' UTR of interleukin 5 (IL5)-receptor gene of human and mouse are polymorphic in their length among humans and different strains of mice. In 20 different human Epstein-Barr virus (EBV)-transformed cell lines, six alleles of IL5R could be distinguished. In the mouse, three different alleles are found. With the human-specific IL5R tandem repeat marker in human-rodent somatic cell hybrids, the IL5R gene was mapped to human Chromosome (Chr) 3 p25-p26. With the mouse-specific IL5R tandem repeat sequence in recombinant inbred strains of mice, the Il5r gene was mapped to the distal part of mouse Chr 6 close to the Raf-1 locus.

Animals↗

Genomic cloning and chromosomal assignment of rat regucalcin gene.

The gene for a Ca(2+)-binding protein regucalcin was cloned from a rat genomic library which was constructed in lambda FIX II by screening with radiolabeled probe (complementary DNA of rat liver regucalcin). Positive clone had 19.9 kb insert of size and contained four exons of the gene coding for a rat regucalcin. These exons included the partial coding sequence (61.2% of open reading frame) and the entire 3'-untranslated region of the gene. The nucleotide sequence of exons completely agreed with that of a rat regucalcin cDNA clone. The sequence analysis of the clone showed that the identifier sequence and two simple repeated sequences exist in the intron of the gene. Moreover, chromosomal location of the rat regucalcin gene was determined by direct R-banding fluorescence in situ hybridization (FISH) method with the 19.9 kb clone containing four exons. The regucalcin gene was localized on rat chromosome Xq11.1-12 proximal end.

Amino Acid Sequence↗

Identification and determination of the relationships of species and strains within the genus Leishmania using single primers in the polymerase chain reaction.

DNA polymorphisms were assessed in different species and strains within the genus Leishmania by amplifying genomic DNA with single non-specific primers. This polymerase chain reaction (PCR) method employed non-random primers which anneal to mini- and microsatellite DNA sequences like the M13 core sequence and the simple repeat sequences (GTG)5 and (GACA)4, and the T3B primer derived from an intergenic spacer for tRNA genes. Distinctive and reproducible sets of amplified DNA fragments were obtained for all Leishmania isolates tested. The number and size of amplification products were found to be characteristic for a given taxon. Highly similar PCR profiles were observed when genomic DNA of representatives of the L. donovani, L. mexicana or L. braziliensis complexes was amplified. By comparing PCR patterns of unidentified Leishmania isolates with those obtained from reference strains it was possible to identify these isolates at the species level. The information of the amplification patterns was used for the construction of phylogenetic trees to measure the genetic relatedness within the genus Leishmania.

Animals↗

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial↗

Repeat-associated phase variable genes in the complete genome sequence of Neisseria meningitidis strain MC58.

Phase variation, mediated through variation in the length of simple sequence repeats, is recognized as an important mechanism for controlling the expression of factors involved in bacterial virulence. Phase variation is associated with most of the currently recognized virulence determinants of Neisseria meningitidis. Based upon the complete genome sequence of the N. meningitidis serogroup B strain MC58, we have identified tracts of potentially unstable simple sequence repeats and their potential functional significance determined on the basis of sequence context. Of the 65 potentially phase variable genes identified, only 13 were previously recognized. Comparison with the sequences from the other two pathogenic Neisseria sequencing projects shows differences in the length of the repeats in 36 of the 65 genes identified, including 25 of those not previously known to be phase variable. Six genes that did not have differences in the length of the repeat instead had polymorphisms such that the gene would not be expected to be phase variable in at least one of the other strains. A further 12 candidates did not have homologues in either of the other two genome sequences. The large proportion of these genes that are associated with frameshifts and with differences in repeat length between the neisserial genome sequences is further corroborative evidence that they are phase variable. The number of potentially phase variable genes is substantially greater than for any other species studied to date, and would allow N. meningitidis to generate a very large repertoire of phenotypes through expression of these genes in different combinations. Novel phase variable candidates identified in the strain MC58 genome sequence include a spectrum of genes encoding glycosyltransferases, toxin related products, and metabolic activities as well as several restriction/modification and bacteriocin-related genes and a number of open reading frames (ORFs) for which the function is currently unknown. This suggests that the potential role of phase variation in mediating bacterium-host interactions is much greater than has been appreciated to date. Analysis of the distribution of homopolymeric tract lengths indicates that this species has sequence-specific mutational biases that favour the instability of sequences associated with phase variation.

Bacterial Proteins↗

Comparative analysis of seeded and vegetative biotype buffalograsses based on phylogenetic relationship using ISSRs, SSRs, RAPDs, and SRAPs.

Buffalograss [ Buchloe dactyloides (Nutt.) Englem.] is the only native grass that is being used extensively as a turfgrass in the Great Plains region. Its low-growth habit, drought resistance, and low-maintenance requirement make it attractive as a turfgrass species. Our objective was to obtain an overview on the genetic relatedness among and within seeded and vegetative biotype buffalograsses using inter-simple sequence repeats (ISSRs), random amplified polymorphic DNA (RAPDs), sequence-related amplified polymorphisms (SRAPs), and simple sequence repeats (SSRs) markers that were derived from related species (maize, pearl millet, sorghum, and sugarcane). Twenty individuals per cultivar were genotyped using 30 markers from each marker system. All buffalograss cultivars were uniquely fingerprinted by all four marker systems. Mean genetic similarities were estimated at 0.52, 0.51, 0.62, and 0.57 using SSRs, ISSRs, SRAPs, and RAPDs, respectively. Two main clusters separating the seeded-biotype from the vegetative-biotype cultivars were produced using UPGMA analysis. Further subgroupings were unequivocal. The Mantel test resulted in a very good fit (SRAP=0.92, ISSR=0.90) to good fit (RAPD=0.86, SSR=0.88) of cophenetic values. Comparing the four marker systems to each other, RAPD and SRAP similarity indices were highly correlated ( r=0.73), while Spearman's rank correlation coefficient between RAPDs and SSRs was r=0.24 and between ISSRs and SSRs was r=0.66. A genotype-assignment analytical approach might be useful for cultivar identification and property rights protection. Polymorphic SRAPs were abundant and demonstrated genetic diversity among closely related cultivars.

Analysis of Variance↗

Evidence on DNA slippage step-length distribution.

A simple model based on a master equation is constructed in order to reveal the details of the mutational events modifying simple sequence repeats in the human genome, A database of simple repeats together with their flanking sequences comprising approximately 10(5) entries from all 24 human chromosomes was constructed. By aligning the pairs of fragments of sequences containing the repeat elements, the matrices that count the number of slippage events were evaluated. The counts were then used as a target to be reproduced by our theoretical model, in which the elongation and shortening of the repeats proceed through a mechanism in which the step lengths exhibit a decaying distribution in the form of an inverse power law rather than through one nucleotide extension or deletion, which was the most frequent supposition in previous studies.

Chromosomes, Human↗

(TG)n uncovers a sex-specific hybridization pattern in cattle.

Screening of a bovine genomic library with the human minisatellite 33.6 probe uncovered a family of clones that, when used to probe Southern blots of bovine genomic DNA digested with the restriction enzyme HaeIII or MboI, revealed sexually dimorphic, but otherwise virtually monomorphic, patterns among the larger DNA fragments to which they hybridized. Characterization of one of these clones revealed that it contains different minisatellite sequences. The sexual dimorphism hybridization pattern observed with this clone was found to be due to multiple copies of two tandemly interspersed repeats: the simple sequence (TG)n and a previously undescribed 29-bp sequence. Both repeats appear to share many genomic loci including autosomal loci. In contrast, Southern analysis of AluI- or HinfI-digested bovine DNA with the (TG)n repeat used as a probe yielded substantial polymorphism. These results show that (i) different minisatellites can be found in a cluster, (ii) both simple and more complex repeated sequences other than the simple quaternary (GATA)n repeat can be sexually dimorphic, and (iii) simple repeats can reveal substantial polymorphism.

Animals↗

Human von Willebrand factor gene and pseudogene: structural analysis and differentiation by polymerase chain reaction.

Structural analysis of the von Willebrand factor gene located on chromosome 12 is complicated by the presence of a partial unprocessed pseudogene on chromosome 22q11-13. The structures of the von Willebrand factor pseudogene and corresponding segment of the gene were determined, and methods were developed for the rapid differentiation of von Willebrand factor gene and pseudogene sequences. The pseudogene is 21-29 kilobases in length and corresponds to 12 exons (exons 23-34) of the von Willebrand factor gene. Approximately 21 kilobases of the gene and pseudogene were sequenced, including the 5' boundary of the pseudogene. The 3' boundary of the pseudogene lies within an 8-kb region corresponding to intron 34 of the gene. The presence of splice site and nonsense mutations suggests that the pseudogene cannot yield functional transcripts. The pseudogene has diverged approximately 3.1% in nucleotide sequence from the gene. This suggests a recent evolutionary origin approximately 19-29 million years ago, near the time of divergence of humans and apes from monkeys. Several repetitive sequences were identified, including 4 Alu, one Line-1, and several short simple sequence repeats. Several of these simple repeats differ in length between the gene and pseudogene and provide useful markers for distinguishing these loci. Sequence differences between the gene and pseudogene were exploited to design oligonucleotide primers for use in the polymerase chain reaction to selectivity amplify sequences corresponding to exons 23-34 from either the von Willebrand factor gene or the pseudogene. This method is useful for the analysis of gene defects in patients with von Willebrand disease, without interference from homologous sequences in the pseudogene.

Amino Acid Sequence↗

A genetic linkage map for tef [Eragrostis tef (Zucc.) Trotter].

Tef [Eragrostis tef (Zucc.) Trotter] is the major cereal crop in Ethiopia. Tef is an allotetraploid with a base chromosome number of 10 (2n = 4x = 40) and a genome size of 730 Mbp. Ninety-four F(9) recombinant inbred lines (RIL) derived from the interspecific cross, Eragrostis tef cv. Kaye Murri x Eragrostis pilosa (accession 30-5), were mapped using restriction fragment length polymorphisms (RFLP), simple sequence repeats derived from expressed sequence tags (EST-SSR), single nucleotide polymorphism/insertion and deletion (SNP/INDEL), intron fragment length polymorphism (IFLP) and inter-simple sequence repeat amplification (ISSR). A total of 156 loci from 121 markers was grouped into 21 linkage groups at LOD 4, and the map covered 2,081.5 cM with a mean density of 12.3 cM per locus. Three putative homoeologous groups were identified based on multi-locus markers. Sixteen percent of the loci deviated from normal segregation with a predominance of E. tef alleles, and a majority of the distorted loci were clustered on three linkage groups. This map will be useful for further genetic studies in tef including mapping of loci controlling quantitative traits (QTL), and comparative analysis with other cereal crops.

Chromosome Mapping↗