PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Simple repeated sequences”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Comparative maps of human 19p13.3 and mouse chromosome 10 allow identification of sequences at evolutionary breakpoints.

A cosmid/bacterial artificial chromosome (BAC) contiguous (contig) map of human chromosome (HSA) 19p13.3 has been constructed, and over 50 genes have been localized to the contig. Genes and anonymous ESTs from approximately 4000 kb of human 19p13.3 were placed on the central mouse chromosome 10 map by genetic mapping and pulsed-field gel electrophoresis (PFGE) analysis. A region of approximately 2500 kb of HSA 19p13.3 is collinear to mouse chromosome (MMU) 10. In contrast, the adjacent approximately 1200 kb are inverted. Two genes are located in a 50-kb region after the inversion on MMU 10, followed by a region of homology to mouse chromosome 17. The synteny breakpoint and one of the inversion breakpoints has been localized to sequenced regions in human <5 kb in size. Both breakpoints are rich in simple tandem repeats, including (TCTG)n, (CT)n, and (GTCTCT)n, suggesting that simple repeat sequences may be involved in chromosome breaks during evolution. The overall size of the region in mouse is smaller, although no large regions are missing. Comparing the physical maps to the genetic maps showed that in contrast to the higher-than-average rate of genetic recombination in gene-rich telomeric region on HSA 19p13.3, the average rate of recombination is lower than expected in the homologous mouse region. This might indicate that a hot spot of recombination may have been lost in mouse or gained in human during evolution, or that the position of sequences along the chromosome (telomeric compared to the middle of a chromosome) is important for recombination rates.

Animals↗

On simple repetitive DNA sequences and complex diseases.

Simple repetitive DNA sequences are abundantly interspersed in eukaryote genomes and therefore useful in genome research and genetic fingerprinting in plants, fungi and animals, including man. Recently, simple repeats were also identified in some prokaryotic genomes. Hence the same probes can be applied for multilocus DNA fingerprinting in medically relevant bacteria. Simple repeats including composite dinucleotide microsatellites are differentially represented in different compartments of eukaryote genomes. Expanded triplet blocks in and around certain genes may, for example, cause so-called trinucleotide diseases in man. As a consequence, simple repetitive sequences should also be characterized with respect to their influences on the DNA structure, gene expression, genomic (in)stability and their development on an evolutionary time scale. Here three examples of microsatellites in the human major histocompatibility complex (HLA) are investigated, a (GT)n microsatellite situated 2 kb 5' off the lymphotoxin alpha (LTA) gene, a (GAA)n block in the 5' part of the HLA-F gene and a composite (GT)n(GA)m stretch in the second intron of HLA-DRBl genes. Grossly differing mutation rates are evident in these elements as well as varying linkage disequilibria. The unfolding of these simple repeats in distant human populations is covered including Caucasians, Bushmen and South American Indians. Furthermore, implications of simple repeat neighboring genes are discussed for the multifactorial diseases multiple sclerosis (MS), rheumatoid arthritis (RA) and early onset pauciarticular arthritis (EOPA). Polymorphisms of HLA-DRBl and T cell receptor beta variable (TCRBV) genes confer susceptibility for these autoimmune diseases as demonstrable by intronic simple repeat variability. Microsatellite polymorphisms within the TNF region reveal linkage disequilibria with HLA-DRBl and different promotor alleles of the TNFA gene. Disease associations with TNFA microsatellite alleles are, on the one hand, secondary to associations with HLA-DRBl genes (in MS) or they represent additional risk factors (in RA, EOPA) on the other hand. Evolutionary persistence, various structural conformations and the specific binding of nuclear proteins to several simple repeat sequences refute the preconceptions of biological insignificance for all of these ubiquitously interspersed elements.

Arthritis↗

The abundance of various polymorphic microsatellite motifs differs between plants and vertebrates.

The abundance of different simple sequence motifs in plants was accessed through data base searches of DNA sequences and quantitative hybridization with synthetic dinucleotide repeats. Database searches indicated that microsatellites are five times less abundant in the genomes of plants than in mammals. The most common plant repeat motif was AA/TT followed by AT/TA and CT/GA. This group comprised about 75% of all microsatellites with a length of more than 6 repeats. The GT/CA motif being the most abundant dinucleotide repeat in mammals was found to be considerably less frequent in plants. To address the question if plant simple repeat sequences are variable as in mammals, (GT)n and (CT)n microsatellites were isolated from B.napus. Five loci were investigated by PCR-analysis and amplified products were obtained for all microsatellites from B. oleracea, B.napus and B.rapa DNA, but only for one primer pair from B.nigra. Polymorphism was detected for all microsatellites.

Animals↗

Modification of breast cancer risk in young women by a polymorphic sequence in the egfr gene.

The regulation of the epidermal growth factor receptor (egfr) gene in human cancer is not yet fully understood. Recent data on a polymorphic CA repeat located at the 5'-regulatory sequence in intron 1 of the egfr gene [egfr CA simple sequence repeat (SSR) I] point to a possible inheritance of cancer risk associated with the egfr gene. Furthermore, we have detected frequent allelic imbalances restricted to the egfr CA SSR I in breast cancer tissue and nontumorous breast tissue adjacent to invasive and in situ breast cancer representing amplifications. Therefore, we conducted a population-based case-control study to assess the relationship between the egfr polymorphism and breast cancer risk. Cases with a first primary breast cancer by age 50 years and age-matched population controls provided information on known and suspected risk factors. The allelic length of the egfr CA SSR was determined in 616 cases and 1072 population-sampled controls. Genotypes were categorized for analysis by allele length. Multivariate logistic regression was used to compare genotype distributions, accounting for other risk factors, and to investigate gene-environment interactions. We found a modifying effect, albeit no main effect, of the allelic length of the egfr polymorphism on breast cancer risk. The presence of two long alleles (>/==" BORDER="0">19 CA) was associated with a significantly elevated odds ratio (OR) of 10.4 [95% confidence interval (CI), 1.85-58.70] among women with a first-degree family history of breast cancer (P = 0.015 for interaction). The risk increase associated with high red meat consumption (OR, 10.68; 95% CI, 1.57-72.58) and the protective effect of high vegetable intake (OR, 0.07; 95% CI, 0.004-1.07) was also most pronounced among carriers of two long alleles (>/==" BORDER="0">19 CA). The length of the egfr CA SSR may increase the risk for familial breast cancers, and its effect could be modulated by dietary factors.

Adult↗

Identification and characterization of microsatellites in Norway spruce (Picea abies K.).

Norway spruce (Picea abies) genomic libraries were screened for presence of dinucleotide AC/GT and AG/CT microsatellites (or simple sequence repeats). On average, one (AG)n microsatellite every 194 kb and one (AC)n microsatellite every 406 kb were found. Forty-six positive clones were sequenced and primers flanking 24 AG microsatellites and 12 AC microsatellites diesigned. Only seven (20%) of them produced the expected single-locus polymorphic pattern when used to amplify Norway spruce DNAs. The other primer pairs gave either multiple bands or bad amplification, or a single monomorphic fragment. Such a small proportion of successful primer pairs was attributed to the high level of complexity of the Norway spruce genome. Dot blot analysis of the clones showed that many of them contained repetitive DNA and that those giving the single-locus polymorphic patterns usually corresponded to single-copy sequences. A family of repetitive DNA that contained AG repeats was identified and was present in about 40,000 copies per haploid genome. Simple Mendelian inheritance was observed for all the polymorphisms tested. The average number of alleles was 13, ranging from 6 to 22, and the expected heterozygosity was 0.79 when seven microsatellites were used to genotype a panel of 18 trees representing different populations. Compared with isozymes, microsatellites are about five times more informative and could provide an extremely valuable source of markers for genome mapping and genetic diversity studies.

Alleles↗

The region of tobacco mosaic virus RNA involved in the nucleation of assembly.

The interaction of TMV RNA with the disk aggregate of TMV protein at the initiation of assembly has been studied by using the techniques of RNA sequencing. The 5' end group has been identified, and shown not to be protected in the early stages of assembly from accessibility to nuclease digestion. A population of RNA fragments of average length 250 nucleotides, originating from a unique region of TMV RNA, is encapsidated by limited assembly, and sufficient sequence information is available to identify certain unusual features. The protected region does not contain highly reiterated simple repeating sequences, but may contain more complicated repeats. The length and complexity of the nucleation region may reflect adaptation to the efficient mediation of the conformational change from disk to helix of TMV protein, besides a requirement for binding to the disk, and this may be an important part of the mechanism of specificity in the nucleation of assembly.

Electrophoresis↗

Organization and complete sequence of identical embryonic and plasmacytoma kappa V-region genes.

We have determined the entire sequence of two immunoglobulin kappa variable region genes derived separately from embryonic and plasmacytoma cells. Both genes are in the germline configuration; both are organized into hydrophobic leader and mature V-region coding sequences separated by a short intervening sequence. Both are identical with one another throughout the 1360 bases determined. In addition, there is a simple repeat sequence of 31 CA doublets that occurs about 300 bases from the 3' side of the coding sequence in each gene.

Amino Acid Sequence↗

The Friedreich ataxia region: characterization of two novel genes and reduction of the critical region to 300 kb.

Friedreich ataxia is a severe neurodegenerative autosomal recessive disorder of unknown biochemical defect. The Friedreich ataxia locus (FRDA) is tightly linked to the centromeric side of the D9S5 locus. We have used 'exon-trapping' to identify two new genes, approximately 100 and 200 kb centromeric to D9S5, respectively. One gene appears ubiquitously expressed while the other is prominently expressed in muscle. The ubiquitous transcript codes for a protein containing a 20 aa repeat reminiscent of simple repeats found in several ribonucleoproteins. Using the single-strand conformation polymorphism (SSCP) procedure, we searched for mutations in affected patients in the coding sequence of the two genes, as well as in a gene that we had previously identified in the same region. Eight polymorphic DNA changes but no causative mutations were found, suggesting that the genes are not candidates for Friedreich ataxia. The discovery of a simple sequence repeat polymorphism in the most centromeric gene allowed the localization within that gene of the breakpoint of a previously described recombination in a Friedreich ataxia family, therefore excluding the two distal genes from the FRDA region. The lack of causative mutations in the three genes and the position of the recombination further delineate the FRDA locus to a 300 kb interval.

Amino Acid Sequence↗

[Identification and registration of corn genotypes using molecular markers].

In plant genetics and breeding, second-generation molecular markers allow detailed characterization of plant genotypes. Unique genotypes at ten simple sequence repeat (SSR) loci were established for 40 maize accessions by means of PCR. For every locus, SSR analysis revealed heterozygotes among simple hybrids, which made it possible to identify the parental forms with a high probability of exclusion of nonparental forms. A system was proposed for registration of maize genotypes in the form of genetic formulas reflecting the allelic state of microsatellite loci, in order to catalog, preserve, and effectively employ the existing maize gene pool in breeding.

Base Sequence↗

The simple repeat poly(dT-dG).poly(dC-dA) common to eukaryotes is absent from eubacteria and archaebacteria and rare in protozoans.

Genomic DNA from a wide variety of prokaryotic and eukaryotic organisms has been assayed for the simple repeat sequence poly(dT-dG).poly(dC-dA) by Southern blotting and DNA slot blot hybridizations. Consistent with findings of others, we have found the simple alternating sequence to be present in multiple copies in all organisms in the animal kingdom (e.g., mammals, reptiles, amphibians, fish, crustaceans, insects, jellyfish, nematodes). The TG element was also found in lower eukaryotes (Saccharomyces cerevisiae, Neurospora crassa, and Dictyostelium discoideum) and at a much lower frequency in protozoans (Oxytricha fallux and Tetrahymena thermophila). The sequence was also repeated in high copy number in a higher plant (Zea mays) as well as at very high levels in a unicellular green alga (Chlamydomonas reinhardi). Although the copy number of the repeat per haploid genome was generally proportional to genome size, there was a greater-than-1,000-fold variation in the number of (TG)25/100-kb genomic DNA. By contrast, no eu-or archaebacterium--including Myxococcus xanthus, whose life cycle is very similar to that of the slime mold Dictyostelium discoideum, and Halobacter volcanii, whose genome contains other repeated sequences--was found whose genomic DNA contained this sequence in detectable amounts. A computer search also failed to find the TG element in human mitochondrial DNA.

Animals↗

Characterization of microsatellites from flow-sorted porcine chromosome 13.

Porcine flow-sorted Chromosome (Chr) 13 was PCR amplified with primers based on porcine short interspersed element (SINE) sequences. The product was cloned, gridded in microtiter plates, and screened with a [GT]10 oligonucleotide which gave 45 positive clones. Sequencing of these clones showed that 36 were unique, and 26 [GT]n microsatellites were characterized. Six other simple repeat sequences, the majority of which were associated with the 3' end of the SINE sequence, were also detected. Twenty-one primers sets were selected, and 13 of these detected useful polymorphisms in the grandparents (n = 26) of the European porcine mapping collaboration (PiGMaP) reference families. These 13 markers were mapped in the "PiGMaP" reference families, and a two-point linkage analysis was performed. The Lod scores indicated that three of the markers were not linked and the remaining 11 formed two linkage groups of two and nine markers respectively. The larger linkage group was also linked to the transferrin locus, permitting assignment of nine markers to porcine Chr 13.

Animals↗

Discrimination of DNA duplexes with matched and mismatched tandem repeats by T4 endonuclease VII.

A simple in situ method for scoring short tandem DNA repeat length has been developed using T4 endonuclease VII. This method measures tandem repeated simple sequences embedded in unique sequences. Single-stranded loops are formed on duplexes containing mismatched (different) numbers of tandem repeats. No single stranded loops are formed on structures containing matched (identical) numbers of tandem repeats. The matched and mismatched loop structures were distinguished and differentially labeled by enzymatic treatment with T4 endonuclease VII.

Bacteriophage T4↗

Microsatellite allele frequencies in humans and chimpanzees, with implications for constraints on allele size.

The distributions of allele sizes at eight simple-sequence repeat (SSR) or microsatellite loci in chimpanzees are found and compared with the distributions previously obtained from several human populations. At several loci, the differences in average allele size between chimpanzees and humans are sufficiently small that there might be a constraint on the evolution of average allele size. Furthermore, a model that allows for a bias in the mutation process shows that for some loci a weak bias can account for the observations. Several alleles at one of the loci (Mfd 59) were sequenced. Differences between alleles of different lengths were found to be more complex than previously assumed. An 8-base-pair deletion was present in the nonvariable region of the chimpanzee locus. This locus contains a previously unrecognized repeated region, which is imperfect in humans and perfect in chimpanzees. The apparently greater opportunity for mutation conferred by the two perfect repeat regions in chimpanzees is reflected in the higher variance in repeat number at Mfd 59 in chimpanzees than in humans. These data indicate that interspecific differences in allele length are not always attributable to simple changes in the number of repeats.

Alleles↗

The locus of sequence-directed and protein-induced DNA bending.

The bending locus of trypanosome kinetoplast DNA, identified by gel electrophoresis, has tracts of a simple repeat sequence (CA5-6 T) symmetrically distributed about it, with a repeat interval of 10 base pairs. The analogous bending induced when catabolite gene activating protein binds to its recognition sequence near the promoter of the Escherichia coli lac operon is centred on a site about 5-7 base pairs away from the centre of the protein binding site.

Animals↗

The topographic organization of repetitive DNA in the human nucleolus.

The nucleolus is a highly specialized nuclear domain where ribosomal DNA (rDNA) is transcribed and preribosomes are assembled. We investigated the molecular organization of the human lymphocyte nucleolus by fluorescence in situ hybridization and confocal laser scanning microscopy and found that transcribed rDNA and nontranscribed ribosomal intergenic spacer (IGS) sequences colocalized to discrete regions frequently on the nucleolar periphery of phytohemagglutinin-stimulated cells. The 5S rDNA gene cluster located on the long arm of chromosome 1 was not regularly associated with the nucleolus. Short interspersed (SINE) Alu elements detected by BLUR 11 were distributed diffusely throughout the nucleus but were severely underrepresented in the nucleolus, whereas an Alu element subcloned from the IGS detected sequences enriched in the nucleolus but sparsely represented in the remainder of the nucleus. In contrast, long interspersed (LINE) Kpn elements, which were located at the nucleolus, were not found in rDNA but were identified outside the ribosomal gene complex on the short arm of at least one acrocentric chromosome. A human chromosome 21-derived alphoid sequence that hybridized to the centromere was localized outside but near the nucleolus, and nonribosomal DNA consisting of a tandemly repeated simple sequence cluster derived from the short arm of chromosome 15 was organized in a compact fashion in the nucleolus. Our study provides new insight into the content and structure of the human nucleolus and illustrates that the unique organization of repetitive DNA on the acrocentric chromosome short arms is reflected in the topographic organization of the nucleolus.

Blotting, Southern↗

Characterization of genetic markers in the 3' end of the apo B gene and their use in family and population studies.

The 3' end of the apo B gene is highly polymorphic. Two point mutations in the coding sequence of the gene create EcoRI (E+, E-) and XbaI (X+, X-) RFLPs. The two loci are in random association and the frequency of the four haplotypes, E+X+, E+X-, E-X+ and E-X- in the normolipidaemic population are 0.68, 0.30, 0.02 and 0.00, respectively. Although the polymorphic nucleotide underlying the EcoRI RFLP creates an amino acid substitution in the apo B protein (Glu----Lys) in a region close to a putative LDL-receptor recognition site(s), we find no statistically significant difference in the frequency of the apo BGlu and apo BLys alleles in hyperlipidaemic patients (familial hypercholesterolaemia, type IIA with no tendon xanthomas, IIB and probably IV) and the normolipidaemic population. In contrast, we confirm previous findings, that the X+ allele is in linkage disequilibrium with a genetic locus that predisposes to the development of higher fasting plasma triglyceride levels than the X- allele. We have characterized a highly polymorphic region immediately 3' to the apo B gene. At least 5 alleles of this locus exist in the population and our family studies show it should be an extremely informative locus to use in studies where polymorphic or mutant apo B alleles are suspected to underly certain forms of familial hyperlipidaemia. DNA sequence analysis of this highly polymorphic locus shows that the variation is entirely attributable to the number of times the simple repeating sequence 5'-TTTATAATTAAAATATTTATAATTAAATAT-3' is present.

Adult↗

Cloning yeast telomeres on linear plasmid vectors.

We have constructed a linear yeast plasmid by joining fragments from the termini of Tetrahymena ribosomal DNA to a yeast vector. Structural features of the terminus region of the Tetrahymena rDNA plasmid maintained in the yeast linear plasmid include a set of specifically placed single-strand interruptions within the cluster of hexanucleotide (C4A2) repeat units. An artificially constructed hairpin terminus was unable to stabilize a linear plasmid in yeast. The fact that yeast can recognize and use DNA ends from the distantly related organism Tetrahymena suggests that the structural features required for telomere replication and resolution have been highly conserved in evolution. The linear plasmid was used as a vector to clone chromosomal telomeres from yeast. One Tetrahymena end was removed by restriction digestion, and yeast fragments that could function as an end on a linear plasmid were selected. Restriction mapping and hybridization analysis demonstrated that these fragments were yeast telomeres, and suggested that all yeast chromosomes might have a common telomere sequence. Yeast telomeres appear to be similar in structure to the rDNA of Tetrahymena, in which specific nicks or gaps are present within a simple repeated sequence near the terminus of the DNA.

Animals↗

Hypervariable microsatellites provide a general source of polymorphic DNA markers for the chloroplast genome.

BACKGROUND: The study of plant populations is greatly facilitated by the deployment of chloroplast DNA markers. Asymmetric inheritance, lower effective population sizes and perceived lower mutation rates indicate that the chloroplast genome may have different patterns of genetic diversity compared to nuclear genomes. Convenient assays that would allow intraspecific chloroplast variability to be detected are required. RESULTS: Eukaryote nuclear genomes contain ubiquitous simple sequence repeat (microsatellite) loci that are highly polymorphic in length; these polymorphisms can be rapidly typed by the polymerase chain reaction (PCR). Using primers flanking simple mononucleotide repeat motifs in the chloroplast DNA of annual and perennial soybean species, we demonstrate that microsatellites in the chloroplast genome also exhibit length variation, and that this polymorphism is due to changes in the repeat region. Furthermore, we have observed a nonrandom geographic distribution of variations at these loci, and have examined the number and location of such repeats within the chloroplast genomes of other species. CONCLUSIONS: PCR-based analysis of mononucleotide repeats may be used to detect both intraspecific and interspecific variability in the chloroplast genomes of seed plants. The analysis of polymorphic microsatellites thus provides an important experimental tool to examine a range of issues in plant genetics.

Base Sequence↗