PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “simple sequence repeat”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Development and validation of whole-genome SSR markers in sugar beet (Beta vulgaris L.).

Sugar beet (Beta vulgaris L.) is an important sugar and cash crop worldwide. To systematically characterize SSR (Simple Sequence Repeat) loci across sugar beet chromosomes and enable the precise identification of germplasm resources, this study conducted a genome-wide scan for SSR loci, analyzed their distribution patterns, and determined their genotypes using resequencing data from 123 sugar beet varieties. The results revealed an abundance of SSR loci in the sugar beet genome, with a total of 135, 379 identified, from which 135, 344 pairs of SSR primers were designed (135, 344 primer pairs successfully designed; 35 loci failed to meet design criteria). Specifically, 31, 748 primer pairs were designed based on SSRs located in unassigned scaffolds, and 103, 596 primer pairs from SSRs assigned to the nine chromosomes. Through bioinformatic analysis, we identified 28, 768 SSR primers located in multi-copy genes with PIC (Polymorphism Information Content) ≥ 0.5, and 2, 326 SSR markers located in single-copy genes residing in various genic regions (among which 543 had PIC ≥ 0.5, with the highest reaching 0.776). PCR (Polymerase Chain Reaction) validation confirmed 20 robust and polymorphic markers producing clear and reproducible bands. Among them, 10 SSR primers located in multi-copy genes exhibited three or more polymorphic types, and 10 markers located in single-copy genes displayed 2-3 polymorphic types. The most polymorphic marker, YCD-4-2, detected 11 polymorphic types across 48 varieties. Furthermore, to explore markers with potential functional significance, we annotated the genes harboring SSR markers located in single-copy genes. The results showed that 1, 264 SSRs located in single-copy genes were localized to 967 genes, which are significantly enriched in pathways related to carbohydrate metabolism, stress responses, and plant-pathogen interactions. The 20 validated markers and the 2, 326 SSRs located in single-copy genes provided in this study can be directly applied to fingerprinting of sugar beet varieties, seed purity testing, and marker-assisted selection, thus representing a practical resource for molecular breeding.

genome-wide↗

Intimate association of microsatellite repeats with retrotransposons and other dispersed repetitive elements in barley.

Simple sequence repeat (SSR)-based genetic markers are being actively developed for the majority of crop plant species. In barley, characterization of 290 dinucleotide repeat-containing clones from SSR-enriched libraries has revealed that a high percentage are associated with cereal retrotransposon-like and other dispersed repetitive elements. Associations found were with BARE-1, WIS2-1A, PREM1 and the dispersed repetitive element R173. Additional similarities between different SSR clones, which have no matches in DNA sequence databases, indicate that this phenomenon is probably widespread in the barley genome. Sequence homologies to the non-coding regions of several cereal genes were also explained by homology to mobile genetic elements. The SSRs found can therefore be classified into two types: (1) those with unique sequences on either flank, and (2) those which are intimately associated with retro-transposons and other dispersed repetitive elements. As the cereal genome is thought to consist largely of this type of DNA, some random association would be expected. However, the conserved positions of the SSRs, relative to repetitive elements, indicate that they have arisen non-randomly. Furthermore, this class of SSRs can be classified into three subtypes: (1) those which are positioned 3' of a transposable element with unique sequence on the other flank, (2) those positioned 5' of a transposable element, and (3) those which have arisen from an internal sequence and so have transposable element sequence on both flanks. The first appear to be analogous to the class of SSRs in mammalian systems which are associated with Alu elements and SINEs (short interspersed elements) and which have been postulated to arise following integration of an extended and polyadenylated retro-transcript into the host genome, followed by mutation of the poly(A) tract and expansion into an SSR. For the second, we postulate that a proto-SSR (A-rich sequence) has acted as a 'landing pad' for transposable element insertion (rather than being the result of insertion), while the third includes those which have evolved as a component of an active transposable element which has spread throughout the genome during bursts of transposition activity. The implications of these associations for genome and SSR evolution in barley are discussed.

Base Sequence↗

DosDNA occurs along yeast chromosomes, regardless of functional significance of the sequence.

Complex genomes contain numerous simple sequence repeats, the biological significance of which remains obscure. Recently it has been shown that several human diseases are the result of changes in such sequences. Thus it has become urgent to undertake a systematic study of their properties. We have set the task of describing as completely as possible the set of sequences which contain bases organized according to symmetrical elements, the dosDNA: defined ordered sequence. Examination of local anomalies in dinucleotide composition serves to identify dosDNA zones in the genome. The study of chromosomes II, III, VIII and XI of Saccharomyces cerevisiae reveals these dosDNA zones comprise about 2% of the genome. They are regularly distributed along the chromosomes, regardless of the functional significance of the sequence. A more detailed analysis of dosDNA segments seems to indicate that simple repeats are the consequence of local properties of the chromosome, and not due to any motif in particular.

Base Sequence↗

TROLL--tandem repeat occurrence locator.

SUMMARY: Tandem Repeat Occurrence Locator (TROLL), is a light-weight Simple Sequence Repeat (SSR) finder based on a slight modification of the Aho-Corasick algorithm. It is fast and only requires a standard Personal Computer (PC) to operate. We report running times of 127 s to find all SSRs of length 20 bp or more on the complete Arabdopsis genome--approx. 130 Mbases divided in five chromosomes--using a PC Athlon 650 MHz with 256 MB of RAM. AVAILABILITY: TROLL is an open source project and is available at http://finder.sourceforge.net.

Algorithms↗

Characterization of mononucleotide repeats in sequenced prokaryotic genomes.

The increasing availability of prokaryotic genome sequences has shown that simple sequence repeats (SSRs) are widespread in prokaryotes and that there is extensive variation in their length, number and distribution. Considering their potential importance in generating genomic diversity, we determined the distribution of a specific group of SSRs, mononucleotide repeats of size between 5 and 13 nt, in 157 sequenced prokaryotic genomes. The data obtained in the present study show that (i) a large number of mononucleotide SSRs is present in all prokaryotic genomes investigated, (ii) shorter repeats are much more abundant than longer repeats, and (iii) in the majority of the genomes, longer mononucleotide SSRs are excluded from coding regions although we identified several organisms where mononucleotide SSRs are not excluded from the coding regions. We also observed that some genomes contain more mononucleotide SSRs than expected, while others contain significantly less. Bacterial genomes that contain much less mononucleotide SSRs than expected are generally larger and more GC-rich, while bacterial genomes that contain much more mononucleotide SSRs than expected are in general smaller and more AT-rich. Finally, we also noted that genomes that contain a high fraction of horizontally transferred genes have a lower mononucleotide SSR density and that A and T are generally overrepresented in mononucleotide SSRs.

Base Sequence↗

Cross-species transferability and mapping of genomic and cDNA SSRs in pines.

Two unigene datasets of Pinus taeda and Pinus pinaster were screened to detect di-, tri- and tetranucleotide repeated motifs using the SSRIT script. A total of 419 simple sequence repeats (SSRs) were identified, from which only 12.8% overlapped between the two sets. The position of the SSRs within their coding sequences were predicted using FrameD. Trinucleotides appeared to be the most abundant repeated motif (63 and 51% in P. taeda and P. pinaster, respectively) and tended to be found within translated regions (76% in both species), whereas dinucleotide repeats were preferentially found within the 5'- and 3'-untranslated regions (75 and 65%, respectively). Fifty-three primer pairs amplifying a single PCR fragment in the source species (mainly P. taeda), were tested for amplification in six other pine species. The amplification rate with other pine species was high and corresponded with the phylogenetic distance between species, varying from 64.6% in P. canariensis to 94.2% in P. radiata. Genomic SSRs were found to be less transferable; 58 of the 107 primer pairs (i.e. 54%) derived from P. radiata amplified a single fragment in P. pinaster. Nine cDNA-SSRs were located to their chromosomes in two P. pinaster linkage maps. The level of polymorphism of these cDNA-SSRs was compared to that of previously and newly developed genomic-SSRs. Overall, genomic SSRs tend to perform better in terms of heterozygosity and number of alleles. This study suggests that useful SSR markers can be developed from pine ESTs.

Base Sequence↗

Plastid genome evolution and phylogenomics with broad taxon sampling: insights into intrafamilial classification of Hamamelidaceae.

Hamamelidaceae, within the order Saxifragales, comprises 27 genera and approximately 120 species. The family has a pantropical and temperate distribution across the Americas, Asia, Africa, and Australia. Previous molecular investigations, constrained by limited taxon sampling and inadequate genetic markers, supported a five-subfamily classification system. However, these studies predominantly focused on Asian taxa, resulting in poor resolution of the evolutionary relationships among American, African, and Australian genera. To address these sampling gaps, we employed near-complete generic sampling (26 of 27 genera) to investigate plastome architecture, structural variation, and phylogenetic relationships. We newly sequenced and assembled 15 plastid genomes representing geographically and taxonomically underrepresented genera and analyzed them alongside 59 publicly available plastomes retrieved from GenBank. Plastid genomes exhibited conserved quadripartite architecture with sizes ranging from 158, 076 bp to 160, 814 bp, minimal structural variation, consistent GC content (37.7-38.2%), and identical gene order. Inverted repeat (IR) regions had limited size variation (26, 211-26, 429 bp). Simple sequence repeat (SSR) distribution (2, 219 loci) showed no clear correlation with the genus-level phylogenetic relationships. We identified ten hypervariable regions, including coding sequences (accD, ycf1, clpP, ndhF, and rpl22) and intergenic spacers (rpl33-rps18, the trnG-UCC intron, trnH-GUG-psbA, accD-psaI, and petA-psbJ), as promising candidate regions for future applications in species delimitation and phylogenetic studies. Phylogenetic analyses revealed largely congruent topologies across datasets and methods, providing improved resolution and strong support for most subfamilial and tribal relationships compared with previous studies. This study highlights the utility of plastid genome data for resolving deep-level phylogenetic relationships within Hamamelidaceae. The genome architecture reflects the high conservation of plastid genomes, while the identified mutation hotspots represent potential resources for future taxonomic and phylogenetic studies. Our results support the existing subfamily classification while improving geographical coverage and generic representation, providing a robust framework for future taxonomic and evolutionary studies of this globally distributed and taxonomically complex family.

Hamamelidaceae↗

Development of microsatellite markers and characterization of simple sequence length polymorphism (SSLP) in rice (Oryza sativa L.).

Microsatellite markers containing simple sequence repeats (SSR) are a valuable tool for genetic analysis. Our objective is to augment the existing RFLP map of rice with simple sequence length polymorphisms (SSLP). In this study, we describe 20 new microsatellite markers that have been assigned to positions along the rice chromosomes, characterized for their allelic diversity in cultivated and wild rice, and tested for amplification in distantly related species. Our results indicate that the genomic distribution of microsatellites in rice appears to be random, with no obvious bias for, or clustering in particular regions, that mapping results are identical in intersubspecific and interspecific populations, and that amplification in wild relatives of Oryza sativa is reliable in species most closely related to cultivated rice but becomes less successful as the genetic distance increases. Sequence analysis of SSLP alleles in three related indica varieties demonstrated the clustering of complex arrays of SSR motifs in a single 300-bp region with independent variation in each. Two microsatellite markers amplified multiple loci that were mapped onto independent rice chromosomes, suggesting the presence of duplicated regions within the rice genome. The availability of increasing numbers of mapped SSLP markers can be expected to increase the power and resolution of genome analysis in rice.

Alleles↗

Identification and characterization of dinucleotide repeat (CA)n markers for genetic mapping in dog.

A large block of simple sequence repeat (SSR) polymorphisms for the dog genome has been isolated and characterized. Screening of primary libraries by conventional hybridization methods as well as by screening of enriched marker-selected libraries led to the isolation of a large number of genomic clones that contained (CA)n repeats. The sequences of 101 clones showed that the size and complexity of (CA)n repeats in the dog genome were similar to those reported for these markers in the human genome. Detailed analysis of a representative subset of these markers revealed that most markers were moderately to highly polymorphic, with PIC values exceeding 0.70 for 33% of the markers tested. An association between higher PIC values and markers containing longer (CA)n repeats was observed in these studies, as previously noted for similar markers in the human genome. A list of primer sequences that tag each characterized marker is provided, and a comprehensive system of nomenclature for the dog genome is suggested.

Alleles↗

Long perfect dinucleotide repeats are typical of vertebrates, show motif preferences and size convergence.

Microsatellites are simple sequence repeats (SSRs) showing complex patterns of length, motif sizes, motif sequences, and repeat perfection. We studied the structure of the dinucleotide SSR population at the genome level by analyzing assembled DNA sequence across species. Three dinucleotide populations were distinguished when SSR genome frequency was analyzed as a function of repeat length and repeat perfection. A population of low-perfection SSRs was identified, which is constituted by short repeats and represents the vast majority of genomic dinucleotide SSRs across eukaryotic genomes. In turn, the highly perfect repeats are 30 to 50 times less frequent and, in addition to short repeats, also contain a long repeat population that is uniquely represented in vertebrate species. Distinctive features of this population include the modal peak in the frequency distribution of repeat length and the strong preferential usage of the repeat motifs AC and AG. These results raise the hypothesis that the ability of carrying a distinct population of long, highly perfect dinucleotide repeats in the genome is a late acquisition in chordate evolution. Our analysis also suggests that different dinucleotide repeat populations have different dynamics and are likely to be underlined by different molecular mechanisms of generation and maintenance in the genome. Thus, these observations imply that caution should be taken in extrapolating results from studies on SSR mutability and on SSR phylogenetic comparisons that do not take into account the stratification of dinucelotide populations in the eukaryotic genome.

Animals↗

Requirements for the dGTP-dependent repeat addition processivity of recombinant Tetrahymena telomerase.

Telomerase is a reverse transcriptase responsible for adding simple sequence repeats to chromosome 3'-ends. The template for telomeric repeat synthesis is carried within the RNA component of the telomerase ribonucleoprotein complex. Telomerases can copy their internal templates with repeat addition processivity, reusing the same template multiple times in the extension of a single primer. For some telomerases, optimal repeat addition processivity requires high micromolar dGTP concentrations, a much higher dGTP concentration than required for processive nucleotide addition within a repeat. We have investigated the requirements for dGTP-dependent repeat addition processivity using recombinant Tetrahymena telomerase. By altering the template sequence, we show that repeat addition processivity retains the same dGTP-dependence even if dGTP is not the first nucleotide incorporated in the second repeat. Furthermore, no dNTP other than dGTP can stimulate repeat addition processivity, even if it is the first nucleotide incorporated in the second repeat. Using structural variants of dGTP, we demonstrate that the stimulation of repeat addition processivity is specific for dGTP base and sugar constituents but requires only a single phosphate group. However, all nucleotides that stimulate repeat addition processivity also inhibit or compete with dGTP incorporation into product DNA. By assaying telomerase complexes reconstituted with a variety of altered templates, we find that repeat addition processivity has an unanticipated template or product sequence specificity. Finally, we show that a novel, nascent product DNA binding site establishes dGTP-dependent repeat addition processivity.

Animals↗

Triplet repeats in human genome: distribution and their association with genes and other genomic regions.

MOTIVATION: Simple sequence repeats (SSRs) or microsatellite repeats are found abundantly in many prokaryotic and eukaryotic genomes. Among SSRs, triplet repeats are of special significance because some of them have been linked to various genetic disorders. The objective of the study is to analyze the triplet repeats of complete human genome and to identify the genes that contain the triplet repeats in their coding region. The analysis will help us to identify the candidate genes that have potential for repeat expansion. RESULTS: We have analyzed triplet repeats in the complete human genome from the publicly available sequences. Our analysis revealed that AGC and CCG repeat were predominantly present in the coding regions of the genome while UTRs and the upstream sequences contained CCG repeats in relative abundance. Analysis of density of triplet repeats (bp/Mb) revealed that AAT and AAC were the abundant repeats whereas ACT and ACG were the rare repeats found in human genome. We could identify about 2135 known or predicted genes that were associated with at least one of the triplet repeat types. A large proportion of putative transcripts that were identified by gene finding programs were found to be associated with triplet repeats. These transcripts will be the candidate genes for analysis of triplet repeat expansion and a possible association with disease phenotypes. Identification of 171 genes which contain a minimum of ten repeat units will be of particular interest in future in correlating their association with any disease phenotype due to the expansion potential of repeats present in them. The list of genes and other details of analysis are given in the online supplementary data (http://www.ingenovis.com/tripletrepeats).

Databases, Nucleic Acid↗

Genomic simple repetitive DNAs are targets for differential binding of nuclear proteins.

The biological meaning of abundant simple repetitive DNA sequences in eukaryote genomes is obscure. Therefore, (GAA)n, (GT)n, and composite (GT)n(GA)m, blocks were characterized for protein binding in the repeat and flanking sequences of cloned genomic DNA fragments. In gel mobility shift and competition assays the binding of nuclear proteins to the repeats was specific (including some flanking single copy sequences). DNase footprinting revealed the target sequences within and adjacent to the repeats. Chemical modifications (OsO4, DEPC) demonstrated non-B DNA structures in the polypurine blocks. The binding of nuclear proteins in and around simple repeat sequences refute biological insignificance of all of these ubiquitously interspersed elements.

Animals↗

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny↗

Sequence tagged microsatellite profiling (STMP): improved isolation of DNA sequence flanking target SSRs.

Sequence tagged microsatellite profiling (STMP) enables the rapid development of large numbers of co-dominant DNA markers, known as sequence tagged microsatellites (STMs). Each STM is amplified by PCR using a single primer specific to the conserved DNA sequence flanking the microsatellite repeat in combination with a universal primer that anchors to the 5'-ends of the microsatellites. It is also possible to convert STMs into conventional microsatellite, or simple sequence repeat (SSR), markers that are amplified using a pair of primers flanking the repeat sequence. Here, we describe a modification of the STMP procedure to significantly improve the capacity to convert STMs into conventional SSRs and, therefore, facilitate the development of highly specific DNA markers for purposes such as marker-assisted breeding. The usefulness of this technique was demonstrated in bread wheat.

Conserved Sequence↗

The abundance of various polymorphic microsatellite motifs differs between plants and vertebrates.

The abundance of different simple sequence motifs in plants was accessed through data base searches of DNA sequences and quantitative hybridization with synthetic dinucleotide repeats. Database searches indicated that microsatellites are five times less abundant in the genomes of plants than in mammals. The most common plant repeat motif was AA/TT followed by AT/TA and CT/GA. This group comprised about 75% of all microsatellites with a length of more than 6 repeats. The GT/CA motif being the most abundant dinucleotide repeat in mammals was found to be considerably less frequent in plants. To address the question if plant simple repeat sequences are variable as in mammals, (GT)n and (CT)n microsatellites were isolated from B.napus. Five loci were investigated by PCR-analysis and amplified products were obtained for all microsatellites from B. oleracea, B.napus and B.rapa DNA, but only for one primer pair from B.nigra. Polymorphism was detected for all microsatellites.

Animals↗

Complete sequence of the 45-kb mouse ribosomal DNA repeat: analysis of the intergenic spacer.

DNA from a single bacterial artificial chromosome clone was used to sequence the mouse ribosomal DNA intergenic spacer from the 3' end of the 45S pre-RNA to the spacer promoter (Accession No. AF441733). This made possible the assembly of a complete mouse ribosomal DNA repeat unit (45309 bp long, TPA Accession No. BK000964). Analysis of the intergenic spacer (IGS) showed a high density of simple sequence repeats and transposable elements. The IGS contains two long sequence blocks, which are repeated tandemly. Some of the sequences in these blocks are also present in other parts of the IGS. A difference in the mutation rate along the mouse IGS was observed. The significance of sequence motifs in the IGS for transcription enhancement, transcription termination, origin of replication, and nucleolar organization is discussed.

Animals↗

Laboratory Information Management Software for genotyping workflows: applications in high throughput crop genotyping.

BACKGROUND: With the advances in DNA sequencer-based technologies, it has become possible to automate several steps of the genotyping process leading to increased throughput. To efficiently handle the large amounts of genotypic data generated and help with quality control, there is a strong need for a software system that can help with the tracking of samples and capture and management of data at different steps of the process. Such systems, while serving to manage the workflow precisely, also encourage good laboratory practice by standardizing protocols, recording and annotating data from every step of the workflow. RESULTS: A laboratory information management system (LIMS) has been designed and implemented at the International Crops Research Institute for the Semi-Arid Tropics (ICRISAT) that meets the requirements of a moderately high throughput molecular genotyping facility. The application is designed as modules and is simple to learn and use. The application leads the user through each step of the process from starting an experiment to the storing of output data from the genotype detection step with auto-binning of alleles; thus ensuring that every DNA sample is handled in an identical manner and all the necessary data are captured. The application keeps track of DNA samples and generated data. Data entry into the system is through the use of forms for file uploads. The LIMS provides functions to trace back to the electrophoresis gel files or sample source for any genotypic data and for repeating experiments. The LIMS is being presently used for the capture of high throughput SSR (simple-sequence repeat) genotyping data from the legume (chickpea, groundnut and pigeonpea) and cereal (sorghum and millets) crops of importance in the semi-arid tropics. CONCLUSION: A laboratory information management system is available that has been found useful in the management of microsatellite genotype data in a moderately high throughput genotyping laboratory. The application with source code is freely available for academic users and can be downloaded from http://www.icrisat.org/gt-bt/lims/lims.asp.

Algorithms↗