PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Simple repeated sequences”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Simple sequence repeats in the Helicobacter pylori genome.

We describe an integrated system for the analysis of DNA sequence motifs within complete bacterial genome sequences. This system is based around ACeDB, a genome database with an integrated graphical user interface; we identify and display motifs in the context of genetic, sequence and bibliographic data. Tomb et aL (1997) previously reported the identification of contingency genes in Helicobacter pylori through their association with homopolymeric tracts and dinucleotide repeats. With this as a starting point, we validated the system by a search for this type of repeat and used the contextual information to assess the likelihood that they mediate phase variation in the associated open reading frames (ORFs). We found all of the repeats previously described, and identified 27 putative phase-variable genes (including 17 previously described). These could be divided into three groups: lipopolysaccharide (LPS) biosynthesis, cell-surface-associated proteins and DNA restriction/modification systems. Five of the putative genes did not have obvious homologues in any of the public domain sequence databases. The reading frame of some ORFs was disrupted by the presence of the repeats, including the alpha(1-2) fucosyltransferase gene, necessary for the synthesis of the Lewis Y epitope. An additional benefit of this approach is that the results of each search can be analysed further and compared with those from other genomes. This revealed that H. pylori has an unusually high frequency of homopurine:homopyrimidine repeats suggesting mechanistic biases that favour their presence and instability.

Base Sequence↗

Genetic variation in five Mediterranean populations of Juniperus phoenicea as revealed by inter-simple sequence repeat (ISSR) markers.

BACKGROUND AND AIMS: The assessment of the genetic variability and the identification of isolated populations within a given species represent important information to plan conservation strategies on a genetic basis. In this work, the genetic variability in five natural populations of Juniperus phoenicea, three from Sardinia, one from Cyprus and the last one in the Maritime Alps was analysed by means of ISSRs, on the hypothesis that the latter could have been a refugial one during the last glaciation. METHODS: ISSRs were chosen because of their ability to detect variation without any prior sequence information. The use of three primers yielded 45 reproducible, polymorphic bands, which were utilized to estimate the basic parameters of genetic variability and diversity. KEY RESULTS: All of the populations analysed harboured an adequate amount of genetic variability, with H(S) = 0.1299. The proportion of genetic diversity between populations has been estimated by G(ST) = 0.12. The three Sardinian populations are separated, as tested by AMOVA, from the Cyprus and the continental ones. CONCLUSIONS: The results indicate that geographical isolation has represented a major barrier to gene flow in Juniperus phoenicea. This work represents a first step towards a full genetic characterization of a conifer from the Mediterranean, a world biodiversity hotspot confronted with climate change, and thus contributes towards the planning of genetics-informed conservation strategies.

Biodiversity↗

Assessing probability of ancestry using simple sequence repeat profiles: applications to maize hybrids and inbreds.

Determination of parentage is fundamental to the study of biology and to applications such as the identification of pedigrees. Limitations to studies of parentage have stemmed from the use of an insufficient number of hypervariable loci and mismatches of alleles that can be caused by mutation or by laboratory error and that can generate false exclusions. Furthermore, most studies of parentage have been limited to comparisons of small numbers of specific parent-progeny triplets thereby precluding large-scale surveys of candidates where there may be no prior knowledge of parentage. We present an algorithm that can determine probability of parentage in circumstances where there is no prior knowledge of pedigree and that is robust in the face of missing data or mistyped data. We present data from 54 maize hybrids and 586 maize inbreds that were profiled using 195 SSR loci including simulations of additional levels of missing and mistyped data to demonstrate the utility and flexibility of this algorithm.

Algorithms↗

Assessing probability of ancestry using simple sequence repeat profiles: applications to maize inbred lines and soybean varieties.

Determining parentage is a fundamental problem in biology and in applications such as identifying pedigrees. Difficulties inferring parentage derive from extensive inbreeding within the population, whether natural or planned; using an insufficient number of hypervariable loci; and from allele mismatches caused by mutation or by laboratory errors that generate false exclusions. Many studies of parentage have been limited to comparisons of small numbers of specific parent-progeny triplets. There have been few large-scale surveys of candidates in which there is no prior knowledge of parentage. We present an algorithm that determines the probability of parentage in circumstances where there is no prior knowledge of pedigree and that is robust in the face of missing data and mistyped data. The focus is parentage of an inbred line having uncertain ancestry. The algorithm is a variation of a previously published hybrid-focused algorithm. We describe the algorithm and demonstrate its performance in determining parentage of 43 inbred varieties of soybean that have been profiled using 236 SSR loci and from seven inbred varieties of maize that were profiled using 70 SSR loci. We include simulations of additional levels of missing and mistyped data to show the algorithm's utility and flexibility.

Data Interpretation, Statistical↗

Simple repeated sequences in human satellite DNA.

In an extensive analysis, using a range of restriction endonucleases, HinfI and TaqI were found to differentiate satellites I, II and III & IV. Satellite I is resistant to digestion by TaqI, but is cleaved by HinfI to yield three major fragments of approximate size 770, 850 and 950bp, associated in a single length of DNA. The 770bp fragment contains recognition sites for a number of other enzymes, whereas the 850 and 950bp fragments are "silent" by restriction enzyme analysis. Satellite II is digested by HinfI into a large number of very small (10-80bp) fragments, many of which also contain TaqI sites. A proportion of the HinfI sites in satellite II have the sequence 5'GA(GC)TC. The HinfI digestion products of satellites III and IV form a complete ladder, stretching from 15bp or less to more than 250bp, with adjacent multimers separated by an increment of 5bp. The ladder fragments do not contain TaqI sites and all HinfI sites have the sequence 5'GA(AT)TC. Three fragments from the HinfI ladder of satellite III have been sequenced, and all consist of a tandemly repeated 5bp sequence, 5'TTCCA, with a non-repeated, G+C rich sequence, 9bp in length, at the 3' end.

Base Composition↗

Rapid, high fidelity analysis of simple sequence repeats on an electronically active DNA microchip.

We describe a method for the discrimination of short tandem repeat (STR) alleles based on active microarray hybridization. An essential factor in this method is electronic hybridization of the target DNA, at high stringency, in <5 min. High stringency is critical to avoid slippage of hybrids along repeat tracts at allele-specific test sites in the array. These conditions are attainable only with hybridization kinetics realized by electronic concentration of DNA. A sandwich hybrid is assembled, in which proper base stacking of juxtaposed terminal nucleotides results in a thermodynamically favored complex. The increased stability of this complex relative to non-stacked termini and/or base pair mismatches is used to determine the identification of STR alleles. This method is capable of simultaneous and precise identification of alleles containing different numbers of repeats, as well as mutations within these repeats. Given the throughput capabilities of microarrays our system has the potential to enhance the use of microsatellites in forensic criminology, diagnostics and genetic mapping.

Alleles↗

Differential distribution of simple sequence repeats in eukaryotic genome sequences.

Complete chromosome/genome sequences available from humans, Drosophila melanogaster, Caenorhabditis elegans, Arabidopsis thaliana, and Saccharomyces cerevisiae were analyzed for the occurrence of mono-, di-, tri-, and tetranucleotide repeats. In all of the genomes studied, dinucleotide repeat stretches tended to be longer than other repeats. Additionally, tetranucleotide repeats in humans and trinucleotide repeats in Drosophila also seemed to be longer. Although the trends for different repeats are similar between different chromosomes within a genome, the density of repeats may vary between different chromosomes of the same species. The abundance or rarity of various di- and trinucleotide repeats in different genomes cannot be explained by nucleotide composition of a sequence or potential of repeated motifs to form alternative DNA structures. This suggests that in addition to nucleotide composition of repeat motifs, characteristic DNA replication/repair/recombination machinery might play an important role in the genesis of repeats. Moreover, analysis of complete genome coding DNA sequences of Drosophila, C. elegans, and yeast indicated that expansions of codon repeats corresponding to small hydrophilic amino acids are tolerated more, while strong selection pressures probably eliminate codon repeats encoding hydrophobic and basic amino acids. The locations and sequences of all of the repeat loci detected in genome sequences and coding DNA sequences are available at http://www.ncl-india.org/ssr and could be useful for further studies.

Animals↗

DNA fingerprints of rice (Oryza sativa) obtained from hypervariable chloroplast simple sequence repeats.

The aim of this research was to develop a convenient polymerase chain reaction-based assay that would allow intraspecific chloroplast variability to be detected. Our approach is based on the detection of length polymorphism within chloroplast mononucleotide microsatellite loci. Information from the fully sequenced rice chloroplast genome was used to identify 12 regions with a minimum of ten uninterrupted mononucleotide repeats. Primers flanking these repeats were used in conjunction with polymerase chain reaction to examine levels of polymorphism in six wild and 14 cultivated rice accessions. A total of six of the primer pairs revealed length polymorphism with between two and five size variants being detected. Diversity indices varied between 0.07 and 0.72. The length variation detected at multiple, physically linked sites was used to identify 15 unique haplotypes with an overall diversity index of 0.90. This level of polymorphism is sufficiently high to allow chloroplast variability to be studied at the intraspecific level. An additional 47 Oryza sativa accessions were also assayed with 31 unique chloroplast haplotypes being detected. The distribution of these haplotypes is described in relation to isozyme groupings and subspecies differentiation. The relevance and implications of these results for plant population genetics and the management of germplasm collections is discussed.

Chloroplasts↗

Inter simple sequence repeat analysis of genetic diversity and relationships in cultivated barley of Nordic and Baltic origin.

This study evaluates putative changes of genetic diversity and relationships of barley in the Nordic and Baltic countries that might have taken place during the last century as a result of commercial breeding. Four ISSR primers were used to analyse 227 accessions, yielding a total of 47 polymorphic loci. Shannon-Weaver diversity values for each locus ranged from 0.012 to 0.693. Overall, there were no significant changes of genetic diversity observed over time. A significant decrease of diversity was, however, observed in material from the southern parts of the Nordic and Baltic countries. In material from the northern parts no decrease of diversity was observed. The genetic diversity of six-rowed barley bred in the middle of the 20th century was low, but there was no significant difference between modern accessions and landraces or old cultivars. The magnitude in changes of genetic diversity differed also in material from different countries of origin. A cluster analysis clearly separated the material into two groups. The first cluster included 86.5% of all six-rowed accessions, whereas the second cluster contained 97.4% of all two-rowed accessions.

Genetic Markers↗

Simple sequence repeat markers distinguish among morphotypes of Sphaeropsis sapinea.

Sphaeropsis sapinea is a fungal endophyte of Pinus spp. that can cause disease following predisposition of trees by biotic or abiotic stresses. Four morphotypes of S. sapinea have been described from within the natural range of the fungus, while only one morphotype has been identified on exotic pines in the Southern Hemisphere. The aim of this study was to develop robust polymorphic markers that could be used in both taxonomic and population studies. Inter-short-sequence-repeat primers containing microsatellite sequences and degenerate anchors at the 5' end were used to target microsatellite-rich areas in an S. sapinea isolate. PCR amplification using an annealing temperature of 49 degrees C resulted in profiles containing 5 to 10 bands. These bands were cloned and sequenced, and new short-sequence-repeat (SSR) primer pairs were designed that flanked microsatellite-rich regions. Eleven polymorphic SSR markers were tested on 40 isolates of S. sapinea representing different morphotypes as well as on 2 isolates of the closely related species Botryosphaeria obtusa. The putative I morphotype was found to be identical to B. obtusa. Otherwise, the markers clearly distinguished the remaining three morphotypes and, furthermore, showed that the C morphotype was more closely related to the A than the B morphotype. The B morphotype was the most genetically diverse, and the isolates could be further divided based on their geographic origins. Sequencing of different alleles from each locus showed that the most polymorphic markers had mutations within a microsatellite sequence.

Alleles↗

Simple sequence repeat diversity in diploid and tetraploid Coffea species.

Thirty-four fluorescently labeled microsatellite markers were used to assess genetic diversity in a set of 30 Coffea accessions from the CENICAFE germplasm bank in Colombia. The plant material included one sample per accession of seven East African accessions representing five diploid species and 23 wild and cultivated tetraploid accessions of Coffea arabica from Africa, Indonesia, and South America. More allelic diversity was detected among the five diploid species than among the 23 tetraploid genotypes. The diploid species averaged 3.6 alleles/locus and had an average polymorphism information content (PIC) value of 0.6, whereas the wild tetraploids averaged 2.5 alleles/locus and had an average PIC value of 0.3 and the cultivated tetraploids (C. arabica cultivars) averaged 1.9 alleles/locus and had an average PIC value of 0.22. Fifty-five percent of the alleles found in the wild tetraploids were not shared with cultivated C. arabica genotypes, supporting the idea that the wild tetraploid ancestors from Ethiopia could be used productively as a source of novel genetic variation to expand the gene pool of elite C. arabica germplasm.

Base Sequence↗

Development and mapping of EST-derived simple sequence repeat markers for hexaploid wheat.

Expressed sequence tags (ESTs) are a valuable source of molecular markers. To enhance the resolution of an existing linkage map and to identify putative functional polymorphic gene loci in hexaploid wheat (Triticum aestivum L.), over 260,000 ESTs from 5 different grass species were analyzed and 5418 SSR-containing sequences were identified. Using sequence similarity analysis, 156 cross-species superclusters and 138 singletons were used to develop primer pairs, which were then tested on the genomic DNA of barley (Hordeum vulgare), maize (Zea mays), rice (Oryza sativa), and wheat. Three-hundred sixty-eight primer pairs produced PCR amplicons from at least one species and 227 primer pairs amplified DNA from two or more species. EST-SSR sequences containing dinucleotide motifs were significantly more polymorphic (74%) than those containing trinucleotides (56%), and polymorphism was similar for markers in both coding and 5' untranslated (UTR) regions. Out of 112 EST-SSR markers, 90 identified 149 loci that were integrated into a reference wheat genetic map. These loci were distributed on 19 of the 21 wheat chromosomes and were clustered in the distal chromosomal regions. Multiple-loci were detected by 39% of the primer pairs. Of the 90 mapped ESTs, putative functions for 22 were identified using BLASTX queries. In addition, 80 EST-SSR markers (104 loci) were located to chromosomes using nullisomic-tetrasomic lines. The enhanced map from this study provides a basis for comparative mapping using orthologous and PCR-based markers and for identification of expressed genes possibly affecting important traits in wheat.

Chromosome Mapping↗