PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “simple sequence repeat”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Genome-wide analysis of microsatellite repeats in humans: their abundance and density in specific genomic regions.

BACKGROUND: Simple sequence repeats (SSRs) are found in most organisms, and occupy about 3% of the human genome. Although it is becoming clear that such repeats are important in genomic organization and function and may be associated with disease conditions, their systematic analysis has not been reported. This is the first report examining the distribution and density of simple sequence repeats (1-6 base-pairs (bp)) in the entire human genome. RESULTS: The densities of SSRs across the human chromosomes were found to be relatively uniform. However, the overall density of SSR was found to be high in chromosome 19. Triplets and hexamers were more predominant in exonic regions compared to intronic and intergenic regions, except for chromosome Y. Comparison of densities of various SSRs revealed that whereas trimers and pentamers showed a similar pattern (500-1,000 bp/Mb) across the chromosomes, di- tetra- and hexa-nucleotide repeats showed patterns of higher (2,000-3,000 bp/Mb) density. Repeats of the same nucleotide were found to be higher than other repeat types. Repeats of A, AT, AC, AAT, AAC, AAG, AGC, AAAC, AAAT, AAAG, AAGG, AGAT predominate, whereas repeats of C, CG, ACT, ACG, AACC, AACG, AACT, AAGC, AAGT, ACCC, ACCG, ACCT, CCCG and CCGG are rare. CONCLUSIONS: The overall SSR density was comparable in all chromosomes. The density of different repeats, however, showed significant variation. Tri- and hexa-nucleotide repeats are more abundant in exons, whereas other repeats are more abundant in non-coding regions.

Dinucleotide Repeats↗

Characterization of telomere DNA from Neurospora crassa.

The nucleotide sequence of the telomere at the right end of linkage group V (VR) in the standard OR23-IV-A strain of the filamentous fungus, Neurospora crassa, reveals the following features. At the chromosome terminus, tandem repeats of the hexanucleotide TTAGGG are present. Immediately centromere-proximal to the simple sequence repeat is a more complex element called Pogo that is reiterated 5-10 times in the genomes of various Neurospora strains. The element possesses several features characteristic of a transposable element: direct repeats of 318 bp flank the element, there is a long internal open reading frame (ORF), and a 3-bp duplication is found at its borders. However, Pogo has other structural features that are more difficult to reconcile with the standard model of a transposable element. A second telomere from Neurospora was also cloned by screening a genomic lambda library with a synthetic oligodeoxyribonucleotide homologous to the simple sequence repeats. This telomere is entirely non-homologous with the VR telomere except for the TTAGGG repeats, has no associated copy of Pogo, and has no nearby ORFs. There are no long stretches of TTAGGG repeats present in the Neurospora genome at non-telomeric sites.

Amino Acid Sequence↗

Comparative Analysis of Chloroplast Genomes Reveals Molecular Evolution and Phylogenetic Relationships in Fraxinus (Fraxinus mandshurica).

Fraxinus mandshurica (Manchurian ash) is an ecologically and economically valuable hardwood tree native to Northeast Asia, yet its genomic resources remain limited. We assembled its complete chloroplast (cp) genome (155,559 bp) using hybrid PacBio and Illumina sequencing and performed comparative, phylogenetic, and evolutionary analyses. The cp genome exhibits a typical quadripartite structure encoding 132 gene copies, comprising 114 unique genes (80 protein-coding, 30 tRNA, and 4 rRNA genes), with 18 genes duplicated in the inverted repeat (IR) regions. Simple sequence repeat analysis revealed dominance of mononucleotide A/T repeats. Phylogenetic analysis of 53 complete cp genomes strongly supported the monophyly of Oleaceae and resolved F. mandshurica as sister to the North American F. nigra, consistent with previously proposed Miocene intercontinental dispersal scenarios between East Asia and North America. Most protein-coding genes were under strong purifying selection (Ka/Ks << 1), whereas petB, rpl2, and several ndh genes showed elevated Ka/Ks values that are suggestive of altered selective constraint but are based on very few substitutions and are therefore not, on their own, evidence of positive selection. Nucleotide diversity (Pi) analysis identified 15 hypervariable intergenic spacers (mean Pi = 0.067), among which trnM-CAU-rps14, ndhJ-ndhK, and petL-petG represent promising candidate barcode regions requiring further validation. This study provides a high-quality, fully annotated cp genome of F. mandshurica and a valuable genomic resource for future phylogenetic, population genetic, and conservation studies of this important genus.

Fraxinus↗

Mapping of digested and undigested random amplified microsatellite polymorphisms in barley.

The broad use of microsatellites as a tool for constructing linkage maps in plants has been limited by the need for sequence data to detect the underlying simple sequence repeats. Therefore, random amplified microsatellite polymorphisms (RAMPs) were studied as an alternative approach for barely mapping. Labelled (GA)n simple sequence repeat primers were combined with RAPD primers of different length and sequence to generate RAMPs. To get additional polymorphisms (called dRAMPs), the obtained products were also analysed after digestion with MseI. There were 0-11 polymorphisms found per primer combination. Sixty RAMPs/dRAMPs identifying 40 new loci were mapped onto a barley RFLP map. The new DNA markers are found on all chromosomes and they increased the length of the barely map by 174 cM to a total of 1270 cM. Interestingly, the RAMPs/dRAMPs caused stretching effects in genome areas where stretching was also observed for AFLPs.

Base Sequence↗

Construction of small-insert genomic DNA libraries highly enriched for microsatellite repeat sequences.

We describe an efficient method for the construction of small-insert genomic libraries enriched for highly polymorphic, simple sequence repeats. With this approach, libraries in which 40-50% of the members contain (CA)n repeats are produced, representing an approximately 50-fold enrichment over conventional small-insert genomic DNA libraries. Briefly, a genomic library with an average insert size of less than 500 base pairs was constructed in a phagemid vector. Amplification of this library in a dut ung strain of Escherichia coli allowed the recovery of the library as closed circular single-stranded DNA with uracil frequently incorporated in place of thymine. This DNA was used as a template for second-strand DNA synthesis, primed with (CA)n or (TG)n oligonucleotides, at elevated temperatures by a thermostable DNA polymerase. Transformation of this mixture into wild-type E. coli strains resulted in the recovery of primer-extended products as a consequence of the strong genetic selection against single-stranded uracil-containing DNA molecules. In this manner, a library highly enriched for the targeted microsatellite-containing clones was recovered. This approach is widely applicable and can be used to generate marker-selected libraries bearing any simple sequence repeat from cDNAs, whole genomes, single chromosomes, or more restricted chromosomal regions of interest.

Animals↗

Mitogenome assembly and phylogenetic relationships of Phalaris arundinacea.

INTRODUCTION: As a perennial herb of Poaceae, Phalaris arundinacea plays key roles in grazing, production, and soil and water conservation because of its well-developed rhizomes and seed dispersal. We assembled and annotated the first mitogenome of P. arundinacea to support evolutionary and taxonomic research. METHODS: We assembled and annotated the first complete mitochondrial genome of P. arundinacea by integrating Illumina short reads with Nanopore long reads via a hybrid assembly strategy. The genome architecture was comprehensively characterized, encompassing codon usage bias, repetitive sequence organization, and inter-organellar genetic exchange with the chloroplast genome. RESULTS AND DISCUSSION: Assembly of the P. arundinacea mitogenome revealed two circular structures with a combined length of 526,717 bp. The genome comprised a set of 37 protein-coding genes (PCGs), 27 tRNAs, and 8 rRNAs, with the rRNA genes exhibiting full assembly (100% coverage). The mitochondrial genome contained 154 forward and 164 palindromic repeats, along with 25 tandem repeats and 124 simple sequence repeats (SSRs). Notably, 102 SSRs were distributed on contig1, predominantly in tetrameric form. Furthermore, 376 RNA editing sites were predicted. A total of 104 fragments were integrated into the mitochondrial genome from the chloroplast, amounting to 55,866 bp of transferred sequence. Finally, phylogenetic analysis of 28 plant mitogenomes placed P. arundinacea closest to species within the genus Poa (P. chaixii and P. pratensis). Comparative analysis of non-synonymous-to-synonymous substitution rate (Ka/Ks) ratios across divergent species revealed that the mitochondrial genome of P. arundinacea underwent stabilizing evolutionary dynamics, characterized by predominant purifying selection with several lineage-specific variations in selective pressure. Our findings support the close phylogenetic relationship between P. arundinacea and species of the genus Poa and provide a reference mitochondrial genome resource for future comparative studies within Phalaris that incorporate broader taxon sampling. These results support deeper phylogenetic investigations of P. arundinacea and facilitate future work on its germplasm characterization and applied use.

Phalaris arundinacea↗

Identification of novel simple sequence length polymorphisms (SSLPs) in mouse by interspersed repetitive element (IRE)-PCR.

Interspersed repetitive element (IRE)-PCR is a useful method for identification of novel human or mouse sequence tagged sites (STSs) from contigs of genomic clones. We describe the use of IRE-PCR with mouse B1 repetitive element primers to generate novel, PCR amplifiable, simple sequence length polymorphisms (SSLPs) from yeast artificial chromosome (YAC) clones containing regions of mouse chromosomes 13 and 14. Forty-two IRE-PCR products were cloned and sequenced from eight YACs. Of these, 29 clones contained multiple simple sequence repeat units. PCR analysis with primers derived from unique sequences flanking the simple sequence repeat units in seven clones showed all to be polymorphic between various mouse strains. This novel approach to SSLP identification represents an efficient method for saturating a genomic interval with polymorphic genetic markers that may expedite the positional cloning of genes for traits and diseases.

Animals↗

Molecular tagging of gene conferring leaf blight resistance using microsatellites in sorghum [Sorghum bicolor (L.) Moench].

Resistance to leaf blight in sorghum [Sorghum bicolor (L.) Moench] accession G-118 was found to segregate as a single dominant trait in a cross to susceptible cultivar, HC-136. Molecular marker(s) linked to the locus for disease resistance was identified using simple sequence repeat (SSR) markers coupled with bulk segregant analysis. Genomic DNA from the parental cultivars and bulks were screened by PCR amplification with 50 simple sequence repeat primer pairs. Out of these, 38 SSR primers produced polymorphism between parents. After screening of these 38 SSRs with resistant and susceptible bulk, one SSR primer, Xtxp 309 produced a unique band of approximately 700 bp only in resistant parent and resistant bulk and a unique band of 450 bp only in susceptible parent and susceptible bulk. Upon screening with individual resistant and susceptible recombinant inbred lines (RILs), marker Xtxp 309 produced amplification in 23 of the 26 resistant RILs and no amplification was produced in any of the 25 susceptible RILs. The same marker Xtxp 309 produced amplification in 21 of the susceptible RILs and 3 of the resistant RILs of 450 bp band. This was found to be located at a distance of 3.12 cM away from the locus governing resistance to leaf blight which was considered to be closely linked and 7.95 cM away from the locus governing susceptibility to leaf blight. This marker may prove useful in MAS for gene introgression, plant genetic diagnostics and gene pyramiding for resistance via genetic transformation for disease resistance in plants.

Fungi↗

Mapping of the interleukin 5 receptor gene to human chromosome 3 p25-p26 and to mouse chromosome 6 close to the Raf-1 locus with polymorphic tandem repeat sequences.

Simple-sequence tandem repeat sequences in the 3' UTR of interleukin 5 (IL5)-receptor gene of human and mouse are polymorphic in their length among humans and different strains of mice. In 20 different human Epstein-Barr virus (EBV)-transformed cell lines, six alleles of IL5R could be distinguished. In the mouse, three different alleles are found. With the human-specific IL5R tandem repeat marker in human-rodent somatic cell hybrids, the IL5R gene was mapped to human Chromosome (Chr) 3 p25-p26. With the mouse-specific IL5R tandem repeat sequence in recombinant inbred strains of mice, the Il5r gene was mapped to the distal part of mouse Chr 6 close to the Raf-1 locus.

Animals↗

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial↗

Repeat-associated phase variable genes in the complete genome sequence of Neisseria meningitidis strain MC58.

Phase variation, mediated through variation in the length of simple sequence repeats, is recognized as an important mechanism for controlling the expression of factors involved in bacterial virulence. Phase variation is associated with most of the currently recognized virulence determinants of Neisseria meningitidis. Based upon the complete genome sequence of the N. meningitidis serogroup B strain MC58, we have identified tracts of potentially unstable simple sequence repeats and their potential functional significance determined on the basis of sequence context. Of the 65 potentially phase variable genes identified, only 13 were previously recognized. Comparison with the sequences from the other two pathogenic Neisseria sequencing projects shows differences in the length of the repeats in 36 of the 65 genes identified, including 25 of those not previously known to be phase variable. Six genes that did not have differences in the length of the repeat instead had polymorphisms such that the gene would not be expected to be phase variable in at least one of the other strains. A further 12 candidates did not have homologues in either of the other two genome sequences. The large proportion of these genes that are associated with frameshifts and with differences in repeat length between the neisserial genome sequences is further corroborative evidence that they are phase variable. The number of potentially phase variable genes is substantially greater than for any other species studied to date, and would allow N. meningitidis to generate a very large repertoire of phenotypes through expression of these genes in different combinations. Novel phase variable candidates identified in the strain MC58 genome sequence include a spectrum of genes encoding glycosyltransferases, toxin related products, and metabolic activities as well as several restriction/modification and bacteriocin-related genes and a number of open reading frames (ORFs) for which the function is currently unknown. This suggests that the potential role of phase variation in mediating bacterium-host interactions is much greater than has been appreciated to date. Analysis of the distribution of homopolymeric tract lengths indicates that this species has sequence-specific mutational biases that favour the instability of sequences associated with phase variation.

Bacterial Proteins↗

Comparative analysis of seeded and vegetative biotype buffalograsses based on phylogenetic relationship using ISSRs, SSRs, RAPDs, and SRAPs.

Buffalograss [ Buchloe dactyloides (Nutt.) Englem.] is the only native grass that is being used extensively as a turfgrass in the Great Plains region. Its low-growth habit, drought resistance, and low-maintenance requirement make it attractive as a turfgrass species. Our objective was to obtain an overview on the genetic relatedness among and within seeded and vegetative biotype buffalograsses using inter-simple sequence repeats (ISSRs), random amplified polymorphic DNA (RAPDs), sequence-related amplified polymorphisms (SRAPs), and simple sequence repeats (SSRs) markers that were derived from related species (maize, pearl millet, sorghum, and sugarcane). Twenty individuals per cultivar were genotyped using 30 markers from each marker system. All buffalograss cultivars were uniquely fingerprinted by all four marker systems. Mean genetic similarities were estimated at 0.52, 0.51, 0.62, and 0.57 using SSRs, ISSRs, SRAPs, and RAPDs, respectively. Two main clusters separating the seeded-biotype from the vegetative-biotype cultivars were produced using UPGMA analysis. Further subgroupings were unequivocal. The Mantel test resulted in a very good fit (SRAP=0.92, ISSR=0.90) to good fit (RAPD=0.86, SSR=0.88) of cophenetic values. Comparing the four marker systems to each other, RAPD and SRAP similarity indices were highly correlated ( r=0.73), while Spearman's rank correlation coefficient between RAPDs and SSRs was r=0.24 and between ISSRs and SSRs was r=0.66. A genotype-assignment analytical approach might be useful for cultivar identification and property rights protection. Polymorphic SRAPs were abundant and demonstrated genetic diversity among closely related cultivars.

Analysis of Variance↗

Evidence on DNA slippage step-length distribution.

A simple model based on a master equation is constructed in order to reveal the details of the mutational events modifying simple sequence repeats in the human genome, A database of simple repeats together with their flanking sequences comprising approximately 10(5) entries from all 24 human chromosomes was constructed. By aligning the pairs of fragments of sequences containing the repeat elements, the matrices that count the number of slippage events were evaluated. The counts were then used as a target to be reproduced by our theoretical model, in which the elongation and shortening of the repeats proceed through a mechanism in which the step lengths exhibit a decaying distribution in the form of an inverse power law rather than through one nucleotide extension or deletion, which was the most frequent supposition in previous studies.

Chromosomes, Human↗

Structures of trinucleotide repeats in human transcripts and their functional implications.

Among the goals of RNA structural and functional genomics is determining structures and establishing the functions of a rich repertoire of simple sequence repeats in transcripts. These repeats are present in transcripts from their 'birth' in the nucleus to their 'death' in cytoplasm and have the potential of being involved in many steps of RNA regulation. The knowledge of their structural features and functional roles will also shed more light on the postulated mechanisms of RNA pathogenesis in a growing list of neurological diseases caused by simple sequence repeat expansions. Here, we discuss several different lines of research to support the hypothesis that the mechanism of RNA pathogenesis may be a more common phenomenon triggered or modulated also by abundant long normal repeats. We propose structures of the repeat regions in transcripts of genes involved in Triplet Repeat Expansion Diseases. We have classified the polymorphic repeat alleles of these genes according to their ability to form hairpin structures in transcripts, and describe the distribution of different structural forms of the repeats in the human population. We have also reported the results of a systematic survey of the human transcriptome to identify mRNAs containing triplet repeats and to classify them according to structural and functional criteria. Based on this knowledge, we discuss the putative wider role of triplet repeat RNA hairpins in human diseases. A hypothetical model is proposed in which long normal RNA hairpins formed by the repeats may also be involved in pathogenesis.

Alleles↗

Diploid hybrid speciation in Penstemon (Scrophulariaceae).

Hybrid speciation has played a significant role in the evolution of angiosperms at the polyploid level. However, relatively little is known about the importance of hybrid speciation at the diploid level. Two species of Penstemon have been proposed as diploid hybrid derivatives based on morphological data, artificial crossing studies, and pollinator behavior observations: Penstemon spectabilis (derived from hybridization between Penstemon centranthifolius and Penstemon grinnellii) and Penstemon clevelandii (derived from hybridization between P. centranthifolius and P. spectabilis). Previous studies were inconclusive regarding the purported hybrid nature of these species because of a lack of molecular markers sufficient to differentiate the parental taxa in the hybrid complex. We developed hypervariable nuclear markers using inter-simple sequence repeat banding patterns to test these classic hypotheses of diploid hybrid speciation in Penstemon. Each species in the hybrid complex was genetically distinct, separated by 10-42 species-specific inter-simple sequence repeat markers. Our data do not support the hybrid origin of P. spectabilis but clearly support the diploid hybrid origin of P. clevelandii. Our results further suggest that the primary reason diploid hybrid speciation is so difficult to detect is the lack of molecular markers able to differentiate parental taxa from one another, particularly with recently diverged species.

Biological Evolution↗

Assessment of genetic polymorphisms in DNA from formalin fixed neurological tissues.

The ability to analyze the genotype of deceased affected members of pedigrees segregating inherited neurological diseases considerably augments the informativeness of such pedigrees. This information has direct application in attempts to isolate disease genes by positional cloning strategies, and for genetic counselling. We show that the genotype at polymorphic simple sequence repeat loci can be determined from genomic DNA isolated from 10 micron thick paraffin embedded, formalin fixed neurological tissues. The critical constraint on this method is the size of the template target bearing the simple sequence repeat, which should ideally be less than 165 base pairs.

Alzheimer Disease↗

Human von Willebrand factor gene and pseudogene: structural analysis and differentiation by polymerase chain reaction.

Structural analysis of the von Willebrand factor gene located on chromosome 12 is complicated by the presence of a partial unprocessed pseudogene on chromosome 22q11-13. The structures of the von Willebrand factor pseudogene and corresponding segment of the gene were determined, and methods were developed for the rapid differentiation of von Willebrand factor gene and pseudogene sequences. The pseudogene is 21-29 kilobases in length and corresponds to 12 exons (exons 23-34) of the von Willebrand factor gene. Approximately 21 kilobases of the gene and pseudogene were sequenced, including the 5' boundary of the pseudogene. The 3' boundary of the pseudogene lies within an 8-kb region corresponding to intron 34 of the gene. The presence of splice site and nonsense mutations suggests that the pseudogene cannot yield functional transcripts. The pseudogene has diverged approximately 3.1% in nucleotide sequence from the gene. This suggests a recent evolutionary origin approximately 19-29 million years ago, near the time of divergence of humans and apes from monkeys. Several repetitive sequences were identified, including 4 Alu, one Line-1, and several short simple sequence repeats. Several of these simple repeats differ in length between the gene and pseudogene and provide useful markers for distinguishing these loci. Sequence differences between the gene and pseudogene were exploited to design oligonucleotide primers for use in the polymerase chain reaction to selectivity amplify sequences corresponding to exons 23-34 from either the von Willebrand factor gene or the pseudogene. This method is useful for the analysis of gene defects in patients with von Willebrand disease, without interference from homologous sequences in the pseudogene.

Amino Acid Sequence↗

A genetic linkage map for tef [Eragrostis tef (Zucc.) Trotter].

Tef [Eragrostis tef (Zucc.) Trotter] is the major cereal crop in Ethiopia. Tef is an allotetraploid with a base chromosome number of 10 (2n = 4x = 40) and a genome size of 730 Mbp. Ninety-four F(9) recombinant inbred lines (RIL) derived from the interspecific cross, Eragrostis tef cv. Kaye Murri x Eragrostis pilosa (accession 30-5), were mapped using restriction fragment length polymorphisms (RFLP), simple sequence repeats derived from expressed sequence tags (EST-SSR), single nucleotide polymorphism/insertion and deletion (SNP/INDEL), intron fragment length polymorphism (IFLP) and inter-simple sequence repeat amplification (ISSR). A total of 156 loci from 121 markers was grouped into 21 linkage groups at LOD 4, and the map covered 2,081.5 cM with a mean density of 12.3 cM per locus. Three putative homoeologous groups were identified based on multi-locus markers. Sixteen percent of the loci deviated from normal segregation with a predominance of E. tef alleles, and a majority of the distorted loci were clustered on three linkage groups. This map will be useful for further genetic studies in tef including mapping of loci controlling quantitative traits (QTL), and comparative analysis with other cereal crops.

Chromosome Mapping↗