PubMed HealthSearch

SEARCH · PubMed Health

Results for “Tandem Repeat Sequences”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Characterization of a porcine variable number tandem repeat sequence specific for the glucosephosphate isomerase locus.

A variable number of tandem repeat from a porcine glucosephosphate isomerase intron has been isolated and sequenced. The repeat has a unit size of 39 bp, is highly conserved and is present in at least 14 copies. Flanking sequences show a sequence periodicity of 53-54 bp and some sequence homology to the 39 bp repeat. A considerable part of the genomic DNA has been lost during subcloning and is considered to be deletion prone or refractory to propagation in E. coli. The tandem repeat is locus specific and detects at least six alleles in BamHI digested porcine DNA. No homology to other tandem repeat sequences has been found.

Animals

Intratubular germ cell neoplasia in infantile yolk sac tumor. Verification by tandem repeat sequence in situ hybridization.

The strong association of intratubular germ cell neoplasia (ITGCN) with adult germ cell testicular tumors is well known, but studies noting the absence of ITGCN in certain germ cell neoplasms such as spermatocytic seminoma, childhood teratoma, and infantile yolk sac tumor (YST) have raised the issue of whether these latter neoplasms follow a different path of tumorigenesis, accounting for their more benign behavior. A case study illustrating the association of ITGCN with infantile YST is presented to challenge this hypothesis. In addition to the usual characteristic features that included strong cytoplasmic glycogen deposits, and focal placental alkaline phosphatase immunoreactivity, the atypical intratubular germ cells manifested triploidy by in situ hybridization using as probe a telomeric tandem repeat sequence, p1-79, specific to chromosome 1. The invasive YST cells, in contrast, showed evidence of tetraploidy by both in situ hybridization and flow and image cytometric studies, excluding the possibility that the atypical intratubular germ cells represented intratubular invasion by adjacent YST. These findings challenge the belief that the infantile YST follows a different path of tumorigenesis than its adult germ cell counterpart and suggest other hypotheses that might better explain its more benign behavior.

Blotting, Southern

von Willebrand disease family studies: comparison of three methods of analysis of the von Willebrand factor gene polymorphism related to a variable number tandem repeat sequence in intron 40.

A region with a variable number of tandem ATCT repeats (VNTR) has previously been localized within intron 40 of the von Willebrand factor (vWF) gene. In the present report we describe the use of this polymorphism as a genetic marker to study the inheritance pattern in five families affected with various types of von Willebrand disease (vWD): types I, IIA, IIB, IIC and the newly characterized variant with totally defective FVIII binding. Three means of investigation previously reported, all using polymerase chain reaction (PCR) amplification of this vWF gene region, were compared in terms of informativeness. The two direct single-step procedures analysing only partial sequences of the VNTR region turned out to be less informative (three studies informative out of five) than the third method characterizing the variability of the whole VNTR sequence. This latter approach, based on the analysis of the Alu I restriction pattern of the VNTR region, was informative in all the families investigated, therefore avoiding the need to combine it with other genetic marker studies for efficient gene tracking. In conclusion, this two-step (PCR and digestion) method is the most informative for the characterization of the inheritance of the different subtypes of vWD and for the prenatal diagnosis of its severe forms.

Alleles

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

ICP22 homolog of equine herpesvirus 1: expression from early and late promoters.

The complete nucleotide sequence of the short region, made up of a unique segment (Us; 6.5 kb) bracketed by a pair of inverted repeat sequences (IR; 12.8 kb each), of the equine herpesvirus 1 (EHV-1) genome has been determined recently in our laboratory. Analysis of the IR segment revealed a major open reading frame (ORF) designated IR4. The IR4 ORF exhibits significant homology to the immediate-early gene US1 (ICP22) of herpes simplex virus type 1 and to the ICP22 homologs of varicella-zoster virus (ORF63), pseudorabies virus (RSp40), and equine herpesvirus 4 (ORF4). The IR4 ORF is located entirely within each of the inverted repeat sequences (nucleotides [nt] 7918 to 9327) and has the potential to encode a polypeptide of 469 amino acids (49,890 Da). Within the IR4 ORF are two reiterated sequences: a 7-nt sequence tandemly repeated 17 times and a 25-nt sequence tandemly repeated 13 times. Nucleotide sequence analyses of IR4 also revealed several potential cis-regulatory sequences, two TATA sequences separated by 287 nt, an in-frame translation initiation codon following each TATA sequence, and a single polyadenylation site. To address the nature of the mRNA species encoded by IR4, we used Northern (RNA) blot and S1 nuclease analyses. RNA mapping data revealed that IR4 has two promoters that are regulated differentially during a lytic infection. A 1.4-kb mRNA appears initially at 2 h postinfection and is an early transcript since its synthesis is not affected by the presence of phosphonoacetic acid, an inhibitor of EHV-1 DNA replication. In contrast, a 1.7-kb mRNA appears at later times postinfection and is designated as a gamma-1 transcript, since its synthesis is significantly reduced by phosphonoacetic acid. These IR4-specific mRNAs are 3' coterminal, have unique 5' termini, and would code for in-frame, overlapping, carboxy-coterminal proteins of 293 and 469 amino acids, respectively. Interestingly, the site of homologous recombination to generate the genome of EHV-1 defective interfering particles that initiate persistent infection occurs between nt 3244 and 3251 of UL3 (ICP27 homolog) and nt 9027 and 9034 of IR4 (ICP22 homolog). Thus, this recombination event would generate a unique ORF that would encode a potential protein whose amino end was derived from the N-terminal 193 amino acids of the ICP22 homolog and whose carboxyl end was derived from the C-terminal 68 amino acids of the ICP27 homolog.

Amino Acid Sequence

P-element homologous sequences are tandemly repeated in the genome of Drosophila guanche.

In Drosophila guanche, P-homologous sequences were found to be located in a tandem repetitive array (copy number: 20-50) at a single genomic site. The cytological position on the polytene chromosomes was determined by in situ hybridization (chromosome O: 85C). Sequencing of one complete repeat unit (3.25 kilobases) revealed high sequence similarity between the central coding region comprising exons 0 to 2 and the corresponding section of the Drosophila melanogaster P element. The rest of the sequence has diverged considerably. Exon 3 has no coding function and the inverted repeats have disappeared. The P homologues of D. guanche apparently have lost their mobility but have retained the coding capacity for a protein similar to the 66-kDa P-element repressor of D. melanogaster. Divergence between different repeat units indicates early amplification of the sequence at this particular genomic site. The presence of a common P-element site at 85C in Drosophila subobscura, Drosophila madeirensis, and D. guanche suggests that clustering of the sequence at this location took place before the phylogenetic radiation of the three species.

Amino Acid Sequence

Genetic mapping of tandemly repeated telomeric DNA sequences in tomato (Lycopersicon esculentum).

A telomere-associated tandemly repeated DNA sequence of tomato, TGR I, has been used to map telomeres on the tomato RFLP linkage map. Mapping was performed by monitoring the segregation of entire arrays of TGR I from a segregating F2 population using pulsed-field gel electrophoresis (PFGE). With this strategy, four telomeres have been mapped to the ends of the short arm of chromosomes 9 and 12 and the long arms of chromosomes 5 and 11, using a saturated RFLP map of tomato containing approximately 1000 RFLP markers. In all four cases, the TGR I locus maps to the end of the chromosome, and the distance between the most distal single-copy RFLP marker and the telomeric TGR I locus was between 1.6 and 9.6 cM. This indicates that the region close to the telomeres does not show an excessive rate of recombination compared to other regions of the genome and that the RFLP map of tomato is essentially complete and covers the entire genome for all practical purposes. Additionally, the mapping technique presented here should be generally applicable to the mapping of other tandemly repeated DNA sequences.

Chromosome Mapping

Highly repeated DNA sequences in birds: the structure and evolution of an abundant, tandemly repeated 190-bp DNA fragment in parrots.

Up to 6.8% of the parrot (Psittaciformes) genome consists of a tandemly repeated, 190-bp sequence (P1) located in the centromere of many if not all chromosomes. Monomer repeats from 10 different psittacine species representing four subfamilies were isolated and cloned. The intraspecific sequence variation ranged from 1.5 to 7%. The interspecific sequence variation ranged from less than 3% between two species of cockatoos to approximately 45% between cockatoos and other parrots. The monomer sequences of all 10 parrot species contained several conserved (> 90%) sequence elements at identical locations within the repeat. A comparison with tandemly repeated DNA sequences in other avian species showed that several of these conserved elements were also present at similar locations within the 184-bp repeat of the Chilean flamingo (Phoenicopterus chilensis), suggesting a great antiquity of the repeat. One of the elements was also found in the tandemly repeated sequences of the crane (Gruidae) and falcon (Falconidae) families. The data were used for the construction of a partial most parsimonious relationship that supports a regional subdivision of the Psittaciformes.

Animals

Feasibility of prenatal diagnosis of beta-thalassemia using two highly polymorphic microsatellites 5' to the beta-globin gene.

Short tandem repeats (STRs) are highly informative loci within the human genome, consisting of short nucleotide sequences tandemly repeated in variable numbers. This results in different alleles of variable length. Herein we describe two STRS located 5' to the beta-globin gene. They can be detected by non radioactive methods and may be used to make prenatal diagnosis of beta-thalassemia.

Base Sequence

A novel relationship between time offsets in capillary electrophoresis and DNA sequence variations in short tandem repeats.

Next-generation sequencing (NGS) provides increased discriminatory power in forensic DNA analysis due to the detection of isoalleles. Differences in sequences between alleles allow for a second layer of differentiation between DNA contributors beyond the number of short tandem repeat (STR) repeat units. However, because NGS is a more time and resource-intensive analysis than conventional capillary electrophoresis (CE), laboratories may benefit from indicators that suggest NGS is likely to provide added value. This study examined whether CE migration offsets, measured as residuals in the OSIRIS analysis software, can differ significantly among STR isoalleles. Residuals represent the time offset between a sample allele peak and its corresponding allelic ladder peak. Paired CE and NGS data from 95 single source samples were analyzed for CE-based residual differences, as the NGS data provided the sequence information of the corresponding isoalleles. Residual values differed significantly among isoalleles at several STR loci. Statistically significant differences were identified at D16S539 and D3S1358, as well as at specific allele lengths within D12S391, D13S317, and D8S1179. These findings demonstrate that CE residual variation can reflect underlying STR sequence differences between contributors. In practice, residual-based metrics could help laboratories to identify casework reference samples where NGS is likely to provide additional discrimination, without the need for processing outside of a routine CE workflow. Due to the potentially large number of isoalleles, community wide efforts to aggregate CE residual differences versus isoallele sequences may be useful in the validation and implementation of this approach to add value to forensic DNA analyses.

Electrophoresis, Capillary

Sequence studies on mouse L-cell satellite DNA by base-specific degradation with T4 endonuclease IV.

The base sequence of mouse L-cell satellite DNA was investigated by degradation of the two separated complementary strands with the base specific enzyme, T4 endonuclease IV. Digestion of the heavy strand DNA released a limited number of oligonucleotides which were separated by ionophoresis/homochromatography, isolated, and sequenced by the 'wandering spot' method. The light strand DNA was resistant to digestion with T4 endonuclease IV and no detectable amounts of oligonucleotides were released. The oligonucleotides obtained from the heavy strand were related in sequence, indicating that mouse satellite DNA derived from a short tandemly repeated sequence. The sequence of part of the original repeat unit is proposed to be C-A-T-T-T-T-T-C. Five major oligonucleotides were identified, all of which differ from the proposed original sequence by single base changes. The five major oligonucleotides occur with about equal frequency and together comprise approximately 50% of the oligonucleotides released by T4 endonuclease IV from the heavy strand DNA. In addition to the five major oligonucleotides, several oligonucleotides were found to occur in lesser amounts. Since these oligonucleotides are related to the major oligonucleotides, it is likely that they have arisen from them by mutation.

Base Sequence

The nucleotide sequence of the variable region in Trypanosoma brucei completes the sequence analysis of the maxicircle component of mitochondrial kinetoplast DNA.

The nucleotide sequence of two non-contiguous DNA fragments of 4.0 and 2.2 kb, respectively, of the kinetoplast maxicircle of Trypanosoma brucei brucei EATRO strain 427 has been determined, completing the sequence analysis of the so-called variable region (see also de Vries et al., 1988, Mol. Biochem. Parasitol. 27, 71-82). Analysis of the entire 8-kb variable region sequence revealed the presence of a 5.2-kb cluster of imperfect, tandemly repeated sequences, flanked by DNA of unique sequence. Both repetitive and unique DNA evolve rapidly, but comparison to the closely related strain EATRO 164 indicated that the repetitive cluster is more prone to sequence and size divergence. The variable region is transcribed into RNAs of varying lengths but appears to be devoid of genes encoding mitochondrial proteins or tRNAs, as judged from computer analysis. Moreover, genes that could encode guide RNAs involved in producing the known edited mitochondrial mRNA sequences are also absent. The repetitive DNA cluster within this region consists of 14 blocks each containing one 130 bp repeat and a variable number of 19 bp repeats. A duplicated sequence was identified (5'-GGGGTTGGTGT) which proved to be identical to the eleven 5'-terminal residues of the universal minicircle dodecamer involved in initiation of leading strand synthesis. This suggests a role for these sequences in the initiation of maxicircle DNA replication. With the data presented in this report, the nucleotide sequence analysis of the 23016 bp maxicircle of T. brucei brucei EATRO strain 427 has been completed.

Animals

Regulation of transcription of the germ-line Ig alpha constant region gene by an ATF element and by novel transforming growth factor-beta 1-responsive elements.

Inasmuch as transcription of unrearranged, or germ-line, Ig CH genes appears to direct switch recombination, understanding the regulation of this transcription is essential for understanding the regulation of class switching. Transforming growth factor-beta 1 (TGF-beta 1) induces germ-line alpha transcripts and increases class switching to IgA in the I.29 mu B lymphoma and in Peyer's patch and splenic B cells. It has been previously demonstrated that induction of germ-line alpha transcripts by TGF-beta occurs at the transcriptional level in I.29 mu cells. We now demonstrate that the DNA segment located 5' to the initiation sites of germ-line alpha RNA drives expression of a luciferase reporter gene construct in transient transfection experiments. Full constitutive expression requires no more than 106 bp of the 5' flanking segment. By creating a series of deletion and substitution mutations, we have demonstrated that an ATF/CRE site residing within this region is very important for constitutive expression of the germ-line alpha promoter, but mutation of this motif does not diminish TGF-beta induction. Inducibility by TGF-beta requires additional sequences residing between -128 to -106 relative to the first RNA initiation site. Two copies of a tandemly repeated sequence 5' CA-CAG(G)CCAGAC 3' (termed Ig alpha TGF-beta-RE) are located in the region from -127 to -105. An oligonucleotide containing multimers of these repeats confers TGF-beta inducibility to a heterologous promoter. An additional copy of the TGF-beta-RE was identified at -41/-30 and its deletion reduces the TGF-beta response. Thus, we conclude that tandem repeats of a novel TGF-beta-RE are the positive regulatory element for the TGF-beta response. Our study provides further evidence that TGF-beta directs class switching to IgA through induction of transcription of the germ-line C alpha gene and demonstrates that TGF-beta can activate the promoter for the germ-line alpha gene.

Activating Transcription Factors

Synthesis of hybrid bacterial plasmids containing highly repeated satellite DNA.

Hybrid plasmid molecules containing tandemly repeated Drosophila satellite DNA were constructed using a modification of the (dA)-(dT) homopolymer procedure of Lobban and Kaiser (1973). Recombinant plasmids recovered after transformation of recA bacteria contained 10% of the amount of satellite DNA present in the transforming molecules. The cloned plasmids were not homogenous in size. Recombinant plasmids isolated from a single colony contained populations of circular molecules which varied both in the length of the satellite region and in the poly(dA)-(dt) regions linking satellite and vector. While subcloning reduced the heterogeneity of these plasmid populations, continued cell growth caused further variations in the size of the repeated regions. Two different simple sequence satellites of Drosophila melanogaster (1.672 and 1.705 g/cm3) were unstable in both recA and recBC hosts and in both pSC101 and pCR1 vectors. We propose that this recA-independent instability of tandemly repeated sequences is due to unequal intramolecular recombination events in replicating DNA molecules, a mechanism analogous to sister chromatid exchange in eucaryotes.

DNA

The genome of human papovavirus BKV.

The complete DNA sequence of human papovavirus BKV(Dun), consisting of 5153 nucleotide pairs, is presented. We describe the segments of the genome which correspond to the replication origin, the tandem repeated sequences, the 5' and 3' ends of the mRNAs, the splice sites, the early and late viral proteins and the putative viral polypeptides. These BKV DNA sequences are compared with analogous regions in the SV40 and Py virus genomes in an attempt to localize viral functions for lytic growth and transformation.

Amino Acid Sequence

High level expression in E. coli of an alternate reading frame of pS2 mRNA that encodes a mimotope of human breast epithelial mucin tandem repeat.

The high molecular weight mucin found in human milk fat globule and on the surface of mammary and other epithelial cells contains a 20 amino acid tandem repeat sequence that is highly immunogenic. We have immunoscreened lambda gt11 cDNA expression libraries from MCF7 cells and lactating breast tissue with 5 anti-mucin monoclonal antibodies. We isolated a group of cDNA clones that had the repeat sequence (HB11-2, HB11-6, HB11-10) and a group that had little or no homology with the repeat sequence (NP4, NP5, HB11-4). A fusion protein produced by NP4 bound preferentially BrE2, while HB11-4 bound only BrE2 and BrE3, NP5 produced a fusion protein that bound only Mc5 and not the other MAbs. Sequencing of the NP5 cDNA revealed it to be distinct from the mucin sequence and instead to have 97% identity with the estrogen induced transcript, pS2. An alternate reading frame was translated by the lambda gt11 fusion gene yielding a 44 amino acid protein having no homology with pS2 protein. Only a short region of homology (5 amino acids) with the breast mucin tandem repeat was found which was shown to be a mimotope for the Mc5 epitope on the breast mucin. High level expression of the NP5 cDNA was achieved by subcloning it into pEX2. The NP5 fusion protein has been useful for developing an assay for the presence of mucin derived antigen in patient serum.

Amino Acid Sequence

Sequence analysis of adenovirus DNA: complete nucleotide sequence of the spliced 5' noncoding region of adenovirus 2 hexon messenger RNA.

The complete nucleotide sequence of the 5' noncoding region of the adenovirus 2 hexon messenger RNA has been established by sequence analysis of reverse transcripts. Such transcripts were generated by extension of specific single-stranded DNA primers with reverse transcriptase after hybridization to purified hexon mRNA. The total length of the 5' noncoding region was determined to be 240 nucleotides, of which the spliced tripartite leader sequence contributes 202 nucleotides including the terminal m7G. The sizes of the different segments of the tripartite leader were estimated by comparing the established mRNA sequence with the genomic sequences for the first and third leader segments, and were found to be 42 nucleotides for the first segment, 71 nucleotides for the second and 89 nucleotides for the third. The estimates are ambiguous, however, due to the presence of tandemly repeated sequences at both ends of the intervening sequence between the third leader segment and the body of the hexon mRNA. The sequence of the leader allows the formation of hydrogen-bonded interactions with the 3' end of 18S ribosomal RNA near the capped 5' end and also close to the initiator AUG.

Adenoviruses, Human