PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “insertion sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Pathogenicity islands of virulent bacteria: structure, function and impact on microbial evolution.

Virulence genes of pathogenic bacteria, which code for toxins, adhesins, invasins or other virulence factors, may be located on transmissible genetic elements such as transposons, plasmids or bacteriophages. In addition, such genes may be part of particular regions on the bacterial chromosomes, termed 'pathogenicity islands' (Pais). Pathogenicity islands are found in Gram-negative as well as in Gram-positive bacteria. They are present in the genome of pathogenic strains of a given species but absent or only rarely present in those of non-pathogenic variants of the same or related species. They comprise large DNA regions (up to 200 kb of DNA) and often carry more than one virulence gene, the G + C contents of which often differ from those of the remaining bacterial genome. In most cases, Pais are flanked by specific DNA sequences, such as direct repeats or insertion sequence (IS) elements. In addition, Pais of certain bacteria (e,g. uropathogenic Escherichia coli, Yersinia spp., Helicobacter pylori) have the tendency to delete with high frequencies or may undergo duplications and amplifications. Pais are often associated with tRNA loci, which may represent target sites for the chromosomal integration of these elements. Bacteriophage attachment sites and cryptic genes on Pais, which are homologous to phage integrase genes, plasmid origins of replication of IS elements, indicate that these particular genetic elements were previously able to spread among bacterial populations by horizontal gene transfer, a process known to contribute to microbial evolution.

Biological Evolution↗

Evaluating and improving cDNA sequence quality with cQC.

SUMMARY: Errors are prevalent in cDNA sequences but the extent to which sequence collections differ in frequencies and types of errors has not been investigated systematically. cDNA quality control, or cQC, was developed to evaluate the quality of cDNA sequence collections and to revise those sequences that differ from a higher quality genomic sequence. After removing rRNA, vector, bacterial insertion sequence and chimeric cDNA contaminants, small-scale nucleotide discrepancies were found in 51% of cDNA sequences from one Arabidopsis cDNA collection, 89% from a second Arabidopsis collection and 75% from a rice collection. These errors created premature termination codons in 4 and 42% of cDNA sequences in the respective Arabidopsis collections and in 7% of the rice cDNA sequences.

5' Untranslated Regions↗

Genome dynamics and diversity of Shigella species, the etiologic agents of bacillary dysentery.

The Shigella bacteria cause bacillary dysentery, which remains a significant threat to public health. The genus status and species classification appear no longer valid, as compelling evidence indicates that Shigella, as well as enteroinvasive Escherichia coli, are derived from multiple origins of E.coli and form a single pathovar. Nevertheless, Shigella dysenteriae serotype 1 causes deadly epidemics but Shigella boydii is restricted to the Indian subcontinent, while Shigella flexneri and Shigella sonnei are prevalent in developing and developed countries respectively. To begin to explain these distinctive epidemiological and pathological features at the genome level, we have carried out comparative genomics on four representative strains. Each of the Shigella genomes includes a virulence plasmid that encodes conserved primary virulence determinants. The Shigella chromosomes share most of their genes with that of E.coli K12 strain MG1655, but each has over 200 pseudogenes, 300 approximately 700 copies of insertion sequence (IS) elements, and numerous deletions, insertions, translocations and inversions. There is extensive diversity of putative virulence genes, mostly acquired via bacteriophage-mediated lateral gene transfer. Hence, via convergent evolution involving gain and loss of functions, through bacteriophage-mediated gene acquisition, IS-mediated DNA rearrangements and formation of pseudogenes, the Shigella spp. became highly specific human pathogens with variable epidemiological and pathological features.

DNA Transposable Elements↗

Genetic rearrangements of the regions adjacent to genes encoding heat-labile enterotoxins (eltAB) of enterotoxigenic Escherichia coli strains.

One of the most common bacterially mediated diarrheal infections is caused by enterotoxigenic Escherichia coli (ETEC) strains. ETEC-derived plasmids are responsible for the distribution of the genes encoding the main toxins, namely, the heat-labile and heat-stable enterotoxins. The origins and transfer modes (intra- or interplasmid) of the toxin-encoding genes have not been characterized in detail. In this study, we investigated the DNA regions located near the heat-labile enterotoxin-encoding genes (eltAB) of several clinical isolates. It was found that the eltAB region is flanked by conserved 236- and 280-bp regions, followed by highly variable DNA sequences which consist mainly of partial insertion sequence (IS) elements. Furthermore, we demonstrated that rearrangements of the eltAB region of one particular isolate, which harbors an IS91R sequence next to eltAB, could be produced by a recA-independent but IS91 sequence-dependent mechanism. Possible mechanisms of dissemination of IS element-associated enterotoxin-encoding genes are discussed.

Bacterial Toxins↗

Construction of an IS946-based composite transposon in Lactococcus lactis subsp. lactis.

An artificial composite transposon was constructed based on the lactococcal insertion sequence IS946. A 3.0-kb element composed of the pC194 cat gene (Cmr) flanked by inversely repeated copies of IS946 was assembled on pBluescript KS+. When subcloned into the shuttle vector pSA3 (Emr), two putative transposons were created on the recombinant plasmid pTRK128: the 3.0-kb Cmr element (Tn-CmA) and an inverse 11.5-kb Emr element (Tn-EmA). pTRK128 was electroporated into the recombination-deficient strain Lactococcus lactis MMS362, which contains the self-transmissible plasmid pRS01. An MMS362 Cmr Emr transformant was used to assay for transposition events via conjugal mobilization of pTRK128-encoded Cmr or Emr to L. lactis LM2345. Transfer of either marker alone occurred at frequencies of ca. 2 x 10(-4) per input donor. Approximately 19% of the Emr transconjugants were Cms, indicating loss of the cat gene marker. No Cmr Ems transconjugants were recovered (n = 550). Plasmid analysis showed that the Cms Emr isolates contained a single large plasmid that was determined to be a cointegrate between pRS01 and the Tn-EmA element. A 32P-labeled pSA3 probe hybridized specifically to pTRK128 sequences and revealed different junction fragments within each of the cointegrate plasmids. DNA sequence analysis of the Tn-EmA::pRS01 junctions from a representative cointegrate verified transposition by Tn-EmA. This represents the first example of a functional composite transposon in the genus Lactococcus and serves as an experimental tool and model for the genetic analyses of transposons in these organisms.

Base Sequence↗

The R region found in the human foamy virus long terminal repeat is critical for both Gag and Pol protein expression.

It has been suggested that sequences located within the 5' noncoding region of human foamy virus (HFV) are critical for expression of the viral Gag and Pol structural proteins. Here, we identify a discrete approximately 151-nucleotide sequence, located within the R region of the HFV long terminal repeat, that activates HFV Gag and Pol expression when present in the 5' noncoding region but that is inactive when inverted or when placed in the 3' noncoding region. Sequences that are critical for the expression of both Gag and Pol include not only the 5' splice site positioned at +51 in the R region, which is used to generate the spliced pol mRNA, but also intronic R sequences located well 3' to this splice site. Analysis of total cellular gag and pol mRNA expression demonstrates that deletion of the R region has little effect on gag mRNA levels but that R deletions that would be predicted to leave the pol 5' splice site intact nevertheless inhibit the production of the spliced pol mRNA. Gag expression can be largely rescued by the introduction of an intron into the 5' noncoding sequence in place of the R region but not by an intron or any one of several distinct retroviral nuclear RNA export sequences inserted into the mRNA 3' noncoding sequence. Neither the R element nor the introduced 5' intron markedly affects the cytoplasmic level of HFV gag mRNA. The poor translational utilization of these cytoplasmic mRNAs when the R region is not present in cis also extended to a cat indicator gene linked to an internal ribosome entry site introduced into the 3' noncoding region. Together these data imply that the HFV R region acts in the nucleus to modify the cytoplasmic fate of target HFV mRNA. The close similarity between the role of the HFV R region revealed in this study and previous data (M. Butsch, S. Hull, Y. Wang, T. M. Roberts, and K. Boris-Lawrie, J. Virol. 73:4847--4855, 1999) demonstrating a critical role for the R region in activating gene expression in the unrelated retrovirus spleen necrosis virus suggests that several distinct retrovirus families may utilize a common yet novel mechanism for the posttranscriptional activation of viral structural protein expression.

Gene Expression Regulation, Viral↗

Sequence analysis of the porcine IFNAR1 and IFNGR2 genes.

A porcine BAC clone harboring the tightly linked IFNAR1 and IFNGR2 genes was identified by comparative analysis of the publicly available porcine BAC end sequences. The complete 168,835 bp insert sequence of this clone was determined. Sequence comparisons of the genomic sequence with EST sequences from public databases were performed and allowed a detailed annotation of the IFNAR1 and IFNGR2 genes. The analyzed genes showed a conserved genomic organization with their known mammalian orthologs, however the sequence conservation of these genes across species was relatively low. In addition to the IFNAR1 and IFNGR2 genes, which were completely sequenced, the analyzed BAC clone also contained parts of an orphan gene encoding a putative transmembrane protein (TMEM50B). In contrast to the IFNAR1 and IFNGR2 genes the sequence conservation of the TMEM50B gene across different mammalian species was extremely high.

Amino Acid Sequence↗

Roles of CCAAT/enhancer-binding protein and its binding site on repression and derepression of acetyl-CoA carboxylase gene.

The gene for acetyl-CoA carboxylase, the rate-limiting enzyme in the biosynthesis of long-chain fatty acids, contains two distinct promoter regions, denoted PI and PII, which control the generation of different forms of mRNA. Multiple forms of acetyl-CoA carboxylase (ACC) mRNA with 5'-end heterogeneity are generated as a result of differential splicing of two primary transcripts formed under the control of these two promoters. PI is responsible for the generation of class I mRNAs of ACC, which are induced in a tissue-specific manner under lipogenic conditions. PII generates class II mRNAs of ACC, which are expressed constitutively. Possible mechanisms for the regulation of PI under normal physiological conditions and agents that activate the promoter have been investigated. PI contains a TATA and a CCAAT box. In addition to these sequences, this promoter contains a 28-CA repeat sequence 220 bases upstream from the transcription initiation site; the presence of this sequence leads to about 70% repression of the basal promoter activity. Repression by the 28-CA repeat sequence requires the GCAAT sequence in the CCAAT box. The negative effect of the 28-CA repeat sequence is relieved by a CCAAT/enhancer-binding protein (C/EBP), which binds to the GCAAT sequence. Insertion of the 28-CA repeat sequence into the thymidine kinase promoter results in repression that can also be relieved by the C/EBP gene product. However, the same sequence exerts no effect on ACC promoter II, which has no CCAAT box. During the differentiation of 30A5 preadipocytes into adipocytes, the expression of class I ACC mRNA and C/EBP mRNA is coordinately increased. Therefore, the presence of the CA repeat in the promoter may be responsible for the inactivity of PI, and C/EBP may be one of the factors that is responsible for the activation of PI under lipogenic conditions. Interaction of the CA repeat and the CCAAT box in the repression and derepression of the ACC gene provides a novel function for the CCAAT box and C/EBP in gene regulation.

Acetyl-CoA Carboxylase↗

Genomic rearrangements correlated with antigenic variation in Trypanosoma brucei.

The capacity of African trypanosomes to express sequentially a large repertoire of different surface antigens during an infection enables the parasite to evade the immune response of its host, and makes attempts to produce a vaccine against the disease difficult. It is evident that point mutations cannot account for antigen diversity. Variable antigens like immunoglobulins are derived from an extensive family of genes of which only one is expressed in a given cell. As somatic tic recombination is involved in the immunoglobulin gene system, this similarity prompted us to search for somatic rearrangements in trypanosome variable antigen genes. We have constructed a recombinant plasmid containing approximately half the DNA sequence coding for a Trypanosoma brucei variable antigen and hybridised the inserted sequences to various restriction enzyme digests of nuclear DNA from different trypanosome clones. Differences in the sizes of restriction tion fragments hybridising to the inserted variable antigen coding sequence show altered positions of enzyme sites relative to this sequence, indicating different arrangements of DNA sequences around this gene in different trypanosome clones.

Animals↗

Promiscuous mitochondrial group II intron sequences in plant nuclear genomes.

Gene translocations from the organelles to the nucleus are postulated by the endosymbiont hypothesis. We here report evidence for sequence insertions in the nuclear genomes of plants that are derived from noncoding regions of the mitochondrial genome. Fragments of mitochondrial group II introns are identified in the nuclear genomes of tobacco and a bean species. The duplicated intron sequences of 75-140 bp are derived from cis- and trans-splicing introns of genes encoding subunits 1 and 5 of the NADH dehydrogenase. The mitochondrial sequences are inserted in the vicinities of a lectin gene, different glucanase genes and a gene encoding a subunit of photosystem II. Sequence similarities between the nuclear and mitochondrial copies are in the range of 80 to 97%, suggesting recent transfer events that occurred in the basic glucanase genes before and in the lectin gene after the gene duplications in the evolution of the nuclear gene families. Overlapping regions of the same introns are in two instances also involved in intramitochondrial sequence duplications.

Base Sequence↗

Transposable element ISHp608 of Helicobacter pylori: nonrandom geographic distribution, functional organization, and insertion specificity.

A new member of the IS605 transposable element family, designated ISHp608, was found by subtractive hybridization in Helicobacter pylori. Like the three other insertion sequences (ISs) known in this gastric pathogen, it contains two open reading frames (orfA and orfB), each related to putative transposase genes of simpler (one-gene) elements in other prokaryotes; orfB is also related to the Salmonella virulence gene gipA. PCR and hybridization tests showed that ISHp608 is nonrandomly distributed geographically: it was found in 21% of 194 European and African strains, 14% of 175 Bengali strains, 43% of 131 strains from native Peruvians and Alaska natives, but just 1% of 223 East Asian strains. ISHp608 also seemed more abundant in Peruvian gastric cancer strains than gastritis strains (9 of 14 versus 15 of 45, respectively; P = 0.04). Two ISHp608 types differing by approximately 11% in DNA sequence were identified: one was widely distributed geographically, and the other was found only in Peruvian and Alaskan strains. Isolates of a given type differed by < or = 2% in DNA sequence, but several recombinant elements were also found. ISHp608 marked with a resistance gene was found to (i) transpose in Escherichia coli; (ii) generate simple insertions during transposition, not cointegrates; (iii) insert downstream of the motif 5"-TTAC without duplicating target sequences; and (iv) require orfA but not orfB for its transposition. ISHp608 represents a widespread family of novel chimeric mobile DNA elements whose further analysis should provide new insights into transposition mechanisms and into microbial population genetic structure and genome evolution.

Amino Acid Sequence↗

Gene trap mutagenesis of hnRNP A2/B1: a cryptic 3' splice site in the neomycin resistance gene allows continued expression of the disrupted cellular gene.

BACKGROUND: Tagged sequence mutagenesis is a process for constructing libraries of sequenced insertion mutations in embryonic stem cells that can be transmitted into the mouse germline. To better predict the functional consequences of gene entrapment on cellular gene expression, the present study characterized the effects of a U3Neo gene trap retrovirus inserted into an intron of the hnRNP A2/B1 gene. The mutation was selected for analysis because it occurred in a highly expressed gene and yet did not produce obvious phenotypes following germline transmission. RESULTS: Sequences flanking the integrated gene trap vector in 1B4 cells were used to isolate a full-length cDNA whose predicted amino acid sequence is identical to the human A2 protein at all but one of 341 amino acid residues. hnRNP A2/B1 transcripts extending into the provirus utilize a cryptic 3' splice site located 28 nucleotides downstream of the neomycin phosphotransferase start codon. The inserted Neo sequence and proviral poly(A) site function as an 3' terminal exon that is utilized to produce hnRNP A2/B1-Neo fusion transcripts, or skipped to produce wild-type hnRNP A2/B1 transcripts. This results in only a modest disruption of hnRNPA2/B1 gene expression. CONCLUSIONS: Expression of the occupied hnRNP A2/B1 gene and utilization of the viral poly(A) site are consistent with an exon definition model of pre-mRNA splicing. These results reveal a mechanism by which U3 gene trap vectors can be expressed without disrupting cellular gene expression, thus suggesting ways to improve these vectors for gene trap mutagenesis.

3' Untranslated Regions↗

Transposition of cyanobacterium insertion element ISY100 in Escherichia coli.

The genome of the cyanobacterium Synechocystis sp. strain PCC6803 has nine kinds of insertion sequence (IS) elements, of which ISY100 in 22 copies is the most abundant. A typical ISY100 member is 947 bp long and has imperfect terminal inverted repeat sequences. It has an open reading frame encoding a 282-amino-acid protein that appears to have partial homology with the transposase encoded by a bacterial IS, IS630, indicating that ISY100 belongs to the IS630 family. To determine whether ISY100 has transposition ability, we constructed a plasmid carrying the IPTG (isopropyl-beta-D-thiogalactopyranoside)-inducible transposase gene at one site and mini-ISY100 with the chloramphenicol resistance gene, substituted for the transposase gene of ISY100, at another site and introduced the plasmid into an Escherichia coli strain already harboring a target plasmid. Mini-ISY100 transposed to the target plasmid in the presence of IPTG at a very high frequency. Mini-ISY100 was inserted into the TA sequence and duplicated it upon transposition, as do IS630 family elements. Moreover, the mini-ISY100-carrying plasmid produced linear molecules of mini-ISY100 with the exact 3' ends of ISY100 and 5' ends lacking two nucleotides of the ISY100 sequence. No bacterial insertion elements have been shown to generate such molecules, whereas the eukaryotic Tc1/mariner family elements, Tc1 and Tc3, which transpose to the TA sequence, have. These findings suggest that ISY100 transposes to a new site through the formation of linear molecules, such as Tc1 and Tc3, by excision. Some Tc1/mariner family elements leave a footprint with an extra sequence at the site of excision. No footprints, however, were detected in the case of ISY100, suggesting that eukaryotes have a system that repairs a double strand break at the site of excision by an end-joining reaction, in which the gap is filled with a sequence of several base pairs, whereas prokaryotes do not have such a system. ISY100 transposes in E. coli, indicating that it transposes without any host factor other than the transposase encoded by itself. Therefore, it may be able to transpose in other biological systems.

Amino Acid Sequence↗

The ovalbumin gene: structural sequences in native chicken DNA are not contiguous.

The sequence organization of the structural ovalbumin gene and flanking sequences in native chicken DNA was studied by restriction mapping and filter hybridization using a nick-translated probe generated from pOV230, a recombinant plasmid that contains a full-length ovalbumin DNA synthesized from ovalbumin mRNA. The structural sequences of the ovalbumin gene in native chicken DNA were found to be noncontiguous because at least two restriction endonucleases that do not cut the structural sequence do cleave the natural gene into multiple fragments by cleaving within nonstructural sequences interspersed between the structural sequences. The observation that all ovalbumin DNA-containing sequences were contained within a single DNA fragment generated by BamHI digestion of total chicken DNA has allowed us to construct an inclusive restriction map of the natural ovalbumin gene which contains at least two "insert regions." These regions may be further subdivided into alternating structural and insert sequences. Both insert regions were located within the peptide-coding regions of the gene and the sizes of these insert regions were estimated to be approximately 1.0 and 1.5 kilobase pairs, respectively.

Animals↗

[Analysis of copies of the suffix-retroposon of Drosophila, produced using PCR on genomic DNA].

Nineteen DNA clones, containing copies of the suffix element retroposon of the Drosophila genome and produced by a polymerase chain reaction (PCR), were sequenced. Insertions of the copies in both orientations into microsatellite sequences (CAACA)n/(TGTTG)n and (TTTGT)n/(CACAAA)n were revealed. It was found that, if the microsatellite sequence has both a decanucleotide GCGGCCCGGG (GC-box) and a colinear alternating sequence (A)5(T)4(A)3(T)2(A)1(t)n (AT-box), the insertion of the suffix occurs in only one orientation. It is suggested, that in such apparently heterochromatic sequences, suffix element copies and probably some other retroposons can be inserted a particular orientation by means of site-specific recombination. An approach to analysis of the molecular organization of such sequences is proposed.

Animals↗

Insertion of non-intron sequence into maize introns interferes with splicing.

Transposable element (TE) insertion into or near plant introns can cause intron skipping and alternative splicing events, resulting in reduced expression. To explore the impact of inserted sequences on splicing, we added non-intron sequence to two maize introns and tested these chimeric introns in a maize transient expression assay. Non-intron sequence inserted into Adh1-S intron 1 and actin intron 3 decreased expression from the luciferase reporter gene; the insertion sites tested were not in intron regions thought to be essential for splicing. Alternatively spliced mRNAs were not observed in transcripts derived from the insertion variants. In contrast, addition of an internal segment of an intron to Adh1-S intron 1 resulted in normal splice site selection and efficient processing. Because the normal intron sequence (including the conserved splice junctions) was retained in all constructs, we hypothesize that added non-intron sequence can interfere with intron recognition and/or splicing.

Actins↗

Transposition of an Alu-containing element induced by DNA-advanced glycosylation endproducts.

Advanced glycosylation endproducts react with DNA and cause mutations and DNA transposition in bacteria. To investigate the mutagenic effect of advanced glycosylation in mammalian cells, plasmid DNA containing the lacI mutagenesis marker was modified by advanced glycosylation endproducts in vitro, transfected into murine lymphoid cells, recovered, and analyzed for mutations, plasmid size changes, and the presence of shared insertion sequences. An 853-bp host-derived DNA sequence, designated INS-1, was identified as an insertion element common to plasmids recovered from multiple independent transfections. Modification of DNA by advanced glycosylation increased by 60-fold the apparent frequency of INS-1 transposition: from 0.025% to 1.5%. The INS-1 element contains a 180-bp region that is homologous to the Alu repetitive sequence family. INS-1 was also observed to be present within larger insertional mutations and, in two cases, an apparently truncated version of INS-1 that lacks the Alu region was identified. These results demonstrate the experimental induction of DNA transposition involving mammalian chromosomal elements and suggest that advanced glycosylation may play a role in the formation of Alu-containing insertions that have been found to disrupt human genes.

Animals↗

Genetic organization and transposition properties of IS511.

IS511 is an endogenous insertion sequence (IS) of the bacterium Caulobacter crescentus strain CB15 and it is the first Caulobacter IS to be characterized at the molecular level. We determined the 1266-bp nucleotide sequence of IS511 and investigated its genetic organization, relationship to other ISs, and transposition properties. IS511 belongs to a distinct branch of the IS3 family that includes ISR1, IS476, and IS1222, based on nucleotide sequence similarity. The nucleotide sequence of IS511 encodes open reading frames (orfs) designated here as orfA and orfB, and their relative organization and amino acid sequences of the predicted protein products are very similar to those of orfAs and orfBs of other IS3 family members. Nuclease S1 protection assays identified an IS511 RNA, and its 5' end maps approximately 16 nucleotides upstream of orfA and about six nucleotides downstream of a sequence that is similar to the consensus sequence of C. crescentus housekeeping promoters. Evidence is presented that IS511 is capable of precise excision from the chromosome, and transposition from the chromosome to a plasmid. Transpositional insertions of IS511 occurred within sequences with a relatively high G + C content, and they were usually, but not always, flanked by a 4-bp direct repeat that matches a sequence at the site of insertion. We also determined the nucleotide sequence flanking the four endogenous IS511 elements that reside in the chromosome of C. crescentus. Our findings demonstrate that IS511 is a transposable IS that belongs to a branch of the IS3 family.

Amino Acid Sequence↗