PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Use of a cryptic splice donor site in the chloramphenicol acetyltransferase (CAT)-SV40 small-t antigen cassette generates alternative transcripts in transgenic rats.

The bacterial gene chloramphenicol acetyltransferase (CAT) is a widely used reporter in both in-vitro and in-vivo studies of genetic regulation. We have recently generated novel rat transgenic lines carrying an arylalkylamine N-acetyltransferase (AA-NAT) promoter-reporter construct in which CAT (with associated SV40 small-t antigen sequence) is the reporter. In addition to the predicted transgene transcript (1.9 kb), we identified an abundant 1.5 kb transcript which derives from an alternative splicing event that utilises a cryptic splice donor site located within the CAT gene. The native CAT open reading frame (ORF) is lost in the 1.5 kb transcript, and a western analysis has shown that protein deriving from an aberrant open reading frame is not expressed at detectable levels.

Alternative Splicing↗

Intervening sequence with conserved open reading frame in eubacterial 23S rRNA genes.

An intervening sequence (IVS) occurred in the 23S rRNA genes (rrl) of some, but not all, strains of four species of the spirochete genus Leptospira and was absent from strains in three other species. The IVS varied in size from 485 to 759 base pairs and replaced bases 1224-1245 in both copies of rrl. The two ends of each IVS shared 22-35 bases of complementarity that could form a stable double helix. The presence of an IVS correlated with a cleaved mature 23S rRNA that probably results from removal of the IVS without religation. The 3' site of cleavage was mapped within the inverted repeat of the IVS. An open reading frame of 121-133 amino acids was conserved in the IVS in all four species, oriented so that the sense strand was in the rRNA transcript. When the open reading frames were compared between species, they predicted polypeptides that showed between 51% and 78% amino acid conservation and similar DNA sequence conservation, indicating selection for protein function.

Amino Acid Sequence↗

Molecular diversity and phylogenetic analysis of mariner-like transposons in the genome of the silkworm Bombyx mori.

Genome-wide screening of mariner-like elements (MLEs) in the silkworm Bombyx mori has revealed the presence of five different types of MLEs (Bmmar1, Bmmar2, Bmmar3, Bmmar4 and Bmmar5). We isolated and characterized sixty copies of the MLEs representing the five Bmmar types. Their nucleotide sequences, nucleotide compositions, deduced transposase sequences, codon preferences, and the copy numbers showed extensive variations. Phylogenetic analysis of the sequences revealed that Bmmar1, Bmmar2, and Bmmar3 have been in the B. mori genome for a long time, while Bmmar4 is probably a recent invader of the genome. Because of the long-term association of Bmmar1 and Bmmar2 with the genome, highly mutated miniature Bmmar1 and Bmmar2 are widespread in the genome, and the footprints of these elements are also present in different silkworm genes. However, miniature copies of Bmmar4 were not detected. This recently acquired element has very few mutations. None of the characterized copies had functional transposase open reading frames. They essentially exist as fossils in the genome.

Amino Acid Sequence↗

Transcription at different salinities of Haloferax mediterranei sequences adjacent to partially modified PstI sites.

Two genomic sequences from the halophilic archaeon Haloferax mediterranei, where we had found PstI restriction-pattern modifications depending on the salinity of the growth medium, have been studied. A markedly salt-dependent differential expression has been detected in the nearby regions. Two of the open reading frames characterized correspond to two of the differentially expressed transcripts. In both cases the PstI sites were included in purine-pyrimidine alternancies suggestive of Z-DNA structures and located in non-coding regions with frequent repetitive motifs. A long alternating adenine-thymine tract also appears in the upstream regions of one of these open reading frames. A possible role of local DNA configuration in osmoregulation in this organism is discussed.

Amino Acid Sequence↗

Transcription of four satellite DNA subfamilies in Diprion pini (Hymenoptera, Symphyta, Diprionidae).

Four satellite DNA subfamilies Ps, Pv, Pv65 and Ec, resulting from the evolution of a common ancestral motif, were isolated and characterized in the genomic DNA of Diprion pini, a phytophagous of Pinus sylvestris. Consensus sequences were 148-312 bp long. Sequence analyses revealed that these satellite subfamilies have evolved from a 45-bp ancestral motif. The amounts of each satellite in the genome (0 - 10%) and the accessibility of the DNA to restriction enzymes were sex dependent. The migration of each monomer in polyacrylamide gels and the electrophoretic migration of d(AT) n > or = 3 residues showed that all four satellite subfamilies are curved. Their transcription was analyzed using reverse transcription and the polymerase chain reaction. Three satellite DNA subfamilies were transcribed on both strands, and in both sexes. However, the female satellite DNAs seem to be more actively transcribed than those of males, indicating that transcription is not constitutive. The lack of any significant open reading frame in satellite monomers indicates that the RNA may function as structural or catalytic RNA rather than encoding protein.

Animals↗

Structures of homologous composite transposons carrying cbaABC genes from Europe and North America.

IS1071 is a class II transposable element carrying a tnpA gene related to the transposase genes of the Tn3 family. Copies of IS1071 that are conserved with more than 99% nucleotide sequence identity have been found as direct repeats flanking a remarkable variety of catabolic gene sequences worldwide. The sequences of chlorobenzoate catabolic transposons found on pBRC60 (Tn5271) in Niagara Falls, N.Y., and on pCPE3 in Bologna, Italy, show that these transposons were formed from highly homologous IS1071 and cbaABC components (levels of identity, > 99.5 and > 99.3%, respectively). Nevertheless, the junction sequences between the IS1071L and IS1071R elements and the internal DNA differ by 41 and 927 bp, respectively, suggesting that these transposons were assembled independently on the two plasmids. The formation of the right junction in both transposons truncated an open reading frame for a putative aryl-coenzyme A ligase with sequence similarity to benzoate- and p-hydroxybenzoate-coenzyme A ligases of Rhodopseudomonas palustris.

Alcaligenes↗

IS1631 occurrence in Bradyrhizobium japonicum highly reiterated sequence-possessing strains with high copy numbers of repeated sequences RSalpha and RSbeta.

From Bradyrhizobium japonicum highly reiterated sequence-possessing (HRS) strains indigenous to Niigata and Tokachi in Japan with high copy numbers of the repeated sequences RSalpha and RSbeta (K. Minamisawa, T. Isawa, Y. Nakatsuka, and N. Ichikawa, Appl. Environ. Microbiol. 64:1845-1851, 1998), several insertion sequence (IS)-like elements were isolated by using the formation of DNA duplexes by denaturation and renaturation of total DNA, followed by treatment with S1 nuclease. Most of these sequences showed structural features of bacterial IS elements, terminal inverted repeats, and homology with known IS elements and transposase genes. HRS and non-HRS strains of B. japonicum differed markedly in the profiles obtained after hybridization with all the elements tested. In particular, HRS strains of B. japonicum contained many copies of IS1631, whereas non-HRS strains completely lacked this element. This association remained true even when many field isolates of B. japonicum were examined. Consequently, IS1631 occurrence was well correlated with B. japonicum HRS strains possessing high copy numbers of the repeated sequence RSalpha or RSbeta. DNA sequence analysis indicated that IS1631 is 2,712 bp long. In addition, IS1631 belongs to the IS21 family, as evidenced by its two open reading frames, which encode putative proteins homologous to IstA and IstB of IS21, and its terminal inverted repeat sequences with multiple short repeats.

Amino Acid Sequence↗

Cloning, DNA sequence, and complementation analysis of the Salmonella typhimurium hemN gene encoding a putative oxygen-independent coproporphyrinogen III oxidase.

Coproporphyrinogen oxidation is a last step in heme biosynthesis. The biochemically characterized eukaryotic coproporphyrinogen III oxidases have an obligate requirement for molecular oxygen, and a similar enzyme is encoded by the hemF gene in Salmonella typhimurium. Anaerobic heme synthesis requires an oxygen-independent coproporphyrinogen oxidase, which is probably encoded by the hemN gene in S. typhimurium. The hemN gene has been cloned from an insertion mutant. The nucleotide sequence was obtained and used for PCR amplification of the wild-type gene. A single open reading frame was identified as the hemN gene on the basis of its interruption by the insertion mutation and plasmid complementation studies of hemF hemN double mutants. The predicted HemN protein has 38% amino acid sequence identity to a putative anaerobic Rhodobacter sphaeroides coproporphyrinogen oxidase. The hemN RNA 5' end and the inferred transcription initiation site were mapped by primer extension. The 52.8-kDa HemN protein is expressed from the second ATG codon of the hemN open reading frame. An open reading frame with an unknown function directly upstream of hemN has a striking amino acid sequence, including 11 acidic residues in a row.

Amino Acid Sequence↗

Characterization of indigoidine biosynthetic genes in Erwinia chrysanthemi and role of this blue pigment in pathogenicity.

In the plant-pathogenic bacterium Erwinia chrysanthemi production of pectate lyases, the main virulence determinant, is modulated by a complex network involving several regulatory proteins. One of these regulators, PecS, also controls the synthesis of a blue pigment identified as indigoidine. Since production of this pigment is cryptic in the wild-type strain, E. chrysanthemi ind mutants deficient in indigoidine synthesis were isolated by screening a library of Tn5-B21 insertions in a pecS mutant. These ind mutations were localized close to the regulatory pecS-pecM locus, immediately downstream of pecM. Sequence analysis of this DNA region revealed three open reading frames, indA, indB, and indC, involved in indigoidine biosynthesis. No specific function could be assigned to IndA. In contrast, IndB displays similarity to various phosphatases involved in antibiotic synthesis and IndC reveals significant homology with many nonribosomal peptide synthetases (NRPS). The IndC product contains an adenylation domain showing the signature sequence DAWCFGLI for glutamine recognition and an oxidation domain similar to that found in various thiazole-forming NRPS. These data suggest that glutamine is the precursor of indigoidine. We assume that indigoidine results from the condensation of two glutamine molecules that have been previously cyclized by intramolecular amide bond formation and then dehydrogenated. Expression of ind genes is strongly derepressed in the pecS background, indicating that PecS is the main regulator of this secondary metabolite synthesis. DNA band shift assays support a model whereby the PecS protein represses indA and indC expression by binding to indA and indC promoter regions. The regulatory link, via pecS, between indigoidine and virulence factor production led us to explore a potential role of indigoidine in E. chrysanthemi pathogenicity. Mutants impaired in indigoidine production were unable to cause systemic invasion of potted Saintpaulia ionantha. Moreover, indigoidine production conferred an increased resistance to oxidative stress, indicating that indigoidine may protect the bacteria against the reactive oxygen species generated during the plant defense response.

Amino Acid Sequence↗

Effects of translocations on transcription from PVT.

We have previously described a transcription unit on human chromosome 8, designated as PVT, that is consistently disrupted by the minority forms of translocations [t(2;8) and t(8;22)] in Burkitt's lymphoma. PVT begins 57 kilobase pairs downstream of the proto-oncogene MYC and is more than 200 kilobase pairs in length. In order to explore the pathogenic impact of translocations affecting PVT, we have characterized further the structure and transcription of the locus. In normal cells, PVT is transcribed into a variety of RNAs, the diversity of which remains unexplained. Alleles of PVT affected by translocations give rise to additional RNAs. These RNAs arise from a fusion of the first exon of PVT on chromosome 8 to the constant region of an immunoglobulin light chain on either chromosome 2 or chromosome 22. We have found no evidence that any of the normal or abnormal transcripts of PVT give rise to a protein. Our results suggest that the pathogenic effects of the variant translocations in Burkitt's lymphoma are not executed by a gene situated in a vicinity of the chromosomal breakpoints. Instead, our data leave open the possibility that the effects of the translocations may be mediated by activation of the relatively distant MYC gene.

Base Sequence↗

Differences between coding and non-coding regions in the Trichomonas vaginalis genome: an actin gene as a locus model(1).

The sequence of a cloned genomic fragment of Trichomonas vaginalis containing a complete actin gene was determined. An uninterrupted open reading frame of 1128 nucleotides was found that codes for an actin gene. Two overlapped consensus promoter sequences for T. vaginalis were found 12 nucleotides upstream the actin initiation codon. In addition to actin, two incomplete open reading frames were found at the 5' and 3' ends of the clone. These two sequences are expressed and showed similarity to adenylate cyclase genes and a yeast hypothetical protein. The overall sequence showed a higher G+C content and a lower frequency of repeated sequences in the coding regions when compared with the non-coding regions. A similar unequal nucleotide distribution was found in various T. vaginalis genes retrieved from data bases.

Actins↗

Phylogenetic relationships among group II intron ORFs.

Group II introns are widely believed to have been ancestors of spliceosomal introns, yet little is known about their own evolutionary history. In order to address the evolution of mobile group II introns, we have compiled 71 open reading frames (ORFs) related to group II intron reverse transcriptases and subjected their derived amino acid sequences to phylogenetic analysis. The phylogenetic tree was rooted with reverse transcriptases (RTs) of non-long terminal repeat retroelements, and the inferred phylogeny reveals two major clusters which we term the mitochondrial and chloroplast-like lineages. Bacterial ORFs are mainly positioned at the bases of the two lineages but with weak bootstrap support. The data give an overview of an apparently high degree of horizontal transfer of group II intron ORFs, mostly among related organisms but also between organelles and bacteria. The Zn domain (nuclease) and YADD motif (RT active site) were lost multiple times during evolution. Differences in domain structures suggest that the oldest ORFs were concise, while the ORF in the mitochondrial lineage subsequently expanded in three locations. The data are consistent with a bacterial origin for mobile group II introns.

Amino Acid Sequence↗

The glgB gene from the thermophile Bacillus caldolyticus encodes a thermolabile branching enzyme.

We have cloned the structural gene for the Bacillus caldolyticus glycogen branching enzyme (glgB) in Escherichia coli. The glgB gene consisted of a 1998 bp open reading frame (ORF) encoding a 78,087 Da protein, which was highly similar to the Bacillus stearothermophilus branching enzyme. The 5' end of a second gene that encoded a protein with extensive similarity to E. coli ADP-glucose pyrophosphorylase (ADPGP) partly overlapped the 3' end of the glgB gene. A putative promoter recognized by Bacillus subtilis RNA polymerase containing the sigma factor H (E-sigma H) preceded the genes. These data suggest that in contrast to the situation observed in B. stearothermophilus, the genes involved in glycogen synthesis in B. caldolyticus are clustered on the chromosome, and are presumably coordinately expressed during the early stages of sporulation. An incomplete third gene started upstream of B. caldolyticus glgB. This gene was highly similar to a gene found directly upstream of B. stearothermophilus glgB, which encodes a putative membrane protein with unknown function. The B. caldolyticus glgB gene was expressed in E. coli and B. subtilis. Surprisingly, the branching enzyme appeared to be thermolabile, the temperature of optimal activity being only 39 degrees C.

1,4-alpha-Glucan Branching Enzyme↗

Killer system of Kluyveromyces lactis: the open reading frame 10 of the pGK12 plasmid encodes a putative DNA binding protein.

ORF 10 of the K2 plasmid from Kluyveromyces lactis encodes a small basic protein (22.3% lysine). The function of its product has been investigated. Western blot analysis, using an antibody against MS2 RNA polymerase/ORF 10 fusion protein, reveals a protein band with an apparent molecular weight of 14 kDa. The protein can bind a DNA-Sepharose column, and is eluted by 350 mM-salt. Immunoprecipitation experiments show that the ORF 10 protein coprecipitates with the linear genomic DNAs of the two killer plasmids (K1 and K2). From Western/Southern blot data, it is possible to conclude that the interaction between protein and DNA occurs directly, rather than via other protein(s). ORF 10 is easily detected by Western blot and its transcript is one of the most abundant of the K2 plasmid, suggesting that this protein may have a structural rather than a regulatory function. This possibility is also suggested by the observed sequence homology between the ORF 10 protein and the family of histone-like proteins.

Amino Acid Sequence↗

Analysis of the leader and capsid coding regions of persistent and neurovirulent strains of Theiler's virus.

Most strains of Theiler's virus (TMEV) cause a persistent infection of the central nervous system of the mouse and a chronic demyelinating disease considered a model for multiple sclerosis. Two strains, on the contrary, cause an acute encephalitis and kill mice in a matter of days. We sequenced the leader and capsid coding region of three persistent (TO4, WW, and Yale) isolates and one neurovirulent (FA) isolate of TMEV. We compared these sequences and those already published for other isolates (DA, BeAn, GDVII, and Vilyuisk). The results suggest that virulent and persistent strains did not evolve as two separate groups, but rather that neurovirulent strains arose from a subgroup of persistent strains. The sequences of viruses isolated in different geographic areas and at different times were highly homologous, a surprising finding for an RNA virus. This suggests that severe constraints are imposed on the genome during the viral life cycle. The sequences of the TO4 and WW strains were identical, suggesting that the latter came from a laboratory contamination. The genomes of all the persistent strains sequenced so far contain an alternate open reading frame in the L region, which has been shown, in the case of the DA strain, to code for an 18-kDa protein called "I".

Amino Acid Sequence↗

Nucleotide sequence of a new isolate of ribgrass mosaic tobamovirus infecting Impatiens New Guinea.

The complete nucleotide sequence of a tobamovirus isolated from Impatiens New Guinea was determined. The genome was 6302 nt long, and its genomic organisation was similar to those of other crucufer tobamoviruses. Sequence comparisons with the corresponding sequences of other crucifer tobamoviruses revealed highest levels of identity with the ribgrass mosaic virus (Shanghai isolate). A small open reading frame putatively encoding a 4.5-kDa protein with a low degree of similarity to the ORF6 of tobacco mosaic virus was found nested in the movement protein gene.

Amino Acid Sequence↗

Cloning and sequence analysis of the Candida utilis HIS3 gene.

A DNA fragment, carrying the Candida utilis HIS3 gene, has been isolated from a genomic DNA library by complementation of the E. coli hisB mutant. Its nucleotide sequence was determined and it predicts a single open reading frame of 675 bp (224 aa). The deduced amino acid sequence is highly homologous to other yeast and fungi HIS3 genes.

Amino Acid Sequence↗

Human platelet glycoproteins V and IX: mapping of two leucine-rich glycoprotein genes to chromosome 3 and analysis of structures.

Human platelet glycoproteins Ib alpha, Ib beta, V, and IX comprise an interrelated set of molecules (the Ib-V-IX system) that together form a surface adhesion receptor for the ligand, von Willebrand factor. To complete the primary structural characterization of the genes involved in this system, we have analyzed cosmid clones for both the glycoprotein V and IX genes and used these clones to localize the two genes by fluorescence in situ hybridization. Both genes were found on the long arm of chromosome 3, but at distinct sites, the GPV gene on 3 band q29 and the GP IX gene on 3 band q21. The transcriptional start site of the GPV gene was defined by "anchored" PCR and primer extension. The GPV gene contains two exons, the first consisting of approximately 37 bases and the second of approximately 3500 bases, interrupted by a single 958 base intron. The GPV transcript has multiple start sites spread over a twenty base region. The 5' flanking region of the GPV gene has a series of potential consensus regulatory elements including GATA, ets, and Sp-1 sites, similar to those found in other described megakaryocyte/platelet genes, including those of the Ib-V-IX system. In assessing the four Ib-V-IX genes as a group, all four have a simple, "intron-depleted" structure with the entire open reading frame of the mature polypeptide located within a single exon.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗