PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Evolution of mobile group I introns: recognition of intron sequences by an intron-encoded endonuclease.

Mobile group I introns are hypothesized to have arisen after invasion by endonuclease-encoding open reading frames (ORFs), which mediate their mobility. Consistent with an endonuclease-ORF invasion event, we report similarity between exon junction sequences (the recognition site for the mobility endonuclease) and intron sequences flanking the endonuclease ORF in the sunY gene of phage T4. Furthermore, we have demonstrated the ability of the intron-encoded endonuclease to recognize and cleave these intron sequences when present in fused form in synthetic constructs. These observations and accompanying splicing data are consistent with models in which the invading endonuclease ORF is provided safe haven within a splicing element. In turn the intron is afforded immunity to the endonuclease product, which imparts mobility to the intron.

Bacteriophage T4↗

Ontological analysis of gene expression data: current tools, limitations, and open problems.

Independent of the platform and the analysis methods used, the result of a microarray experiment is, in most cases, a list of differentially expressed genes. An automatic ontological analysis approach has been recently proposed to help with the biological interpretation of such results. Currently, this approach is the de facto standard for the secondary analysis of high throughput experiments and a large number of tools have been developed for this purpose. We present a detailed comparison of 14 such tools using the following criteria: scope of the analysis, visualization capabilities, statistical model(s) used, correction for multiple comparisons, reference microarrays available, installation issues and sources of annotation data. This detailed analysis of the capabilities of these tools will help researchers choose the most appropriate tool for a given type of analysis. More importantly, in spite of the fact that this type of analysis has been generally adopted, this approach has several important intrinsic drawbacks. These drawbacks are associated with all tools discussed and represent conceptual limitations of the current state-of-the-art in ontological analysis. We propose these as challenges for the next generation of secondary data analysis tools.

Algorithms↗

Characterization of new proteins found by analysis of short open reading frames from the full yeast genome.

We have analysed short open reading frames (between 150 and 300 base pairs long) of the yeast genome (Saccharomyces cerevisiae) with a two-step strategy. The first step selects a candidate set of open reading frames from the DNA sequence based on statistical evaluation of DNA and protein sequence properties. The second step filters the candidate set by selecting open reading frames with high similarity to other known sequences (from any organism). As a result, we report ten new predicted proteins not present in the current sequence databases. These include a new alcohol dehydrogenase, a protein probably related to the cell cycle, as well as a homolog of the prokaryotic ribosomal protein L36 likely to be a mitochondrial ribosomal protein coded in the nuclear genome. We conclude that the analysis of short open reading frames leads to biologically interesting discoveries, even though the quantitative yield of new proteins is relatively low.

Alcohol Dehydrogenase↗

Characterization of an open reading frame involved in site-specific integration of filamentous phage Cf1t from Xanthomonas campestris pv. citri.

Cf1t is a single-stranded DNA filamentous phage; a 1.9-kb segment of DNA from Cf1t was found to be responsible for site-specific integration into Xanthomonas campestris pv. citri (XW47), in the absence of any Xanthomonas origin of replication. Deletion analysis and introduction of amber stop codons into this fragment from Cf1t revealed an open reading frame (ORF344) which was involved in the integration function. The predicted amino-acid sequence of ORF344 bears no homology with conserved sequences of the integrase family.

Amino Acid Sequence↗

A superantigen encoded in the open reading frame of the 3' long terminal repeat of mouse mammary tumour virus.

Mice express a collection of superantigens, which bind to class II major histocompatibility proteins and interact with T cells bearing particular V beta chains as part of their alpha beta receptors. These superantigens have been suggested to be encoded by exogenous or endogenous mouse mammary tumour viruses. One such superantigen is now shown to be encoded in the open reading frame of the long terminal repeat of a mammary tumour virus, a gene of previously unknown function.

Amino Acid Sequence↗

The leptin receptor promoter controls expression of a second distinct protein.

The leptin receptor (OB-R) is a single membrane- spanning protein that mediates the weight-regulatory effects of leptin (OB protein). Several mRNA splice variants have been described which either encode OB-R proteins with cytoplasmic domains of different length or the OB-R and B219/OBR variants, which have different 5'-untranslated regions. Here we report evidence for the synthesis of a human mRNA splice variant of the OB-R gene that potentially encodes a novel protein, leptin receptor gene-related protein (OB-RGRP), which displays no sequence similarity to the leptin receptor itself. This OB-RGRP transcript contains the first two OB-R gene 5'-untranslated exons, but then is alternatively spliced to two novel exons which were mapped to a yeast artificial chromosome containing the leptin receptor gene. First identified by analysis of a large human expressed sequence tag database, the OB-RGRP transcript has now also been found in human and mouse tissues by the use of PCR. Preliminary experiments suggest that OB-RGRP and the OB-R variants share similar patterns of expression that are distinct from that of the B219/OBR variant. OB-RGRP is highly homologous to putative open reading frames in both yeast and Caenorhabditis elegans , suggesting a phylogenetically conserved role for this novel protein.

Alternative Splicing↗

Identification of the repressor-encoding gene of the Lactobacillus bacteriophage A2.

The repressor gene of the Lactobacillus phage A2 has the following properties: it (i) encodes a 224-residue polypeptide with DNA binding and RecA cleavage motifs, (ii) is expressed in lysogenic cultures, and (iii) confers superinfection immunity on the host. Adjacent, but divergently transcribed, lies another open reading frame whose product resembles the lambda Cro protein. In the 161-bp intergenic segment, putative promoters and operators have been detected.

Amino Acid Sequence↗

Genomic sequence of C1, the first streptococcal phage.

C(1), a lytic bacteriophage infecting group C streptococci, is one of the earliest-isolated phages, and the method of bacterial classification known as phage typing was defined by using this bacteriophage. We present for the first time a detailed analysis of this phage by use of electron microscopy, protein profiling, and complete nucleotide sequencing. This virus belongs to the Podoviridae family of phages, all of which are characterized by short, noncontractile tails. The C(1) genome consists of a linear double-stranded DNA molecule of 16,687 nucleotides with 143-bp inverted terminal repeats. We have assigned functions to 9 of 20 putative open reading frames based on experimental substantiation or bioinformatic analysis. Their products include DNA polymerase, holin, lysin, major capsid, head-tail connector, neck appendage, and major tail proteins. Additionally, we found one intron belonging to the HNH endonuclease family interrupting the apparent lysin gene, suggesting a potential splicing event yielding a functional lytic enzyme. Examination of the C(1) DNA polymerase suggests that this phage utilizes a protein-primed mechanism of replication, which is prominent in the phi29-like members of Podoviridae. Consistent with this evidence, we experimentally determined that terminal proteins are covalently attached to both 5' termini, despite the fact that no homology to known terminal proteins could be elucidated in any of our open reading frames. Likewise, comparative genomics revealed no close evolutionary matches, suggesting that the C(1) bacteriophage is a unique member of the Podoviridae.

Animals↗

Stress responses in alfalfa (Medicago sativa L.) 11. Molecular cloning and expression of alfalfa isoflavone reductase, a key enzyme of isoflavonoid phytoalexin biosynthesis.

The major phytoalexin in alfalfa is the isoflavonoid (-)-medicarpin (or 6aR, 11aR)-medicarpin. Isoflavone reductase (IFR), the penultimate enzyme in medicarpin biosynthesis, is responsible for introducing one of two chiral centers in (-)-medicarpin. We have isolated a 1.18 kb alfalfa cDNA (pIFRalf1) which, when expressed in Escherichia coli, converts 2'-hydroxyformononetin stereospecifically to (3R)-vestitone, as would be predicted for IFR from alfalfa. The calculated molecular weight of the polypeptide (35,400) derived from the 954 bp open reading frame compares favorably to estimated Mrs determined for IFR proteins purified from other legumes. The transcript (1.4 kb) is highly induced in elicited alfalfa cell cultures. The kinetics of induction are consistent with the appearance of IFR activity, the accumulation of medicarpin, and the observed induction of other enzymes in the pathway. Low levels of IFR transcripts were found in healthy plant parts (roots and nodules) which accumulate low levels of a medicarpin glucoside. IFR appears to be encoded by a single gene in alfalfa. The cloning of IFR opens up the possibility of genetic manipulation of phytoalexin biosynthesis in alfalfa by altering isoflavonoid stereochemistry.

Amino Acid Sequence↗

Characterization of two minicircular plasmid-like DNAs isolated from date-palm mitochondria.

We report here the identification and characterization of two minicircular plasmid-like DNAs isolated from mitochondria of a moroccan date-palm variety. Both molecules were cloned and used as probes in Southern analyses of mitochondrial and total-cellular DNA. Evidence was obtained that these plasmid-like DNAs cross-hybridized but did not show any homology to nuclear, chloroplastic, or main mitochondrial genomes. Sequence analysis revealed that both minicircles, 1,346- and 1,160-bp long, share several stretches of homology, the most important consisting of three identical clusters of lengths 42, 47 and 38 bp. In contrast, no major homology was observed with the other higher-plant plasmid-like DNAs reported so far. Sequence analysis also revealed the presence, in the same strand of one of the minicircles, of two open reading frames potentially encoding proteins 89 and 86 amino acids in length. Interestingly, Northern analyses, using single strands of each minicircle as probes, showed the presence of two transcripts hybridizing only with the strand bearing these two open reading frames. However, computer-assisted comparison of the predicted polypeptide sequences with a protein-sequence library failed to detect any significant homology to known sequences.

Base Sequence↗

Characterization of an acetyl-CoA C-acetyltransferase (thiolase) gene from Clostridium acetobutylicum ATCC 824.

Thiolase (Thl) is an important enzyme at the junction in the pathway leading to the production of either acids (acetate or butyrate) or solvents (acetone, butanol or ethanol) during the growth of Clostridium acetobutylicum ATCC 824. Cloning and expression of the Thl-encoding gene (thl) has been described [Petersen and Bennett, Appl. Environ. Microbiol. 57 (1991) 2735-2741], as has the purification and properties of the enzyme [Wiesenborn et al., Appl. Environ. Microbiol. 54 (1988) 2717-2722]. Here, we present the complete nucleotide sequence (1.9 kb) of thl. The gene encodes a protein of 392 amino acids (aa) (41,237 Da), which mass is in agreement with previous findings using the purified protein. Primer extension analysis has defined the promoter region, and a stem-loop structure found at the end of thl indicates that it is not part of an operon. The aa sequence of Thl showed homology to those of four other beta-ketothiolases: (i) PhbC of Alcaligenes eutrophus, (ii) PhbA of Chromatium vinosum, (iii) PhbA of Thiocystis violacea and (iv) PhbA of Zoogloea ramigera. The C terminus of an open reading frame found upstream from the Thl sequence is similar to OrfX of Bacillus subtilis and to NfrC of Escherichia coli.

Acetyl-CoA C-Acetyltransferase↗

Characterization of the tetracycline resistance plasmid pMD5057 from Lactobacillus plantarum 5057 reveals a composite structure.

The 10,877bp tetracycline resistance plasmid pMD5057 from Lactobacillus plantarum 5057 was completely sequenced. The sequence revealed a composite structure containing DNA from up to four different sources. The replication region had homology to other plasmids of lactic acid bacteria while the tetracycline resistance region, containing a tet(M) gene, had high homology to sequences from Clostridium perfringens and Staphylococcus aureus. Within the tetracycline resistance region a Lactobacillus IS-element was found. The remaining part of the plasmid contained three open reading frames with unknown functions. The composite structure with several truncated genes suggests a recent assembly of the plasmid. This is the first sequence of an antibiotic resistance plasmid isolated from L. plantarum.

Base Sequence↗

Isolation and sequence analysis of the cDNA encoding subunit C of human CCAAT-binding transcription factor.

Sets of cDNA clones corresponding to genes with different expression specificity, 'brain-specific', 'common', 'liver-specific', were identified by differential hybridization of a human fetal brain cDNA library with total cDNA probes of human fetal brain and human fetal liver. Nucleotide sequence analysis revealed that one of the 'common' clones contained the cDNA encoding subunit C of human CCAAT-binding transcription factor. The isolated human CBF-C cDNA is 1977 nt long and consists of the full-length 3'-untranslated region (781 nt) with a poly (A) tail at the 3' end 185 nucleotides of 5'-untranslated region and the open reading frame (1011 nt), encoding a 337-aa protein with 91.7% homology with the translated region of rat CBF-C cDNA.

Amino Acid Sequence↗

Pertussis toxin export requires accessory genes located downstream from the pertussis toxin operon.

Pertussis toxin, a major virulence factor of Bordetella pertussis, is an oligomeric protein composed of five different subunits that are exported individually to the periplasmic space by the signal peptide-dependent pathway. After assembly, the protein is exported from the periplasm to the extracellular compartment. We show that pertussis toxin secretion across the outer membrane requires the gene product of at least one gene (ptlC) that is located downstream from the pertussis toxin operon. The amino acid sequence of PtlC shows a high degree of homology to VirB4, a protein encoded by the virB operon, which contains 11 open reading frames that are involved in the transfer of T-DNA from Agrobacterium tumefaciens to the plant cells. This is a novel mechanism of protein export in Gram-negative bacteria.

Agrobacterium tumefaciens↗

Sequence of a gene cluster from Malonomonas rubra encoding components of the malonate decarboxylase Na+ pump and evidence for their function.

Malonate decarboxylation in Malonomonas rubra involves the formation of malonyl-S-[acyl-carrier protein] from acetyl-S-[acyl-carrier protein] and malonate, carboxyltransfer to a biotin protein and its decarboxylation that is coupled to delta mu Na+ generation. The genes encoding components of the malonate decarboxylase enzyme system have been cloned and sequenced. These are located within a gene cluster of approximately 11 kb comprising 14 genes that have been termed madYZGBAECDHKFLMN in the given order. Upstream of madY an open reading frame pointing into the opposite direction of the mad genes was found with structural similarities to insertion-sequence elements. The upstream region also contains DNA regions which are typical for an Escherichia coli sigma 70 promoter. Within 950 bp downstream of madN no other open reading frame was found. This region contains a putative terminator sequence. The intergenic regions within the mad gene cluster are short (usually < 70 bp, maximum 302 bp) and ribosome binding sites were defined before all 14 genes. Thus, this DNA region could form a transcriptional unit and all 14 genes could be translated into proteins. The genes madABCDEF encode the structural proteins of the malonate decarboxylase as yet identified. By comparing protein and DNA sequences and by data bank searches for related proteins with known function the following assignments could be made: MadA represents the acyl-carrier-protein-transferase component. MadB is the integral membrane-bound carboxybiotin protein decarboxylase, MadC and MadD are the two subunits of the carboxyltransferase, MadE is the acyl carrier protein and MadF is the biotin protein. Sequence comparison further indicates that MadH could be involved in the acetylation of the phosphoribosyl-dephospho-CoA prosthetic group and MadG could be involved in its biosynthesis. MadL and MadM are membrane proteins that could function as malonate carrier. The function of the madY,Z,K and N gene products is as yet unknown.

Amino Acid Sequence↗