PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Unusual transcriptional and translational regulation of the bacteriophage Mu mom operon.

The bacteriophage Mu mom gene encodes a novel DNA modification that protects the viral genome against a wide variety of restriction endonucleases. Expression of mom is subject to a series of unusual regulatory controls. Transcription requires the action of a phage-encoded protein, C, which binds (probably as a dimer) the mom promoter from -33 to -52 (with respect to the transcription start site) in two adjacent DNA major grooves on one face of the helix. No apparent direct interaction between C and the host RNA polymerase (RNAP) is evident; however, C binding alters mom DNA conformation. In the absence of C, RNAP binds the mom promoter at a site that results in transcription in a direction away from the mom gene. The function of this transcription is unknown. An additional layer of transcriptional regulation complexity is due to the fact that the host Dam DNA-(N6-adenine)methyltransferase is required. Dam methylation of three closely spaced upstream GATC sequences is necessary to prevent binding by the host protein, OxyR, which acts as a repressor. Repression is not mediated by inhibition of C binding, but rather through interference with C-mediated recruitment of RNAP to the correct site. Translation of mom is regulated by the phage Com protein. Com is only 62 amino acids long and contains a zinc finger-like structure (coordinated by four cysteine residues) in the amino terminal domain. Com binds mom mRNA 5' to the mom open reading frame, whose translation start signals are contained in a stem-loop translation-inhibition-structure. Com binding to its target site (5' to and adjacent to the translation-inhibition-structure) results in a stable change in RNA secondary structure that exposes the translation start signals.

Amino Acid Sequence↗

Human PON2 gene at 7q21.3: cloning, multiple mRNA forms, and missense polymorphisms in the coding sequence.

We report the cloning and characterization of human PON2, a paraoxonase-related gene-2 that is physically linked with PON1 and PON3 on 7q2l.3. PON2 is ubiquitously expressed and we identified several mRNA forms produced by alternative splicing, or by the use of a second transcription start site. We also describe two polymorphisms in the coding sequences that, in the protein deduced from the longest open reading frame, predict an alanine-to-glycine substitution at residue 147 and a serine-to-cysteine substitution at residue 310.

Amino Acid Sequence↗

Folate biosynthesis pseudogenes, PsifolP and PsifolK, and an O-sialoglycoprotein endopeptidase gene homolog in the phytoplasma genome.

Phytoplasmas are wall-less phytopathogenic prokaryotes of small genome sizes that are obligate parasites of insect vectors and plant hosts. We have cloned a clover phyllody (CPh) phytoplasma DNA locus containing five potential coding sequences. Two were identified as pseudogenes (PsifolP and PsifolK) homologous to folP and folK genes, which encode dihydropteroate synthase (DHPS) and 6-hydroxymethyl-7,8-dihydropterin pyrophosphokinase (HPPK), respectively, in other bacteria. Evolution of the phytoplasma presumably involved loss of functions through the formation of these and other pseudogenes during adaptation to obligate parasitism. The findings suggest that the phytoplasma lacks capacity for de novo folate biosynthesis and possesses a transport system for absorption of preformed folate from host cells. The PsifolP-PsifolK region was flanked by three open reading frames (ORFs) encoding a DegV family protein, a hypothetical protein with a P60-like lipoprotein domain homologous with the P60-like Mycoplasma hominis protein, and a glycoprotease (Gcp) protein that possibly functions as a host adaptation or virulence factor.

Amino Acid Sequence↗

Characterisation of the diol dehydratase pdu operon of Lactobacillus collinoides.

The three genes (pduCDE) encoding the diol dehydratase of Lactobacillus collinoides were sequenced. They exhibited strong identities with the ddrABC and pduCDE genes of Klebsiella oxytoca and Salmonella enterica, respectively. These genes are part of a putative operon with at least four other genes. An eighth open reading frame was identified as homologous to the pocR gene (encoding the operon regulatory protein). Although the enzyme was detected in exponential growth phase, PduCDE activity was increased at the end of exponential phase in presence of 1,2-propanediol.

Bacterial Proteins↗

Molecular characterization of the fragilysin pathogenicity islet of enterotoxigenic Bacteroides fragilis.

Enterotoxigenic strains of Bacteroides fragilis produce an extracellular metalloprotease toxin (termed fragilysin) which is cytopathic to intestinal epithelial cells and induces fluid secretion and tissue damage in ligated intestinal loops. We report here that the fragilysin gene is contained within a small genetic element termed the fragilysin pathogenicity islet. The pathogenicity islet of B. fragilis VPI 13784 was defined as 6,033 bp in length and contained nearly perfect 12-bp direct repeats near its ends. Sequencing across the ends of the pathogenicity islet from two additional enterotoxigenic strains, along with PCR analysis of 20 additional enterotoxigenic strains, revealed that the islet is inserted at a specific site on the B. fragilis chromosome. The site of integration in three nontoxigenic strains contained a 17-bp GC-rich sequence which was not present in toxigenic strains and may represent a target sequence for chromosomal integration. In addition to the fragilysin gene, we identified an open reading frame encoding a predicted protein with a size and structural features similar to those of fragilysin. The deduced amino acid sequence was 28.5% identical and 56.3% similar to fragilysin and contained a nearly identical zinc-binding motif and methionine-turn region.

Amino Acid Sequence↗

Phage operon involved in sensitivity to the Lactococcus lactis abortive infection mechanism AbiD1.

Phage bIL66 is unable to grow on Lactococcus lactis cells harboring the abortive infection gene abiD1. Spontaneous phage mutants able to grow on AbiD1 cells were used to study phage-Abi interaction. A 1.33-kb DNA segment of a mutant phage allowed growth of AbiD1s phages in AbiD1 cells when present in trans. Sequence analysis of this segment revealed an operon composed of four open reading frames, designated orf1 to orf4. The operon is transcribed 10 min after infection from a promoter presenting an extended -10 consensus sequence but no -35 sequence. Analysis of four independent AbiD1r mutants revealed a different point mutation localized in orf1, implying that this open reading frame is needed for sensitivity to AbiD1. However, the sensitivity is partly suppressed when orf3 is expressed in trans on a high-copy-number plasmid, suggesting that AbiD1 acts by decreasing the concentration of an available orf3 product.

Amino Acid Sequence↗

DNA sequence of a gene in Escherichia coli encoding a putative tripartite transcription factor with receiver, ATPase and DNA binding domains.

We have sequenced downstream of the last previously sequenced gene of the glucitol operon (gutABDMRQ) in E. coli and have found that gutQ is the last gene of this operon. Downstream of the gutQ gene is found a palindromic unit (PU or REP sequence), followed by a large open reading frame of 1515 (or possibly 1590) bps transcribed in the direction opposite to that of the gut operon. This open reading frame encodes a protein of 504 (or possibly 529) amino acids with a tripartite structure. The N-terminal "receiver" domain of 187 (or possibly 212) residues is homologous to the FhlA protein of E. coli, a transcriptional activator of formate hydrogen lyase. It may possess a short domain at its extreme N-terminus exhibiting sequence similarity to carbohydrate binding proteins. The central ATPase domain (236 residues) exhibits greatest sequence similarity to the HydG protein of E. coli, a transcriptional activator of labile hydrogenase. The C-terminal DNA binding domain (81 residues) is homologous to NtrX of Azorhizobium caulinodans, a protein involved in transcriptional regulation of nitrogen fixation. Sequence comparisons with well-characterized transcription factors suggest that ORF504 encodes a protein that hydrolyzes ATP to generate the open transcriptional initiation complex of sigma 54-dependent promoters, possibly in response to redox conditions and/or ligand binding. We propose that this tripartite transcription factor arose by fusion of gene fragments encoding its three constituent modules.

Adenosine Triphosphatases↗

[Nucleotide sequence of the pNB2 plasmid from thermophilic bacteria Clostridium thermosaccharolyticum].

The complete nucleotide sequence of pNB2, a 1.9-kilobases cryptic plasmid from thermophilic Clostridium thermosaccharolyticum has been determined. The plasmid consists of 1882 base pairs and has a G+C composition of 27.2%. The sequence contains three open reading frames capable of coding for polypeptides two of which were identified in maxicell Escherichia coli extracts. Our future studies are directed toward a construction of pNB2-derivatives as vectors for Clostridia.

Amino Acid Sequence↗

Alternative splicing of the APC gene and its association with terminal differentiation.

The human tumor suppressor gene, APC, is composed of at least 21 exons, 7 of which are alternatively expressed. Sixteen APC transcripts that differ in their 5'-most regions and arise by the alternative inclusion of 6 of these exons have been identified by reverse transcription-PCR analysis of RNA prepared from human, mouse, and rat cell lines and tissues. Tissue-specific differences were observed in the expression of APC transcripts without exon 1, a coding region for a heptad repeat that supports APC homodimerization. Transcripts without exon 1 were observed at high levels in postmitotic, differentiated tissues and in two cell lines following the induction of differentiation. Sequence analysis of these novel open reading frames predicts APC proteins with different amino-terminal domains and therefore potentially different abilities to associate with other proteins. Our findings suggest that the alternative splicing of APC leads to alternative forms of APC proteins with potentially unique functions in growth control and differentiation.

Alternative Splicing↗

Isolation and characterization of idsA: the gene for the short chain isoprenyl diphosphate synthase from Methanobacterium thermoautotrophicum.

The gene that encodes the bifunctional short chain isoprenyl diphosphate synthase (idsA) for synthesis of farnesyl diphosphate and geranylgeranyl diphosphate in Methanobacterium thermoautotrophicum, a strict archaebacterial anaerobe, was isolated from a genomic DNA library by colony-lift hybridization and sequenced. Amino acid sequences were obtained for the N-terminus of the enzyme and for internal peptide fragments generated by proteolysis and treatment with cyanogen bromide. Degenerate primers based on the amino acid sequences were used in PCR to synthesize a 220-bp probe from genomic DNA. The probe was radiolabeled and used to isolate idsA. DNA sequencing revealed a 975-bp open reading frame located within an operon. The encoded 325-amino-acid protein contained five conserved regions found in eubacterial and eukaryotic farnesyl diphosphate and geranylgeranyl diphosphate synthases, including aspartate-rich motifs commonly found in prenyltransferases.

Alkyl and Aryl Transferases↗

Genomic organization of the X-linked inhibitor of apoptosis and identification of a novel testis-specific transcript.

Here we report the genomic organization and mapping of the X-linked inhibitor of apoptosis gene (BIRC4, also known as XIAP and hILP) and the identification of a closely related transcript. BIRC4 is located on Xq25 and is composed of seven exons. The intron/exon structure is highly conserved between the mouse homologue and its human counterpart. Four bands cross-react with a BIRC4 coding region probe on a genomic Southern blot. One of these cross-reactive bands encodes an intronless gene that expresses a 2.2-kb transcript solely in the testis. This testis-specific transcript contains a putative open reading frame (ORF) that is homologous to the carboxy-terminal end of BIRC4; overexpression of this ORF shows protective effects against BAX-induced apoptosis.

Animals↗

Purification of a soluble hepatitis E open reading frame 2-derived protein with unique antigenic properties.

The second open reading frame (ORF2) of hepatitis E virus (HEV) is predicted to encode a 73-kDa capsid protein (1). When full-length ORF2 was expressed in insect cells (Spodoptera frugiperda (Sf9)) using a recombinant baculovirus, two distinct HEV polypeptides were observed: a full-length insoluble 73-kDa protein, and a soluble 56.5-kDa protein. Following purification and sequence analysis, it was determined that the 56.5-kDa protein was derived from endoproteolytic cleavage site that was between the Thr and Ala residues located at amino acids 111 and 112 in the ORF2 sequence with the carboxy terminus corresponding to residue 636 of the ORF2 sequence. Comparative ELISA data using human acute-phase antisera demonstrated that the 56.5-kDa protein served as a highly reactive antigen in detecting anti-HEV antibodies. These data suggest that the 56.5-kDa protein may serve as a particularly useful antigen for both diagnostic and vaccine purposes.

Amino Acid Sequence↗

Nucleotide sequence of the yeast glutathione S-transferase cDNA.

The nucleotide sequence (658 bp) of the cDNA coding for glutathione S-transferase Y-2 of yeast Issatchenkia orientalis was obtained. The cDNA clone contains an open reading frame of 570 nucleotides encoding a polypeptide comprising 190 amino acids with a molecular weight of 21,520. The primary amino acid sequence of the enzyme exhibits only 25.0% and 21.1% identity with 177 and 151 amino acid residues of maize glutathione S-transferase I and rat glutathione S-transferase Yb2, respectively.

Amino Acid Sequence↗

Phylogenetic analysis of ORF5 and ORF7 sequences of porcine reproductive and respiratory syndrome virus (PRRSV) from PRRS-positive Italian farms: a showcase for PRRSV epidemiology and its consequences on farm management.

We investigated the dynamics of porcine reproductive and respiratory syndrome virus (PRRSV) variability in a range of swine PRRS-positive farms located in Northern Italy, to provide insights into the epidemiology and diffusion of the virus, particularly throughout the entire swine production chain. In this context, we also examined the effectiveness and the critical points of a recently developed gilts acclimatization program in swine breeder farms. To achieve these aims, we designed new primers and determined 64 complete open reading frame 5 (ORF5) sequences, representing Italian PRRSV field strains and the European vaccine Porcilis strain (Intervet); in addition, the more conserved ORF7 of 11 PRRSV strains were sequenced. The domains' prediction of their putative protein sequences was performed as well. Based on these sequences, phylogenetic trees were inferred which revealed a high degree of variability among the PRRSV Italian strains. The outcomes of the phylogenetic analysis showed that the most frequent source of infection in PRRS-positive farms (sow herds, nursery sites, fattening units) was the introduction of animals carrying a new variant and not the modification of already present variants; moreover, the integration of data from phylogenetic analysis and from the clinical and serological status of the swine herds suggested that the acclimatization program could be a valid tool to stabilize the PRRS clinical picture in farms, only when applied in combination with rigorous bio-security routine management and avoid the incoming of new PRRSV variants.

Amino Acid Sequence↗

The main protein of the aggregation factor responsible for species-specific cell adhesion in the marine sponge Microciona prolifera is highly polymorphic.

Species-specific cell recognition in sponges, the oldest living metazoans, is based on a proteoglycan-like aggregation factor. We have screened individual sponge cDNA libraries, identifying multiple related forms for the aggregation factor core protein (MAFp3). Northern blots show the presence in several human tissues of transcripts strongly binding a MAFp3-specific probe. The open reading frame for MAFp3 is not interrupted in the 5' direction, revealing variable protein sequences that contain numerous introns equally spaced. We have studied tissue histocompatibility within a sponge population, finding 100% correlation between rejection behavior and the individual-specific restriction fragment length polymorphism pattern using aggregation factor-related probes. PCR amplifications with specific primers showed that at least some of the MAFp3 forms are allelic and distribute in the population used. A pronounced polymorphism is also observed when analyzing purified aggregation factor in polyacrylamide gels. Protease digestion of the polymorphic glycosaminoglycan-containing bands indicates that glycans are also responsible for the variability. The data presented reveal a high polymorphism of aggregation factor components, which matches the elevated sponge alloincompatibility, suggesting an involvement of the cell adhesion system in sponge allogeneic reactions.

Amino Acid Sequence↗

An organism-specific method to rank predicted coding regions in Trypanosoma brucei.

Genome annotation in differently evolved organisms presents challenges because the lack of sequence-based homology limits the ability to determine the function of putative coding regions. To provide an alternative to annotation by sequence homology, we developed a method that takes advantage of unusual trypanosomatid biology and skews in nucleotide composition between coding regions and upstream regions to rank putative open reading frames based on the likelihood of coding. The method is 93% accurate when tested on known genes. We have applied our method to the full complement of open reading frames on Chromosome I of Trypanosoma brucei, and we can predict with high confidence that 226 putative coding regions are likely to be functional. Methods such as the one described here for discriminating true coding regions are critical for genome annotation when other sources of evidence for function are limited.

Animals↗

Characterization and sequence of Escherichia coli pabC, the gene encoding aminodeoxychorismate lyase, a pyridoxal phosphate-containing enzyme.

In Escherichia coli, p-aminobenzoate (PABA) is synthesized from chorismate and glutamine in two steps. Aminodeoxychorismate synthase components I and II, encoded by pabB and pabA, respectively, convert chorismate and glutamine to 4-amino-4-deoxychorismate (ADC) and glutamate, respectively. ADC lyase, encoded by pabC, converts ADC to PABA and pyruvate. We reported that pabC had been cloned and mapped to 25 min on the E. coli chromosome (J. M. Green and B. P. Nichols, J. Biol. Chem. 266:12971-12975, 1991). Here we report the nucleotide sequence of pabC, including a portion of a sequence of a downstream open reading frame that may be cotranscribed with pabC. A disruption of pabC was constructed and transferred to the chromosome, and the pabC mutant strain required PABA for growth. The deduced amino acid sequence of ADC lyase is similar to those of Bacillus subtilis PabC and a number of amino acid transaminases. Aminodeoxychorismate lyase purified from a strain harboring an overproducing plasmid was shown to contain pyridoxal phosphate as a cofactor. This finding explains the similarity to the transaminases, which also contain pyridoxal phosphate. Expression studies revealed the size of the pabC gene product to be approximately 30 kDa, in agreement with that predicted by the nucleotide sequence data and approximately half the native molecular mass, suggesting that the native enzyme is dimeric.

Amino Acid Sequence↗

Identification of the ubiD gene on the Escherichia coli chromosome.

The open reading frame at 86.7 min on the Escherichia coli chromosome, "yigC," complemented a ubiD mutant strain, AN66, indicating that yigC is the ubiD gene. The gene product, a 497-amino-acid-residue protein, showed extensive homology to the UPF 00096 family of proteins in the Swiss-Prot database.

Amino Acid Sequence↗