PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Cloning and characterization of the gene for the metalloprotease enterotoxin of Bacteroides fragilis.

The Bacteroides fragilis enterotoxin is an extracellular zinc metalloprotease that has been implicated in diarrheal disease of humans and animals. This toxin causes fluid accumulation in intestinal loops and is cytotoxic for HT-29 cells, an intestinal carcinoma cell line. Here we report the cloning and sequencing of the toxin gene (bftP). bftP is 1191 nucleotides coding for a 397 amino acid protein of 44.4 kDa. The toxin has a signal peptide of 18 amino acids that is typical of many lipoproteins followed by a 379 amino acid protoxin. The portion of the protoxin found in culture filtrates and stools begins at amino acid 212. An additional open reading frame located immediately upstream shows some sequence identity with cobra cytotoxins. If expressed, the ORF protein product could also play a role in the virulence of B. fragilis.

Amino Acid Sequence↗

mucS, a gene involved in activation of galactoglucan (EPS II) synthesis gene expression in Rhizobium meliloti.

In addition to the exopolysaccharide succinoglycan, Rhizobium meliloti can produce a galactoglucan exopolysaccharide, EPS II. The production of EPS II occurs in certain mutant strains, in strains containing extra copies of EPS II synthesis genes, or in the wild-type strain under phosphate-limiting conditions. We have identified a gene, mucS, that is in a locus required for EPS II induction by extra gene copies and by phosphate limitation, and that activates the expression of at least one other EPS II synthesis gene. mucS lies within a cluster of EPS II synthesis genes and contains an open reading frame of 190 amino acids. MucS does not show any significant similarity to known genes and may represent a new type of regulatory protein.

Amino Acid Sequence↗

Open reading frames 1 and 2 of adenovirus region E4 are conserved between human serotypes 2 and 5.

The E4 region of human adenovirus type 2 is predicted to encode seven proteins as judged from its nucleotide sequence and the pattern of differential splicing of its transcript. Two of the open reading frames (ORFs), ORF1 and ORF2, had been identified as being disrupted in the recently published sequence of the related serotype 5 virus. These ORFs were resequenced and found to be intact in the wt300 strain of adenovirus type 5.

Adenovirus E4 Proteins↗

Transcriptional analysis of the restriction and modification genes of bacteriophage P1.

Bacteriophage P1 res and mod genes encode the restriction and modification polypeptides of the Type III restriction enzyme EcoP1. Northern blot analysis using res- and mod-specific probes revealed the presence of two separate transcripts in strains harbouring the EcoP1 restriction and modification genes. Furthermore, by constructing a series of fusions with a promoter less lacZ gene, we show that both the res and mod genes are transcribed from separate promoters. A more detailed investigation of the mod promoter region revealed two promoters located some 70 and 140bp upstream from the translational start codon. In addition, another pair of promoters and a further separate promoter are located more than 500bp upstream from this start codon. Two short open reading frames are located between these distal and proximal promoter clusters. Transcription of the res gene is initiated from within the mod open reading frame from two adjacent promoters. In addition a functional promoter is located on the antisense strand close to the res promoter region. The relationship between the transcription units of the res and mod genes is discussed.

Amino Acid Sequence↗

Localization and nucleotide sequences of genes mediating site-specific recombination of the SLP1 element in Streptomyces lividans.

SLP1 is a 17.2-kbp genetic element indigenous to the Streptomyces coelicolor chromosome. During conjugation, SLP1 can undergo excision and subsequent site-specific integration into the chromosomes of recipient cells. We report here the localization, nucleotide sequences, and initial characterization of the genes mediating these recombination events. A region of SLP1 adjacent to the previously identified site of integration, attP, was found to be sufficient to promote site-specific integration of an unrelated Streptomyces plasmid. Nucleotide sequence analysis of a 2.2-kb segment of this region reveals two open reading frames that are adjacent to and transcribed toward the attP site. One of these, the 1,365-bp int gene of SLP1, encodes a predicted 50.6-kDa basic protein having substantial amino acid sequence similarity to a family of site-specific recombinases that includes the Escherichia coli bacteriophage lambda integrase. A linker insertion in the 5' end of the cloned int gene prevents integration, indicating that Int is essential for promoting integration. An open reading frame (orf61) lying immediately 5' to int encodes a predicted 7.1-kDa basic peptide showing limited sequence similarity to the excisionase (xis) genes of other site-specific recombination systems.

Amino Acid Sequence↗

EvDTree: structure-dependent substitution profiles based on decision tree classification of 3D environments.

BACKGROUND: Structure-dependent substitution matrices increase the accuracy of sequence alignments when the 3D structure of one sequence is known, and are successful e.g. in fold recognition. We propose a new automated method, EvDTree, based on a decision tree algorithm, for automatic derivation of amino acid substitution probabilities from a set of sequence-structure alignments. The main advantage over other approaches is an unbiased automatic selection of the most informative structural descriptors and associated values or thresholds. This feature allows automatic derivation of structure-dependent substitution scores for any specific set of structures, without the need to empirically determine best descriptors and parameters. RESULTS: Decision trees for residue substitutions were constructed for each residue type from sequence-structure alignments extracted from the HOMSTRAD database. For each tree cluster, environment-dependent substitution profiles were derived. The resulting structure-dependent substitution scores were assessed using a criterion based on the mean ranking of observed substitution among all possible substitutions and in sequence-structure alignments. The automatically built EvDTree substitution scores provide significantly better results than conventional matrices and similar or slightly better results than other structure-dependent matrices. EvDTree has been applied to small disulfide-rich proteins as a test case to automatically derive specific substitutions scores providing better results than non-specific substitution scores. Analyses of the decision tree classifications provide useful information on the relative importance of different structural descriptors. CONCLUSIONS: We propose a fully automatic method for the classification of structural environments and inference of structure-dependent substitution profiles. We show that this approach is more accurate than existing methods for various applications. The easy adaptation of EvDTree to any specific data set opens the way for class-specific structure-dependent substitution scores which can be used in threading-based remote homology searches.

Alanine↗

[Nucleotide sequence of two exons of the human T-lymphocyte CD4 receptor gene].

Genome DNA encoding the N-part of human CD4 gene located in the 15 kilobase (kb) Sau3a restriction fragment was cloned and nucleotide sequence of a part (3430 b.p.) of this fragment determined. Exons 2 and 3, intron 2, and partially introns 2 and 3 of this gene were located in the sequenced fragment. Six Alu repeats and open reading frames (ORFs) coding for proteins very close to C5 and C3 components of the complement were detected in this fragment.

Amino Acid Sequence↗

The open reading frame 1 of the L1Tc retrotransposon of Trypanosoma cruzi codes for a protein with apurinic-apyrimidinic nuclease activity.

The deduced amino acid sequence of the open reading frame 1 (ORF1) of the L1Tc non-site-specific non-long terminal repeat retrotransposon of Trypanosoma cruzi exhibits a significant homology with the consensus sequence of the class II family of the endonuclease apurinic-apyrimidinic (AP) proteins. The analysis of the activity of the 40-kDa recombinant protein, named NL1Tc, obtained from the expression of the L1Tc ORF1 in an Escherichia coli "in vitro" expression system revealed that the sequence codes for a protein with endonuclease activity specific for apurinic-apyrimidinic (AP) sites. Data are also presented showing that in vivo expression of the NL1Tc protein conferred viability by complementation to E. coli exonuclease III deletion mutants (BW286 strain). We propose that the biological function of the AP endonuclease activity of the NL1Tc protein may be connected with the introduction into the DNA of free 3' ends that could be used as primers for the integration, along the T. cruzi genome, of the L1Tc element and that the nicking could be a general mechanism for the retrotransposition of non-site-specific non-long terminal repeat retrotransposons.

Amino Acid Sequence↗

The sequence of a 17.5 kb DNA fragment on the left arm of yeast chromosome XI identifies the protein kinase gene ELM1, the DNA primase gene PRI2, a new gene encoding a putative histone and seven new open reading frames.

A 17.5 kb DNA fragment of chromosome XI, located between the genetic loci mif2 and mak11 was sequenced and analysed. Ten open reading frames were identified. Two of them are the previously sequenced genes ELM1 and PRI2, two (YKL253 and YKL256) show homologies to proteins from other organisms and one (YKL262) to yeast and mouse histone.

Amino Acid Sequence↗

Murine fibroblast growth factor receptor 1 gene generates multiple messenger RNAs containing two open reading frames via alternative splicing.

The arrangement of exons and introns encoding 5'-side of murine fibroblast growth factor (FGF) receptor 1 (FGFR-1) gene was mapped. A large intron with a size of 14 kb was identified between exon 1 and exon 2. In addition, all FGFR-1 subtypes including a unique variant form with 12 amino acids insertion and two amino acids deletion were observed to be able to be generated through alternative splicing. Furthermore, complete sequencing of the 5'-region of FGFR-1 mRNA revealed that a relatively large open reading frame precedes the major open reading frame encoding FGFR-1. These results indicate that FGFR-1 mRNAs are uniquely translated from an internal translation start site.

3T3 Cells↗

Nucleotide sequence analysis of a 30-kb region of the bovine herpesvirus 1 genome which exhibits a colinear gene arrangement with the UL21 to UL4 genes of herpes simplex virus.

We report the nucleotide sequence of the 19-kb HindIII fragment B of bovine herpesvirus 1 (BHV-1) DNA and adjacent parts of the HindIII A and L fragments, which together span a still completely uncharted 30-kb region located between the glycoprotein H gene and the right end of the unique long segment. The analysis revealed 17 complete open reading frames (ORFs) and 2 ORFs that were interrupted by potential splice donor and acceptor sites. All of these ORFs exhibited strong amino acid sequence homology to the gene products of other alphaherpesviruses. The BHV-1 ORFs were arranged colinearly with the prototype sequence of herpes simplex virus 1 in the range of the UL21 to UL4 genes. Colinearity was also observed with the genes of betaherpesviruses and gamma herpesviruses, although not all ORFs exhibited clear sequence homology. The possible functions of the proteins encoded within the sequenced region are assessed and features found are discussed. Unexpected findings include the following: high amino acid sequence conservation among alphaherpesviruses despite large differences in G + C content, ranging from 45% for varicella zoster virus to 72% for BHV-1; high similarity with other UL20 proteins at the predicted structural level in spite of relatively low amino acid homology; and a 2-kb open reading frame overlapping UL19 in the opposite sense and exhibiting high amino acid similarity to the same area of pseudorabies virus.

Amino Acid Sequence↗

Structure and expression of a light-inducible shoot-specific rice gene.

By differential screening of a cDNA library of two-week-old rice seedlings cDNA clones were obtained, corresponding to shoot-specific mRNAs. By sequence analysis two of these clones were found to be rbcS cDNA clones. The mRNA corresponding to a third cDNA clone (COS5) displayed an expression pattern similar to the expression pattern of rbcS genes. The mRNA (800 bases) was light-inducible and encoded by a single-copy gene. The genomic clone (GOS5) was isolated and the intron/exon structure was determined by comparing the nucleotide sequences of the mRNA and the genomic clone. The gene contains two introns. Transcription start sites were determined by S1-nuclease mapping and primer extension. The start site obtained by both methods is located 87 bp upstream of the translation start site and 23 bp downstream of TATA box-like sequence. In the 5' non-coding region motifs can be found that are homologous to sequences in promoters that are light- or UV-inducible or confer leaf-specific expression. The open reading frame present in GOS5 codes for a protein (15 kDa) that contains a putative chloroplast transit peptide and does not show any significant homology to protein sequences in the NBRF protein database.

Amino Acid Sequence↗

Molecular characterization and heterogeneity of feline immunodeficiency virus isolates.

We have molecularly cloned the complete genomic DNA of TM2 strain of feline immunodeficiency virus (FIV) isolated in Japan and compared its nucleotide and the deduced amino acid sequence with those of previously described U.S. isolates, FIV Petaluma and FIV PPR. The infectious molecular clone of FIV TM2 is different from FIV Petaluma in host cell range; the clone can not infect Crandell feline kidney cells which were permissive for FIV Petaluma. The amino acid sequence homologies, in gag, pol, and env genes between FIV TM2 and Petaluma were 90%, 87%, and 81%, respectively. On the other hand, comparative analysis of each gene between FIV Petaluma and PPR showed 96,95, and 85%, respectively. These results suggested that the genomic diversity was present among FIV strains isolated from geographically distant areas. Interestingly, tat- and rev-like short open reading frames contained inframe stop codons in the FIV Petaluma but not in the FIV TM2.

Amino Acid Sequence↗

The ZDS1 and ZDS2 proteins require the Sir3p component of yeast silent chromatin to enhance the stability of short linear centromeric plasmids.

Yeast artificial chromosome (YAC) clones of Saccharomyces cerevisiae containing a centromere, origin of replication, two telomeres and a >50 kb insert of DNA are maintained as normal yeast chromosomes. However, short linear centromeric plasmids of 10-15 kb in size (short YACs) are missegregated at a much higher frequency than long YACs or 10-15 kb circular centromeric plasmids. A search for genes that stabilized short linear centromeric plasmids when present in multiple copies per cell uncovered ZDS1, which reduced the rate at which cells lost the short YAC, increased the fraction of cells that maintained the short YAC and decreased the number of short YACs per cell. Multiple copies of ZDS2, a homolog of ZDS1, had similar effects. Genes near yeast telomeres are transcriptionally silenced by the recruitment of proteins encoded by the SIR2, SIR3 and SIR4 genes (Sir2p, Sir3p and Sir4p). Multiple copies of ZDS1 and ZDS2 caused an increase in telomeric silencing. In addition, ZDS1 and ZDS2 both required the open reading frame encoding the N-terminal 174 amino acids of Sir3p to stabilize short YACs. Thus, the short YAC stability assay revealed a silencing-independent function for the Sir3p N-terminus. Two-hybrid analysis indicated that Zds1p and Zds2p interact with Sir2p, Sir3p, Sir4p or the yeast telomere binding protein Rap1p. Deletion of both ZDS1 and ZDS2 made short YACs, but not a 100 kb YAC, extremely unstable and also caused a 70 bp increase in the length of the telomeric TG1-3 repeats. These data indicate that short YACs can be stabilized by trans-acting factors and suggest that the proteins encoded by ZDS1 and ZDS2 alter short YAC stability by interacting with proteins that function at the telomere.

Adaptor Proteins, Signal Transducing↗

Cloning and nucleotide sequence of the major capsid protein from Lactococcus lactis ssp. cremoris bacteriophage F4-1.

The gene (mcp) coding for the major capsid protein (MCP) of the Lactococcus lactis ssp. cremoris bacteriophage F4-1 has been cloned and its nucleotide sequence determined. The mcp gene was localized, by Western blotting with rabbit antiserum against intact bacteriophage, within a 3.3-kb HindIII-Spe I fragment and the sequence of the entire region determined. The 35-kDa MCP is coded for by a 905-bp open reading frame preceded by a putative ribosome-binding site. Deletion analysis and N-terminal sequencing of the MCP confirmed the identification of the gene coding for this bacteriophage MCP.

Amino Acid Sequence↗

A putative bioactive conformation for the altered peptide ligand of myelin basic protein and inhibitor of experimental autoimmune encephalomyelitis [Arg91, Ala96] MBP87-99.

[Arg(91), Ala(96)] MBP(87-99) is an altered peptide ligand (APL) of myelin basic protein (MBP), shown to actively inhibit experimental autoimmune encephalomyelitis (EAE), which is studied as a model of multiple sclerosis (MS). The APL has been rationally designed by substituting two of the critical residues for recognition by the T-cell receptor. A conformational analysis of the APL has been sought using a combination of 2D NOESY nuclear magnetic resonance (NMR) experiments and detailed molecular dynamics (MD) calculations, in order to comprehend the stereoelectronic requirements for antagonistic activity, and to propose a putative bioactive conformation based on spatial proximities of the native peptide in the crystal structure. The proposed structure presents backbone similarity with the native peptide especially at the N-terminus, which is important for major histocompatibility complex (MHC) binding. Primary (Val(87), Phe(90)) and secondary (Asn(92), Ile(93), Thr(95)) MHC anchors occupy the same region in space, whereas T-cell receptor (TCR) contacts (His(88), Phe(89)) have different orientation between the two structures. A possible explanation, thus, of the antagonistic activity of the APL is that it binds to MHC, preventing the binding of myelin epitopes, but it fails to activate the TCR and hence to trigger the immunologic response. NMR experiments coupled with theoretical calculations are found to be in agreement with X-ray crystallography data and open an avenue for the design and synthesis of novel peptide restricted analogues as well as peptide mimetics that rises as an ultimate goal.

Amino Acid Sequence↗

Kinetoplast DNA minicircles of Leishmania donovani express a protein product.

We describe an unprecedented finding of an open reading frame present in the variable region in one of the minicircle sequence classes of a human pathogenic strain of Leishmania donovani (MHOM/IN/90/RMRI 68) which is transcribed and translated. The encoded protein showed homologies to known transport proteins.

Amino Acid Sequence↗

Clonal deletion of V beta 14-bearing T cells in mice transgenic for mammary tumour virus.

Autoreactive T lymphocytes are clonally deleted during maturation in the thymus. Deletion of T cells expressing particular receptor V beta elements is controlled by poorly defined autosomal dominant genes. A gene has now been identified by expression of transgenes in mice which causes deletion of V beta 14+ T cells. The gene lies in the open reading frame of the long terminal repeat of the mouse mammary tumour virus.

Amino Acid Sequence↗