PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Genetic variability and evolution of the Schistosoma genome analysed by using random amplified polymorphic DNA markers.

The usefulness of random amplified polymorphic DNA markers (RAPD) was assayed in an attempt to discriminate among species, strains and individuals within the genus Schistosoma. Depending on the species, 40-50 arbitrary decamer oligonucleotides were used as primers to amplify total DNA by the polymerase chain reaction (PCR). An important polymorphism was observed among 5 species, allowing a phylogenetic tree to be outlined. These differences can be used for rapid and accurate identification. A limited but easily detectable polymorphism was revealed among 3 strains of a single species (Schistosoma mansoni). Minor differences were observed among individuals of a single strain. A RAPD marker allows sexual discrimination between individuals from the terminal spined-egg species group. Although a limited number of strains have been examined, the results already indicate clearly that RAPD markers constitute a powerful tool for the analysis of genetic variability. This new tool will considerably extend the information available from morphology, isozyme and limited restriction fragment length polymorphism data and opens the way to genetic analysis of these species.

Animals↗

Cloning and nucleotide sequence of the cDNA encoding human erythrocyte-specific AMP deaminase.

The nucleotide sequence of cDNA encoding human erythrocyte AMP deaminase has been determined by screening of human spleen cDNA library and by utilizing polymerase chain reaction (PCR) techniques. The 3.7 kb cDNA contains an open reading frame of 2301 bp which encodes 767 amino acids chain resulting in 89 kDa protein. The polyadenylation consensus signal (5'-AATAAA) located at 1212 bp 3' downstream from the stop codon. The homologies to human and rat muscle-specific AMP deaminases showed 64.1% and 65.2% identities, respectively, at the nucleotide level in the area of open reading frame, and 60.2% and 59.8% similarities at the deduced amino acid level.

AMP Deaminase↗

Characterization of the genes encoding a phosphate-regulated two component sensory system in the marine cyanobacterium Synechococcus sp. WH7803.

An oligomer probe was designed to detect the presence of a putative phoB gene in the genome of the marine, phycoerythrin-containing cyanobacterium Synechococcus sp. WH7803. A 2.2 kb PstI fragment, identified using this probe, was cloned and the complete nucleotide sequence determined. The fragment contained two open reading frames encoding polypeptides which display all the sequence features expected of the response regulator and histidine protein kinase elements of a two component sensory system. Northern analysis confirmed that transcription of these genes was induced by phosphate limitation. On the basis of the sequence similarities and the regulation of their transcription by the availability of inorganic phosphate (Pi) these open reading frames were designated as phoB and phoR, respectively.

Amino Acid Sequence↗

Identification and molecular characterization of two tandemly located flagellin genes from Aeromonas salmonicida A449.

Two tandemly located flagellin genes, flaA and flaB, with 79% nucleotide sequence identity were identified in Aeromonas salmonicida A449. The fla genes are conserved in typical and atypical strains of A. salmonicida, and they display significant divergence at the nucleotide level from the fla genes of the motile species Aeromonas hydrophila and Aeromonas veronii biotype sobria. flaA and flaB encode unprocessed flagellins with predicted Mrs of 32,351 and 32,056, respectively. When cloned under the control of the Ptac promoter, flaB was highly expressed when induced in Escherichia coli DH5alpha, and the FlaB protein was detectable even in the uninduced state. In flaA clones containing intact upstream sequence, FlaA was barely detectable when uninduced and poorly expressed on induction. The A. salmonicida flagellins are antigenically cross-reactive with the A. hydrophila TF7 flagellin(s) and evolutionarily closely related to the flagellins of Pseudomonas aeruginosa and Vibrio anguillarum. Electron microscopy showed that A. salmonicida A449 expresses unsheathed polar flagella at an extremely low frequency under normal laboratory growth conditions, suggesting the presence of a full complement of genes whose products are required to make flagella; e.g., immediately downstream of flaA and flaB are open reading frames encoding FlaG and FlaH homologs.

Aeromonas↗

Sequencing of a 23 kb fragment from Saccharomyces cerevisiae chromosome VI.

Plasmid clone gapB and lambda phage clone 4682, which contain fragments of Saccharomyces cerevisiae chromosome VI, were analysed. A 23 kb sequence was determined and ten open reading frames (ORFs) were revealed. Among them, five ORFs were identical to five yeast genes (SEC4, MSH4, SPB4, DEG1 and NIC96), two were identical to transposable elements (TYA and TYB), one (gapBorfF003) was highly homologous to a yeast expressed sequence tag, and another (4682orfF002) was predicted to be a nuclear protein. Sequence data have been submitted to DDBJ/EMBL/GenBank data library under Accession Number D44604 (clone gapB) and D44600 (clone 4682), respectively.

Amino Acid Sequence↗

Sequencing and analysis of 51 kb on the right arm of chromosome XV from Saccharomyces cerevisiae reveals 30 open reading frames.

We have sequenced a region of 51 kb of the right arm from chromosome XV of Saccharomyces cerevisiae. The sequence contains 30 open reading frames (ORFs) of more than 100 amino acid residues. Thirteen new genes have been identified. Thirteen ORFs correspond to known yeast genes. One delta element and one tRNA gene were identified. Upstream of the RPO31 gene, encoding the largest subunit of RNA polymerase III, lies a Abf1p binding site. The nucleotide sequence data reported in this paper are available in the EMBL, GenBank and DDBJ nucleotide sequence databases under the Accession Number X90518.

Base Sequence↗

The first intron of human c-fms proto-oncogene contains a processed pseudogene (RPL7P) for ribosomal protein L7.

During sequence analysis of the first intron of the human c-fms oncogene, we identified an open reading frame encoding the ribosomal protein L7 (RPL7). The presence of this sequence within intron 1 of the c-fms gene was confirmed by Southern blot hybridization and by sequence analysis of two independent cosmid clones (cos2-e and cos1-22) that span the human genomic c-fms locus. The RPL7 sequence was detected in a region of sequence overlapped by the cos2-e and cos1-22 cosmid clones but oriented opposite to the c-fms gene. We demonstrated that the sequence is identical to the full-length RPL7 cDNA sequence, but lacks any recognizable introns, has a 30-bp poly(A) tail, and is bracketed by two perfect direct repeats of 14 bp. We also showed that despite the fact that the 5' flanking region of the RPL7 sequence contains a potential TATA box upstream of an intact open reading frame, this pseudogene (RPL7P) is not actively transcribed.

Base Sequence↗

Bluetongue virus evolution: sequence analyses of the genomic S1 segments and major core protein VP7.

The S1 segments, encoding the group-specific antigen, VP7, from the five United States prototype BTV serotypes were cloned as full-length entities. The nucleotide and deduced amino acid sequences of segment S1 of BTV-2 were determined and compared with BTV-10, -11, -13, and -17, completing the sequencing of this cognate gene segment from all five US BTV serotypes. Each segment is 1156 bp long and contains an open reading frame encoding the 349-amino acid VP7 protein. Most (greater than 94%) of the amino acids of VP7 among the serotypes are conserved, including the location (position 255) of a single lysine residue. Secondary structure analyses of VP7 predict a putative eight-stranded beta-barrel between amino acid positions 150 and 250, a structure similar to that observed in ssRNA viruses. The S1 genes are flanked by conserved 5' and 3' noncoding regions. Stem-loop structures are predicted at the 3' end of each gene (nucleotide positions 1058-1097). The S1 segments of BTV-2, -10, -11, and -17 have greater than 93% of the nucleotides conserved, while less than 80% of their bases are identical with BTV-13. Analyses of nucleotide mismatches in each codon position of the VP7 open reading frame, transition frequencies, and evolutionary distances show that of the five, BTV-13 is the most distantly related and that BTV-10 and -17 are the most closely related serotypes. Evolutionary distance calculations of segment L2 from BTV-10, -11, and -17 concur with these observations. Comparison of this relationship with hybridization data of segment M3, which codes for VP5, suggests that BTV-17 has evolved by a combination of genetic drift and genomic reassortment. The data also indicate that the five US BTV serotypes are derived from two distinct gene pools. Evolution distances were used to estimate an evolution rate of 2.2 x 10(-3) nucleotide substitution/site/year for BTV segment S1. This rate is similar to the genes of retroviruses and implies an absence of RNA polymerase proofreading activity for dsRNA viruses.

Amino Acid Sequence↗

Characterization of a retrotransposon-like element from Entamoeba histolytica.

The protozoan parasite Entamoeba histolytica is the causative agent of amoebiasis. The genome organization of this organism is not well understood. We had earlier reported the presence of a multicopy sequence, HMc, in E. histolytica. Subsequent analysis showed that HMc is a member of a retrotransposon family that we have named the E. histolytica retrotransposon-like element (EhRLE). Four other members of this family have been characterized. The EhRLE family is distributed across all chromosomes of the parasite. There are 140 copies, which show minor sequence variation with respect to one another (2--4% from the consensus sequence). From a sequence analysis of five members of the EhRLE family, the complete EhRLE unit is estimated to be 4086 bp in length. It has a 27-mer inverted repeat at its ends. A pairwise comparison with sequences in the database showed a highly significant match of a part of EhRLE with reverse transcriptases (RT), especially those encoded by non-long terminal repeat retrotransposons. There are stop codons in all the five EhRLEs, but a continuous open reading frame of 464 amino acids could be reconstructed by comparing the sequences of several EhRLEs. The reconstructed sequence showed a much better identity with RT as compared with any of the original EhRLE sequences. The non-pathogenic species, Entamoeba dispar, also contains this element, with 85% sequence identity with EhRLE. The data suggest that EhRLE may be a retrotransposon, but many of its members are probably nonfunctional due to the accumulation of mutations.

Amino Acid Sequence↗

Identification, expression, and characterization of Escherichia coli guanine deaminase.

Using the human cDNA sequence corresponding to guanine deaminase, the Escherichia coli genome was scanned using the Basic Local Alignment Search Tool (BLAST), and a corresponding 439-residue open reading frame of unknown function was identified as having 36% identity to the human protein. The putative gene was amplified, subcloned into the pMAL-c2 vector, expressed, purified, and characterized enzymatically. The 50.2-kDa protein catalyzed the conversion of guanine to xanthine, having a K(m) of 15 microM with guanine and a k(cat) of 3.2 s(-1). The bacterial enzyme shares a nine-residue heavy metal binding site with human guanine deaminase, PG[FL]VDTHIH, and was found to contain approximately 1 mol of zinc per mol of subunit of protein. The E. coli guanine deaminase locus is 3' from an open reading frame which shows homology to a bacterial purine base permease.

Amino Acid Sequence↗

[Analysis of the nucleotide sequence of a fragment (92-100%) of the CELO avian adenovirus genome].

The nucleotide sequence of 92-100% of the adenovirus CELO (FAV1), strain Phelps, genome has been determined. The computer analysis of the sequences revealed a ClaI site methylated by m*Ecodam. A recognition site for XbaI and two sites for PstI, not found in the corresponding genome of CELO, strain Ote, have been determined. Three extensive (more than 100 amino acid residues) open reading frames exist, coding for the polypeptides with molecular weights of 31.5, 19.3 and 14.5 kD (276, 178 and 128 amino acids, respectively). Some shorter open reading frames have been detected as well within the sequences studied.

Amino Acid Sequence↗

Nucleotide sequence analysis of an 8887 bp region of the left arm of yeast chromosome XIV, encompassing the centromere sequence.

The nucleotide sequencing of 8887 bp of the left arm of chromosome XIV is described. The sequence includes the centromeric region. Both strands were sequenced with an average redundancy of 5.09 per base pair. The overall G+C content is 37.3% (39.2% for putative coding regions versus 32.5% for non-coding regions). Six open reading frames (ORFs) greater than 100 amino acids were detected, all of which are completely confined to the 8.9 kbp region. Codon frequencies of the six ORFs agree with codon usage in Saccharomyces cerevisiae and all show the characteristics of low-level expressed genes. Comparison of the translated sequences with protein sequences in data bases suggests the presence of two ORFs (N2014 and N2007) encoding ribosomal proteins, the latter of which is the previously sequenced MRP7 gene. Another ORF (N2012) could encode a membrane-associated protein since it contains secretory signal sequence and two presumed transmembrane helices. This protein might be involved in mitochondrial energy transfer. ORF N2016 is immediately adjacent to the centromere, suggesting that it corresponds to the SPO1 gene, which is very tightly linked to the centromere at the left arm side of chromosome XIV (Mortimer et al., 1989).

Base Composition↗

The difficulty of identifying genes in anonymous vertebrate sequences.

The identification of genes in newly determined vertebrate genomic sequences can range from a trivial to an impossible task. In a statistical preamble, we show how "insignificant" are the individual features on which gene identification can be rigorously based: promoter signals, splice sites, open reading frames, etc. The practical identification of genes is thus ultimately a tributary of their resemblance to those already present in sequence databases, or incorporated into training sets. The inherent conservatism of the currently popular methods (database similarity search, GRAIL) will greatly limit our capacity for making unexpected biological discoveries from increasingly abundant genomic data. Beyond a very limited subset of trivial cases, the automated interpretation (i.e. without experimental validation) of genomic data, is still a myth. On the other hand, characterizing the 60,000 to 100,000 genes thought to be hidden in the human genome by the mean of individual experiments is not feasible. Thus, it appears that our only hope of turning genome data into genome information must rely on drastic progresses in the way we identify and analyse genes in silico.

Amino Acid Sequence↗

The class 2 selenophosphate synthetase gene of Drosophila contains a functional mammalian-type SECIS.

Synthesis of monoselenophosphate, the selenium donor required for the synthesis of selenocysteine (Sec) is catalyzed by the enzyme selenophosphate synthetase (SPS), first described in Escherichia coli. SPS homologs were identified in archaea, mammals and Drosophila. In the latter, however, an amino acid replacement is present within the catalytic domain and lacks selenide-dependent SPS activity. We describe the identification of a novel Drosophila homolog, Dsps2. The open reading frame of Dsps2 mRNA is interrupted by an UGA stop codon. The 3'UTR contains a mammalian-like Sec insertion sequence which causes translational readthrough in both transfected Drosophila cells and transgenic embryos. Thus, like vertebrates, Drosophila contains two SPS enzymes one with and one without Sec in its catalytic domain. Our data indicate further that the selenoprotein biosynthesis machinery is conserved between mammals and fly, promoting the use of Drosophila as a genetic tool to identify components and mechanistic features of the synthesis pathway.

3' Untranslated Regions↗

The sequence of an 8 kb segment on the left arm of chromosome II from Saccharomyces cerevisiae identifies five new open reading frames of unknown functions, two tRNA genes and two transposable elements.

The DNA sequence of an 8079 bp ClaI fragment located at 40 kb from the centromere on the left arm of chromosome II from Saccharomyces cerevisiae has been determined. Sequence analysis reveals five new open reading frames, tRNA(Gly) and tRNA(Leu) genes as well as sigma and truncated delta elements. The disruption of the three larger open reading frames shows that they are not essential for mitotic growth.

Amino Acid Sequence↗

Characterization of the Acinetobacter plasmid, pRAY, and the identification of regulatory sequences upstream of an aadB gene cassette on this plasmid.

Primer extension analyses carried out to identify the transcription start site of an aadB gene, which is part of a gene cassette recombined at a secondary site on an Acinetobacter plasmid, pRAY, suggest that transcription control signals in Acinetobacter are similar but not identical to their counterparts in Escherichia coli. pRAY was sequenced. An AT-rich region, containing eight copies of the consensus sequence, AAAAAATAT, previously shown to be present in the origins of replication of other Acinetobacter plasmids, was predicted to be the origin of pRAY. The translation product of one of the 10 open reading frames identified on pRAY shows homology to the mobilization protein, MbeA.

Acinetobacter↗

Characterization of the radish mitochondrial orfB locus: possible relationship with male sterility in Ogura radish.

The orfB locus of the normal (fertile) and Ogura (male-sterile) radish mitochondrial genomes has been characterized in order to determine if this region, which has previously been correlated with cytoplasmic male sterility (CMS) in Brassica napus cybrids (Bonhomme et al. 1991; Temple et al. 1992), could also be involved in radish CMS. In normal radish, orfB is expressed as a 600-nucleotide (nt) transcript. In Ogura radish, orfB is present as the second gene of a 1200-nt transcript that also contains a 138-codon open reading frame (orf138). Sequences showing similarity to orf138 are present in normal radish, but are not expressed.

Amino Acid Sequence↗