PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Thermodynamic databases for proteins and protein-nucleic acid interactions.

Thermodynamic data regarding proteins and their interactions are important for understanding the mechanisms of protein folding, protein stability, and molecular recognition. Although there are several structural databases available for proteins and their complexes with other molecules, databases for experimental thermodynamic data on protein stability and interactions are rather scarce. Thus, we have developed two electronically accessible thermodynamic databases. ProTherm, Thermodynamic Database for Proteins and Mutants, contains numerical data of several thermodynamic parameters of protein stability, experimental methods and conditions, along with structural, functional, and literature information. ProNIT, Thermodynamic Database for Protein-Nucleic Acid Interactions, contains thermodynamic data for protein-nucleic acid binding, experimental conditions, structural information of proteins, nucleic acids and the complex, and literature information. These data have been incorporated into 3DinSight, an integrated database for structure, function, and properties of biomolecules. A WWW interface allows users to search for data based on various conditions, with different display and sorting options, and to visualize molecular structures and their interactions. These thermodynamic databases, together with structural databases, help researchers gain insight into the relationship among structure, function, and thermodynamics of proteins and their interactions, and will become useful resources for studying proteins in the postgenomic era.

DNA↗

Evidence of positive Darwinian selection in putative meningococcal vaccine antigens.

Meningococcal meningitidis is a life-threatening disease. In Europe and the United States the majority of cases are caused by virulent meningococcal strains belonging to serogroup B. Presently there is no effective vaccine against serogroup B strains, as traditional vaccine antigens such as polysaccharide capsules are unusable as they lead to autoimmunity. The year 2000 saw the publication of the complete genome of Neisseria meningitidis MC58, a virulent serogroup B bacterium. Working in conjunction with the sequencing project, researchers endeavored to locate highly conserved membrane-associated proteins that elicit an immune response. It is hoped that these proteins will provide a basis for novel vaccines against serogroup B strains. A number of potential vaccine antigens have been located and are presently in phase I clinical trials. Recently many reports pertaining to the evidence of positive Darwinian selection in membrane proteins of pathogens have been reported. This study utilized in silico methods to test for evidence of historical positive Darwinian selection in seven such vaccine candidates. We found that two of these proteins show signatures of adaptive evolution, while the remaining proteins show evidence of strong purifying selection. This has significant implications for the design of a vaccine against serogroup B strains, as it has been shown that vaccines that target epitopes that are under strong purifying selection are better than those that target variable epitopes.

Amino Acid Sequence↗

Analysis of the molecular evolutionary history of the ascorbate peroxidase gene family: inferences from the rice genome.

Ascorbate peroxidase (APx) is a class I peroxidase that catalyzes the conversion of H(2)O(2) to H(2)O and O(2) using ascorbate as the specific electron donor. This enzyme has a key function in scavenging reactive oxygen species (ROS) and the protection against toxic effects of ROS in higher plants, algae, and Euglena. Here we report the identification of an APx multigene family in rice and propose a molecular evolutionary relationship between the diverse APx isoforms. In rice, the APx gene family has eight members, which encode two cytosolic, two putative peroxisomal, and four chloroplastic isoforms, respectively. Phylogenetic analyses were conducted using all APx protein sequences available in the NCBI databases. The results indicate that the different APx isoforms arose by a complex evolutionary process involving several gene duplications. The structural organization of APx genes also reflects this process and provides evidence for a close relationship among proteins located in the same subcellular compartment. A molecular evolutionary pathway, in which cytosolic and peroxisomal isoforms diverged early from chloroplastic ones, is proposed.

Amino Acid Sequence↗

Structural and functional characterization of AtPTR3, a stress-induced peptide transporter of Arabidopsis.

A T-DNA tagged mutant line of Arabidopsis thaliana, produced with a promoter trap vector carrying a promoterless gus (uidA) as a reporter gene, showed GUS induction in response to mechanical wounding. Cloning of the chromosomal DNA flanking the T-DNA revealed that the insert had caused a knockout mutation in a PTR-type peptide transporter gene named At5g46050 in GenBank, here renamed AtPTR3. The gene and the deduced protein were characterized by molecular modelling and bioinformatics. Molecular modelling of the protein with fold recognition identified 12 transmembrane spanning regions and a large loop between the sixth and seventh helices. The structure of AtPTR3 resembled the other PTR-type transporters of plants and transporters in the major facilitator superfamily. Computer analysis of the AtPTR3 promoter suggested its expression in roots, leaves and seeds, complex hormonal regulation and induction by abiotic and biotic stresses. The computer-based hypotheses were tested experimentally by exposing the mutant plants to amino acids and several stress treatments. The AtPTR3 gene was induced by the amino acids histidine, leucine and phenylalanine in cotyledons and lower leaves, whereas a strong induction was obtained in the whole plant upon exposure to salt. Furthermore, the germination frequency of the mutant line was reduced on salt-containing media, suggesting that the AtPTR3 protein is involved in stress tolerance in seeds during germination.

Amino Acids↗

Multiple gene organization of pufferfish Fugu rubripes tropomyosin isoforms and tissue distribution of their transcripts.

The Japanese pufferfish, torafugu (Fugu rubripes), has a haploid genome of about 400 Mb in size, which has been sequenced to approximately 90% coverage. Here we identified six Fugu tropomyosin (TPM) gene sequences by using the BLASTN program and the sequence of the white croaker TPM1 gene in our collection against the draft assembly of the Fugu genomic sequence database. TPM2, TPM3 and TPM4 genes were identified together with a set of two potentially duplicated genes of TPM1 (TPM1-1 and TPM1-2) as described in our previous report and TPM4 (TPM4-1 and TPM4-2) newly found in this study. The expression patterns of these Fugu TPM genes were determined by reverse transcription polymerase chain reaction (RT-PCR). A phylogenetic tree was constructed using the deduced amino acid sequences, which were encoded by the exons common to all vertebrate TPM genes. This indicated that the Fugu TPM1 and TPM4 genes had resulted from a gene duplication in the fish evolutionary lineage.

Alternative Splicing↗

Identification and phylogenetic analyses of the protein arginine methyltransferase gene family in fish and ascidians.

Protein arginine methyltransferases (PRMT) involved in the regulations of signal transduction, protein subcellular localization, and transcription have been mostly studied in mammals and yeast. In this study orthologues of eight human PRMT genes (PRMT1-7 and HRMT1L3) were identified in both puffer fish Fugu rubripes and zebrafish Danio rerio. The fish PRMT genes appear to be conserved with their mammalian orthologues at the levels of amino acid sequences as well as genomic structures. All vertebrate PRMT genes contain 10-16 coding exons except PRMT6 that contains only one coding exon. Western blot analyses of zebrafish tissue extracts confirmed the expression of some PRMT proteins in zebra fish. We further identified six PRMT members (PRMT1, 3-7) in an invertebrate chordate Ciona intestinalis. Genomic structures of the PRMT orthologues are no more conserved in the ascidians, as PRMT3 and PRMT5 contain only one coding exon while PRMT6 contains six exons. PRMT2 and HRMT1L3 that are missing in Ciona appear to be vertebrate-specific. HRMT1L3 is a PRMT1 paralogue with highly conserved sequences and exact exon junctions, whereas the PRMT2 orthologues are very diverged. Different PRMT orthologues are likely to evolve at different rates and the PRMT1 orthologues appear to be most conserved through evolution. Furthermore, phylogenetic analyses using the core regions of various PRMT genes show that PRMT5 with the type II PRMT activity is separated in one branch. All other PRMT genes including PRMT1, 2, 3, 4, 6, 7 and HRMT1L3 clustered in the other branch, probably represent the genes for the type I activity.

Amino Acid Sequence↗

Structure, expression, and mapping of two nodule-specific genes identified by mining public soybean EST databases.

Numerous nodule-specific genes, which are involved in the root nodule development and function, have been known and are still being discovered. Here, we reported the structure, expression, and genetic map location of two novel nodule-specific genes. First, two EST groups, one obtained from a nodule library and the other from all aboveground tissue libraries, were clustered with regard to in silico expression profiles. We compiled a pool of 103 putative nodule-specific sequence clusters. Then, two representative ESTs were selected for further experimental analyses. According to the full-length cDNA sequences, one was an EST of a novel nodule-specific polygalacturonase gene, GmPGN, and the other was an EST of a new short nodule-specific gene, GmEKN. The results of expression analyses of the GmPGN cDNAs indicated that GmPGN expression was not detectable in any of the soybean tissues except in the nodule tissue and may be regulated via alternative splicing. GmEKN expression was the most strongly detected in the nodule. The predicted GmEKN protein is both glutamic acid- and lysine-rich, and is also highly hydrophilic. Genetic mapping located GmPGN near the known quantitative trait locus conferring resistance to soybean cyst nematode on soybean molecular linkage group (MLG) B1, and GmEKN on MLG A2. These results provide useful information for the use of these genes in research on the orchestration of numerous genes in nodule development and function.

Amino Acid Sequence↗

Discovery of differentially expressed genes in human breast cancer using subtracted cDNA libraries and cDNA microarrays.

Identifying novel and known genes that are differentially expressed in breast cancer has important implications in understanding the biology of breast tumorigenesis and developing new diagnostic and therapeutic agents. In this study we have combined two powerful technologies, PCR-based cDNA subtraction and cDNA microarray, as a high throughput methodology designed to identify cDNA clones that are breast tumor- and tissue-specific and are overexpressed in breast tumors. Approximately 2000 cDNA clones generated from the subtracted breast tumor library were arrayed on the microarray chips. The arrayed target cDNAs were then hybridized with 30 pairs of fluorescent-labeled cDNA probes generated from breast tumors and normal tissues to determine the tissue distribution and tumor specificity. cDNA clones showing overexpression in breast tumors by microarray were further analysed by DNA sequencing, GenBank and EST database searches, and quantitative real time PCR. We identified several known genes, including mammaglobin, cytokeratin 19, fibronectin, and hair-specific type II keratin, which have previously been shown to be overexpressed in breast tumors and may play an important role in the malignance of breast. We also discovered B726P which appears to be an isoform of NY-BR-1, a breast tissue-specific gene. Two additional clones discovered, B709P and GABA(A) receptor pi subunit, were not previously described for their overexpression profile in breast tumors. Thus, combining PCR-based cDNA subtraction and cDNA microarray allowed for an efficient way to identify and validate genes with elevated mRNA expression levels in breast cancer that may potentially be involved in breast cancer progression. These differentially expressed genes may be of potential utility as therapeutic and diagnostic targets for breast cancer.

Amino Acid Sequence↗

Distribution patterns of over-represented k-mers in non-coding yeast DNA.

MOTIVATION: Over-represented k-mers in genomic DNA regions are often of particular biological interest. For example, over-represented k-mers in co-regulated families of genes are associated with the DNA binding sites of transcription factors. To measure over-representation, we introduce a statistical background model based on single-mismatches, and apply it to the pooled 500 bp ORF Upstream Regions (USRs) of yeast. More importantly, we investigate the context and spatial distribution of over-represented k-mers in yeast USRs. RESULTS: Single and double-stranded spatial distributions of most over-represented k-mers are highly non-random, and predominantly cluster into a small number of classes that are robust with respect to over-representation measures. Specifically, we show that the three most common distribution patterns can be related to DNA structure, function, and evolution and correspond to: (a) homologous ORF clusters associated with sharply localized distributions; (b) regulatory elements associated with a symmetric broad hill-shaped distribution in the 50-200 bp USR; and (c) runs of As, Ts, and ATs associated with a broad hill-shaped distribution also in the 50-200 bp USR, with extreme structural properties. Analysis of over-representation, homology, localization, and DNA structure are essential components of a general data-mining approach to finding biologically important k-mers in raw genomic DNA and understanding the 'lexicon' of regulatory regions.

Amino Acid Motifs↗

Remote homology detection: a motif based approach.

MOTIVATION: Remote homology detection is the problem of detecting homology in cases of low sequence similarity. It is a hard computational problem with no approach that works well in all cases. RESULTS: We present a method for detecting remote homology that is based on the presence of discrete sequence motifs. The motif content of a pair of sequences is used to define a similarity that is used as a kernel for a Support Vector Machine (SVM) classifier. We test the method on two remote homology detection tasks: prediction of a previously unseen SCOP family and prediction of an enzyme class given other enzymes that have a similar function on other substrates. We find that it performs significantly better than an SVM method that uses BLAST or Smith-Waterman similarity scores as features.

Algorithms↗

Identification of rDNA-specific non-LTR retrotransposons in Cnidaria.

Ribosomal RNA genes are abundant repetitive sequences in most eukaryotes. Ribosomal DNA (rDNA) contains many insertions derived from mobile elements including non-long terminal repeat (non-LTR) retrotransposons. R2 is the well-characterized 28S rDNA-specific non-LTR retrotransposon family that is distributed over at least 4 bilaterian phyla. R2 is a large family sharing the same insertion specificity and classified into 4 clades (R2-A, -B, -C, and -D) based on the N-terminal domain structure and the phylogeny. There is no observation of horizontal transfer of R2; therefore, the origin of R2 dates back to before the split between protostomes and deuterostomes. Here, we in silico identified 1 R2 element from the sea anemone Nematostella vectensis and 2 R2-like retrotransposons from the hydrozoan Hydra magnipapillata. R2 from N. vectensis was inserted into the 28S rDNA like other R2, but the R2-like elements from H. magnipapillata were inserted into the specific sequence in the highly conserved region of the 18S rDNA. We designated the Hydra R2-like elements R8. R8 is inserted at 37 bp upstream from R7, another 18S rDNA-specific retrotransposon family. There is no obvious sequence similarity between targets of R2 and R8, probably because they recognize long DNA sequences. Domain structure and phylogeny indicate that R2 from N. vectensis is the member of the R2-D clade, and R8 from H. magnipapillata belongs to the R2-A clade despite its different sequence specificity. These results suggest that R2 had been generated before the split between cnidarians and bilaterians and that R8 is a retrotransposon family that changed its target from the 28S rDNA to the 18S rDNA.

Amino Acid Sequence↗

NADP-dependent isocitrate dehydrogenase from the halophilic archaeon Haloferax volcanii: cloning, sequence determination and overexpression in Escherichia coli.

A gene encoding NADP-dependent Ds-threo-isocitrate dehydrogenase was isolated from Haloferax volcanii genomic DNA by using a combination of polymerase chain reaction and screening of a lambda EMBL3 library. Analysis of the nucleotide sequence revealed an open reading frame of 1260 bp encoding a protein of 419 amino acids with 45837 Da molecular mass. This sequence is highly similar to previously sequenced isocitrate dehydrogenases. In the alignment of the amino acid sequences with those from several archaeal and mesophilic NADP-dependent isocitrate dehydrogenases, the residues involved in dinucleotide binding and isocitrate binding are well conserved. We have developed methods for the expression in Escherichia coli and purification of the enzyme from H. volcanii. This expression was carried out in E. coli as inclusion bodies using the cytoplasmic expression vector pET3a. The enzyme was refolded by solubilisation in 8 M urea followed by dilution into a buffer containing EDTA, MgCl(2) and 3 M NaCl. Maximal activity was obtained after several hours incubation at room temperature.

Amino Acid Sequence↗

Random sequencing of Paramecium somatic DNA.

We report a random survey of 1 to 2% of the somatic genome of the free-living ciliate Paramecium tetraurelia by single-run sequencing of the ends of plasmid inserts. As in all ciliates, the germ line genome of Paramecium (100 to 200 Mb) is reproducibly rearranged at each sexual cycle to produce a somatic genome of expressed or potentially expressed genes, stripped of repeated sequences, transposons, and AT-rich unique sequence elements limited to the germ line. We found the somatic genome to be compact (>68% coding, estimated from the sequence of several complete library inserts) and to feature uniformly small introns (18 to 35 nucleotides). This facilitated gene discovery: 722 open reading frames (ORFs) were identified by similarity with known proteins, and 119 novel ORFs were tentatively identified by internal comparison of the data set. We determined the phylogenetic position of Paramecium with respect to eukaryotes whose genomes have been sequenced by the distance matrix neighbor-joining method by using random combined protein data from the project. The unrooted tree obtained is very robust and in excellent agreement with accepted topology, providing strong support for the quality and consistency of the data set. Our study demonstrates that a random survey of the somatic genome of Paramecium is a good strategy for gene discovery in this organism.

Amino Acid Sequence↗

Characterization of a ComE3 homologue essential for DNA transformation in Helicobacter pylori.

To find genes involved in natural competence in Helicobacter pylori, we used a bioinformatics database search and found two transformation-related open reading frames (ORFs): a comE3 homologue (HP1361 ORF) of Bacillus subtilis and a comL homologue (HP1378 ORF) of Neisseria gonorrhoeae. We failed to obtain an HP1378 ORF knockout mutant, while an HP1361 ORF knockout mutant was obtained by transposon shuttle mutagenesis. The DNA transformation abilities of both natural transformation and electroporation were severely impaired (frequency, <10(-9)) in the HP1361(-) mutant. Complementation with a pHel2 vector carrying the HP1361 ORF restored the capabilities of natural competence (to a frequency of 4.21 x 10(-7)) and electroporation (to 3.62 x 10(-7)). The HP1361(-) mutant showed impairment in DNA binding and uptake. The results suggest that HP1361 is a comE3 homologue and is required for DNA binding and uptake during DNA transformation.

Amino Acid Sequence↗

SeqVISTA: a graphical tool for sequence feature visualization and comparison.

BACKGROUND: Many readers will sympathize with the following story. You are viewing a gene sequence in Entrez, and you want to find whether it contains a particular sequence motif. You reach for the browser's "find in page" button, but those darn spaces every 10 bp get in the way. And what if the motif is on the opposite strand? Subsequently, your favorite sequence analysis software informs you that there is an interesting feature at position 13982-14013. By painstakingly counting the 10 bp blocks, you are able to examine the sequence at this location. But now you want to see what other features have been annotated close by, and this information is buried several screenfuls higher up the web page. RESULTS: SeqVISTA presents a holistic, graphical view of features annotated on nucleotide or protein sequences. This interactive tool highlights the residues in the sequence that correspond to features chosen by the user, and allows easy searching for sequence motifs or extraction of particular subsequences. SeqVISTA is able to display results from diverse sequence analysis tools in an integrated fashion, and aims to provide much-needed unity to the bioinformatics resources scattered around the Internet. Our viewer may be launched on a GenBank record by a single click of a button installed in the web browser. CONCLUSION: SeqVISTA allows insights to be gained by viewing the totality of sequence annotations and predictions, which may be more revealing than the sum of their parts. SeqVISTA runs on any operating system with a Java 1.4 virtual machine. It is freely available to academic users at http://zlab.bu.edu/SeqVISTA.

Amino Acid Sequence↗

Identification of a novel gene family that includes the interferon-inducible human genes 6-16 and ISG12.

BACKGROUND: The human 6-16 and ISG12 genes are transcriptionally upregulated in a variety of cell types in response to type I interferon (IFN). The predicted products of these genes are small (12.9 and 11.5 kDa respectively), hydrophobic proteins that share 36% overall amino acid identity. Gene disruption and over-expression studies have so far failed to reveal any biochemical or cellular roles for these proteins. RESULTS: We have used in silico analyses to identify a novel family of genes (the ISG12 gene family) related to both the human 6-16 and ISG12 genes. Each ISG12 family member codes for a small hydrophobic protein containing a conserved ~80 amino-acid motif (the ISG12 motif). So far we have detected 46 family members in 25 organisms, ranging from unicellular eukaryotes to humans. Humans have four ISG12 genes: the 6-16 gene at chromosome 1p35 and three genes (ISG12(a), ISG12(b) and ISG12(c)) clustered at chromosome 14q32. Mice have three family members (ISG12(a), ISG12(b1) and ISG12(b2)) clustered at chromosome 12F1 (syntenic with human chromosome 14q32). There does not appear to be a murine 6-16 gene. On the basis of phylogenetic analyses, genomic organisation and intron-alignments we suggest that this family has arisen through divergent inter- and intra-chromosomal gene duplication events. The transcripts from human and mouse genes are detectable, all but two (human ISG12(b) and ISG12(c)) being upregulated in response to type I IFN in the cell lines tested. CONCLUSIONS: Members of the eukaryotic ISG12 gene family encode a small hydrophobic protein with at least one copy of a newly defined motif of approximately 80 amino-acids (the ISG12 motif). In higher eukaryotes, many of the genes have acquired a responsiveness to type I IFN during evolution suggesting that a role in resisting cellular or environmental stress may be a unifying property of all family members. Analysis of gene-function in higher eukaryotes is complicated by the possibility of functional redundancy between family-members. Genetic studies in organisms (e.g. Dictyostelium discoideum) with just one family member so far identified may be particularly helpful in this respect.

Amino Acid Sequence↗

Seven different genes encode a diverse mixture of isoforms of Bet v 1, the major birch pollen allergen.

BACKGROUND: Pollen of the European white birch (Betula pendula, syn. B. verrucosa) is an important cause of hay fever. The main allergen is Bet v 1, member of the pathogenesis-related class 10 (PR-10) multigene family. To establish the number of PR-10/Bet v 1 genes and the isoform diversity within a single tree, PCR amplification, cloning and sequencing of PR-10 genes was performed on two diploid B. pendula cultivars and one interspecific tetraploid Betula hybrid. Sequences were attributed to putative genes based on sequence identity and intron length. Information on transcription was derived by comparison with homologous cDNA sequences available in GenBank/EMBL/DDJB. PCR-cloning of multigene families is accompanied by a high risk for the occurrence of PCR recombination artifacts. We screened for and excluded these artifacts, and also detected putative artifact sequences among database sequences. RESULTS: Forty-four different PR-10 sequences were recovered from B. pendula and assigned to thirteen putative genes. Sequence homology suggests that three genes were transcribed in somatic tissue and seven genes in pollen. The transcription of three other genes remains unknown. In total, fourteen different Bet v 1-type isoforms were identified in the three cultivars, of which nine isoforms were entirely new. Isoforms with high and low IgE-reactivity are encoded by different genes and one birch pollen grain has the genetic background to produce a mixture of isoforms with varying IgE-reactivity. Allergen diversity is even higher in the interspecific tetraploid hybrid, consistent with the presence of two genomes. CONCLUSION: Isoforms of the major birch allergen Bet v 1 are encoded by multiple genes, and we propose to name them accordingly. The present characterization of the Bet v 1 genes provides a framework for the screening of specific Bet v 1 genes among other B. pendula cultivars or Betula species, and for future breeding for trees with a reduced allergenicity. Investigations towards sensitization and immunotherapy should anticipate that patients are exposed to a mixture of Bet v 1 isoforms of different IgE-reactivity, even if pollen originates from a single birch tree.

Allergens↗

Database and analysis system for cDNA clones obtained from full-length enriched cDNA libraries.

We have developed an efficient sequence-analysis system and a database system for clones obtained from full-length enriched cDNA libraries made by using the oligo-capping method. We developed a semi-automatic analysis system for 5'- and 3'-end sequences. It pre-processes raw sequences (vector cut and accurate-sequence region extraction), clusters the sequences, searches for similarities through public databases, annotates completeness of clones and analyzes the ORFs in the sequences. Newly developed or improved programs are used in each step. A new program, ESTiMateFull is used to evaluate and to predict the sequence-fullness based on comparisons with mRNA and EST sequences, respectively. The ATGpr program is used to predict sequence-fullness based on statistical information. The combination of full-length enriched cDNA clones and ATGpr fullness prediction resulted in 70% accuracy in the specificity and the sensitivity of the fullness predictions. For the ORFs predicted by the ATGpr, the signal peptides are predicted and a motif search is performed by our new system. We also developed a program that assembles our sequences with dbEST sequences and developed a system to retrieve clones by the characteristics of the ORFs. As keywords, combination of various results of the analyses can be used for retrieval. And various results such as ORF features and database search results can be shown on the same screen by multiple displays. Full-length clones having interesting functions can thus be retrieved efficiently by using this system.

Amino Acid Sequence↗