PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Mfold web server for nucleic acid folding and hybridization prediction.

The abbreviated name, 'mfold web server', describes a number of closely related software applications available on the World Wide Web (WWW) for the prediction of the secondary structure of single stranded nucleic acids. The objective of this web server is to provide easy access to RNA and DNA folding and hybridization software to the scientific community at large. By making use of universally available web GUIs (Graphical User Interfaces), the server circumvents the problem of portability of this software. Detailed output, in the form of structure plots with or without reliability information, single strand frequency plots and 'energy dot plots', are available for the folding of single sequences. A variety of 'bulk' servers give less information, but in a shorter time and for up to hundreds of sequences at once. The portal for the mfold web server is http://www.bioinfo.rpi.edu/applications/mfold. This URL will be referred to as 'MFOLDROOT'.

Base Sequence↗

The protein identification resource (PIR).

The Protein Identification Resource consists of an integrated computer system composed of a number of protein and nucleic acid sequence databases and software designed for the identification and analysis of protein sequences and their corresponding coding sequences. The PIR serves the scientific community through on-line access, distributing magnetic tapes, and performing off-line sequence identification services for researchers.

Amino Acid Sequence↗

Gene transfer in the evolution of parasite nucleotide biosynthesis.

Nucleotide metabolic pathways provide numerous successful targets for antiparasitic chemotherapy, but the human pathogen Cryptosporidium parvum thus far has proved extraordinarily refractory to classical treatments. Given the importance of this protist as an opportunistic pathogen afflicting immunosuppressed individuals, effective treatments are urgently needed. The genome sequence of C. parvum is approaching completion, and we have used this resource to critically assess nucleotide biosynthesis as a target in C. parvum. Genomic analysis indicates that this parasite is entirely dependent on salvage from the host for its purines and pyrimidines. Metabolic pathway reconstruction and experimental validation in the laboratory further suggest that the loss of pyrimidine de novo synthesis is compensated for by possession of three salvage enzymes. Two of these, uridine kinase-uracil phosphoribosyltransferase and thymidine kinase, are unique to C. parvum within the phylum Apicomplexa. Phylogenetic analysis suggests horizontal gene transfer of thymidine kinase from a proteobacterium. We further show that the purine metabolism in C. parvum follows a highly streamlined pathway. Salvage of adenosine provides C. parvum's sole source of purines. This renders the parasite susceptible to inhibition of inosine monophosphate dehydrogenase, the rate-limiting enzyme in the multistep conversion of AMP to GMP. The inosine 5' monophosphate dehydrogenase inhibitors ribavirin and mycophenolic acid, which are already in clinical use, show pronounced anticryptosporidial activity. Taken together, these data help to explain why widely used drugs fail in the treatment of cryptosporidiosis and suggest more promising targets.

Adenosine↗

Likelihood analysis of the chalcone synthase genes suggests the role of positive selection in morning glories (Ipomoea).

Chalcone synthase (CHS) is a key enzyme in the biosynthesis of flavonoides, which are important for the pigmentation of flowers and act as attractants to pollinators. Genes encoding CHS constitute a multigene family in which the copy number varies among plant species and functional divergence appears to have occurred repeatedly. In morning glories (Ipomoea), five functional CHS genes (A-E) have been described. Phylogenetic analysis of the Ipomoea CHS gene family revealed that CHS A, B, and C experienced accelerated rates of amino acid substitution relative to CHS D and E. To examine whether the CHS genes of the morning glories underwent adaptive evolution, maximum-likelihood models of codon substitution were used to analyze the functional sequences in the Ipomoea CHS gene family. These models used the nonsynonymous/synonymous rate ratio (omega = d(N)/ d(S)) as an indicator of selective pressure and allowed the ratio to vary among lineages or sites. Likelihood ratio test suggested significant variation in selection pressure among amino acid sites, with a small proportion of them detected to be under positive selection along the branches ancestral to CHS A, B, and C. Positive Darwinian selection appears to have promoted the divergence of subfamily ABC and subfamily DE and is at least partially responsible for a rate increase following gene duplication.

Acyltransferases↗

Proteome profiling of Populus euphratica Oliv. upon heat stress.

BACKGROUND AND AIMS: Populus euphratica is a light-demanding species ecologically characterized as a pioneer. It grows in shelter belts along riversides, being part of the natural desert forest ecosystems in China and Middle Eastern countries. It is able to survive extreme temperatures, drought and salt stress, marking itself out as an important plant species to study the mechanisms responsible for survival of woody plants under heat stress. METHODS: Heat effects were evaluated through electrolyte leakage on leaf discs, and LT(50) was determined to occur above 50 degrees C. Protein accumulation profiles of leaves from young plants submitted to 42/37 degrees C for 3 d in a phytotron were determined through 2D-PAGE, and a total of 45 % of up- and downregulated proteins were detected. Matrix-assisted laser desorption ionization-time of flight (MALDI-TOF)/TOF analysis, combined with searches in different databases, enabled the identification of 82 % of the selected spots. KEY RESULTS: Short-term upregulated proteins are related to membrane destabilization and cytoskeleton restructuring, sulfur assimilation, thiamine and hydrophobic amino acid biosynthesis, and protein stability. Long-term upregulated proteins are involved in redox homeostasis and photosynthesis. Late downregulated proteins are involved mainly in carbon metabolism. CONCLUSIONS: Moderate heat response involves proteins related to lipid biogenesis, cytoskeleton structure, sulfate assimilation, thiamine and hydrophobic amino acid biosynthesis, and nuclear transport. Photostasis is achieved through carbon metabolism adjustment, a decrease of photosystem II (PSII) abundance and an increase of PSI contribution to photosynthetic linear electron flow. Thioredoxin h may have a special role in this process in P. euphratica upon moderate heat exposure.

Cell Membrane↗

IMGT/LIGM-DB, the IMGT comprehensive database of immunoglobulin and T cell receptor nucleotide sequences.

IMGT/LIGM-DB is the IMGT comprehensive database of immunoglobulin (IG) and T cell receptor (TR) nucleotide sequences from human and other vertebrate species. It was created in 1989 by LIGM, Montpellier, France and is the oldest and the largest database of IMGT. IMGT/LIGM-DB includes all germline (non-rearranged) and rearranged IG and TR genomic DNA (gDNA) and complementary DNA (cDNA) sequences published in generalist databases. IMGT/LIGM-DB allows searches from the Web interface according to biological and immunogenetic criteria through five distinct modules depending on the user interest. For a given entry, nine types of display are available including the IMGT flat file, the translation of the coding regions and the analysis by the IMGT/V-QUEST tool. IMGT/LIGM-DB distributes expertly annotated sequences. The annotations hugely enhance the quality and the accuracy of the distributed detailed information. They include the sequence identification, the gene and allele classification, the constitutive and specific motif description, the codon and amino acid numbering, and the sequence obtaining information, according to the main concepts of IMGT-ONTOLOGY. They represent the main source of IG and TR gene and allele knowledge stored in IMGT/GENE-DB and in the IMGT reference directory. IMGT/LIGM-DB is freely available at http://imgt.cines.fr.

Animals↗

Optimum growth temperature and the base composition of open reading frames in prokaryotes.

The purine-loading index (PLI) is the difference between the numbers of purines (A+G) and pyrimidines (T+C) per kilobase of single-stranded nucleic acid. By purine-loading their mRNAs organisms may minimize unnecessary RNA-RNA interactions and prevent inadvertent formation of "self" double-stranded RNA. Since RNA-RNA interactions have a strong entropy-driven component, this need to minimize should increase as temperature increases. Consistent with this, we report for 550 prokaryotic species that optimum growth temperature is related to the average PLI of open reading frames. With increasing temperature prokaryotes tend to acquire base A and lose base C, while keeping bases T and G relatively constant. Accordingly, while the PLI increases, the (G+C)% decreases. The previously observed positive correlation between (G+C)% and optimum growth temperature, which applies to RNA species whose structure is of major importance for their function (ribosomal and transfer RNAs) does not apply to mRNAs, and hence is unlikely to apply generally to genomic DNA.

Bacteria↗

Identification and quantification of glycerolipids in cotton fibers: reconciliation with metabolic pathway predictions from DNA databases.

The lipid profiles of cotton fiber cells were determined from total lipid extracts of elongating and maturing cotton fiber cells to see whether the membrane lipid composition changed during the phases of rapid cell elongation or secondary cell wall thickening. Total FA content was highest or increased during elongation and was lower or decreased thereafter, likely reflecting the assembly of the expanding cell membranes during elongation and the shift to membrane maintenance (and increase in secondary cell wall content) in maturing fibers. Analysis of lipid extracts by electrospray ionization and tandem MS (ESI-MS/MS) revealed that in elongating fiber cells (7-10 d post-anthesis), the polar lipids-PC, PE, PI, PA, phosphatidylglycerol, monogalactosyldiacylglycerol, digalactosyldiacylglycerol, and phosphatidylglycerol-were most abundant. These same glycerolipids were found in similar proportions in maturing fiber cells (21 dpa). Detailed molecular species profiles were determined by ESI-MS/MS for all glycerolipid classes, and ESI-MS/MS results were consistent with lipid profiles determined by HPLC and ELSD. The predominant molecular species of PC, PE, PI, and PA was 34:3 (16:0, 18:3), but 36:6 (18:3,18:3) also was prevalent. Total FA analysis of cotton lipids confirmed that indeed linolenic (18:3) and palmitic (16:0) acids were the most abundant FA in these cell types. Bioinformatics data were mined from cotton fiber expressed sequence tag databases in an attempt to reconcile expression of lipid metabolic enzymes with lipid metabolite data. Together, these data form a foundation for future studies of the functional contribution of lipid metabolism to the development of this unusual and economically important cell type.

Chromatography, High Pressure Liquid↗

Large-scale identification and analysis of genome-wide single-nucleotide polymorphisms for mapping in Arabidopsis thaliana.

Genetic markers such as single nucleotide polymorphisms (SNPs) are essential tools for positional cloning, association, or quantitative trait locus mapping and the determination of genetic relationships between individuals. We identified and characterized a genome-wide set of SNP markers by generating 10,706 expressed sequence tags (ESTs) from cDNA libraries derived from 6 different accessions, and by analysis of 606 sequence tagged sites (STS) from up to 12 accessions of the model flowering plant Arabidopsis thaliana. The cDNA libraries for EST sequencing were made from individuals that were stressed by various means to enrich for transcripts from genes expressed under such conditions. SNPs discovered in these sequences may be useful markers for mapping genes involved in interactions with the biotic and abiotic environment. The STS loci are distributed randomly over the genome. By comparison with the Col-0 genome sequence, we identified a total of 8051 SNPs and 637 insertion/deletion polymorphisms (InDel). Analysis of STS-derived SNPs shows that most SNPs are rare, but that it is possible to identify intermediate frequency framework markers that can be used for genetic mapping in many different combinations of accessions. A substantial proportion of SNPs located in ORFs caused a change of the encoded amino acid. A comparison of the density of our SNP markers among accessions in both the EST and STS datasets, revealed that Cvi-0 is the most divergent accession from Col-0 among the 12 accessions studied. All of these markers are freely available via the internet.

Arabidopsis↗

Functional second genes generated by retrotransposition of the X-linked ribosomal protein genes.

We have identified a new class of ribosomal protein (RP) genes that appear to have been retrotransposed from X-linked RP genes. Mammalian ribosomes are composed of four RNA species and 79 different proteins. Unlike RNA constituents, each protein is typically encoded by a single intron- containing gene. Here we describe functional autosomal copies of the X-linked human RP genes, which we designated RPL10L (ribosomal protein L10-like gene), RPL36AL and RPL39L after their progenitors. Because these genes lack introns in their coding regions, they were likely retrotransposed from X-linked genes. The identities between the retrotransposed genes and the original X-linked genes are 89-95% in their nucleotide sequences and 92-99% in their amino acid sequences, respectively. Northern blot and PCR analyses revealed that RPL10L and RPL39L are expressed only in testis, whereas RPL36AL is ubiquitously expressed. Although the role of the autosomal RP genes remains unclear, they may have evolved to compensate for the reduced dosage of X-linked RP genes.

5' Flanking Region↗

Discrimination of Mycobacterium tuberculosis complex bacteria using novel VNTR-PCR targets.

The lack of a convenient high-resolution strain-typing method has hampered the application of molecular epidemiology to the surveillance of bacteria of the Mycobacterium tuberculosis complex, particularly the monitoring of strains of Mycobacterium bovis. With the recent availability of genome sequences for strains of the M. tuberculosis complex, novel PCR-based M. tuberculosis-typing methods have been developed, which target the variable-number tandem repeats (VNTRs) of minisatellite-like mycobacterial interspersed repetitive units (MIRUs), or exact tandem repeats (ETRs). This paper describes the identification of seven VNTR loci in M. tuberculosis H37Rv, the copy number of which varies in other strains of the M. tuberculosis complex. Six of these VNTRs were applied to a panel of 100 different M. bovis isolates, and their discrimination and correlation with spoligotyping and an established set of ETRs were assessed. The number of alleles varied from three to seven at the novel VNTR loci, which differed markedly in their discrimination index. There was positive correlation between spoligotyping, ETR- and VNTR-typing. VNTR-PCR discriminates well between M. bovis strains. Thirty-three allele profiles were identified by the novel VNTRs, 22 for the ETRs and 29 for spoligotyping. When VNTR- and ETR-typing results were combined, a total of 51 different profiles were identified. Digital nomenclature and databasing were intuitive. VNTRs were located both in intergenic regions and annotated ORFs, including PPE (novel glycine-asparigine-rich) proteins, a proposed source of antigenic variation, where VNTRs potentially code repeating amino acid motifs. VNTR-PCR is a valuable tool for strain typing and for the study of the global molecular epidemiology of the M. tuberculosis complex. The novel VNTR targets identified in this study should additionally increase the power of this approach.

Alleles↗

A comparison of group II introns of plastid tRNALysUUU genes encoding maturase protein.

All higher plant plastid genomes have six classes of tRNA genes containing introns. One of those is the tRNALysUUU gene, which encodes maturase protein. In the case of liverwort species from the genus Porella and mosses from the genus Plagiomnium, the maturase coding gene (matK) represents a truncated form of other plant matK genes: several subdomains of the reverse transcriptase-like domain and so-called domain X are not present in these ORFs. These ORFs probably represent pseudogenes of the matK gene. The analysis of codon usage within the matK gene revealed the presence of strong A/T pressure. The use of codons with the third letter being U or A varies from 71-93%. The comparison of maturase amino acid sequences at the family level shows a high identity between species. However, when liverwort and angiosperm maturase sequences are compared, the percentage of identity drops dramatically. The calculated values of the number of nucleotide substitutions vary considerably, even when liverwort species are compared pairwise. The phenetic tree of relationships between plant species on the basis of tRNALysUUU intron sequences concur with the generally accepted plant phylogeny.

Algorithms↗

[Phylogenic analysis of the Sox gene family of vertebrate].

Sox genes of vertebrate are highly evolutionarily conserved, which encode different transcriptional factors involved in various developmental processes. Sox family is characterized by a sequence-specific DNA binding HMG-box containing about 79 amino acids. To realize the complexity of genes of the Sox family in the structures, functions and their evolutionary relationships, in the present study by utilizing all available complete nucleotide/protein sequence data of vertebrate Sox genes, we performed multi-sequence comparison and construction of phylogenic tree, and the grouping of Sox family members and the pattern of their molecular evolution was also analyzed.

Animals↗

Characterization of a strain of Sphingobacterium sp. and its degradation to herbicide mefenacet.

A bacterium (designated strain Y1) degrading acetanilide herbicide mefenacet was isolated from aerobic sludge. Based on the analyses of partial 16S rRNA gene, cellular fatty acid and BIOLOG-GN, and general physiological and biochemical characteristics, strain Y1 was identified as Sphingobacterium multivolum. Strain Y1 was able to degrade mefenacet used as sources of carbon and energy. Degradation of mefenacet was accompanied by producing the metabolites N-methylaniline and an unidentified compound with molecular weight 205, indicating a metabolic pathway of mefenacet initiated by hydrolysis of amido bond.

Acetamides↗

TassDB: a database of alternative tandem splice sites.

Subtle alternative splice events at tandem splice sites are frequent in eukaryotes and substantially increase the complexity of transcriptomes and proteomes. We have developed a relational database, TassDB (TAndem Splice Site DataBase), which stores extensive data about alternative splice events at GYNGYN donors and NAGNAG acceptors. These splice events are of subtle nature since they mostly result in the insertion/deletion of a single amino acid or the substitution of one amino acid by two others. Currently, TassDB contains 114 554 tandem splice sites of eight species, 5209 of which have EST/mRNA evidence for alternative splicing. In addition, human SNPs that affect NAGNAG acceptors are annotated. The database provides a user-friendly interface to search for specific genes or for genes containing tandem splice sites with specific features as well as the possibility to download large datasets. This database should facilitate further experimental studies and large-scale bioinformatics analyses of tandem splice sites. The database is available at http://helios.informatik.uni-freiburg.de/TassDB/.

Alternative Splicing↗

Large-scale correlation of DNA accession numbers to the cDNAs in the FANTOM full-length mouse cDNA clone set.

Oligonucleotide-based microarrays, such as GeneChip, are widely used to determine the large-scale gene expression profiles. However, GeneChip only provides information on the identity of the molecules, and the investigator must obtain each cDNA clone for further analyses. In this study, we devised a program which enables us to correlate a large number of DNA accession numbers to the FANTOM (functional annotation of the mouse) full-length mouse cDNA clone set, and made a correlative table between mouse GeneChip clones and FANTOM clones. This allows easy identification of the corresponding FANTOM clone for each GeneChip clone, even if the sequence of the GeneChip clone does not directly match the FANTOM clone. Using this table, for example, a large number of in situ hybridization probes can be synthesized easily, because the FANTOM clones are flanked by T3/T7 promoters on both ends. In addition, we further developed a program which retrieves the amino acid sequence (AA Seq) for each clone, even for the FANTOM clones that lack the AA Seq description, and classifies the proteins automatically. As an example, we devised a correlation table with predictions of the secretory or transmembrane molecules. The correlation table is useful for a large-scale screening of molecules involved in cell-cell communication in various biological processes. The full correlation table for the GeneChip clones is available at http://www.kjm.keio.ac.jp/past/55/3/correlation_table1.html.

Animals↗

ATP-dependent L-cysteine:1D-myo-inosityl 2-amino-2-deoxy-alpha-D-glucopyranoside ligase, mycothiol biosynthesis enzyme MshC, is related to class I cysteinyl-tRNA synthetases.

Mycothiol is a novel thiol produced only by actinomycetes and is the major low molecular weight thiol in mycobacteria. The mycothiol biosynthetic pathway has been postulated to involve ATP-dependent ligation of L-cysteine (Cys) with 1D-myo-inosityl 2-amino-2-deoxy-alpha-D-glucopyranoside; GlcN-Ins) catalyzed by MshC to produce Cys-GlcN-Ins. The ligase activity was purified approximately 2400-fold from Mycobacterium smegmatis and two proteins of slightly different M(r) approximately 47000 were identified with MshC activity. The N-terminal sequence of the smaller protein revealed that it was coded by a gene in the databases for M. smegmatis and M. tuberculosis previously designated as cysS2. The larger protein was coded by the same gene in M. smegmatis but included an eight amino acid N-terminal extension involving a different start codon. The ligase was found to have K(m) values of 40 +/- 3 and 72 +/- 9 microM for Cys and GlcN-Ins, respectively. The cysS2 gene was thought to encode a second cysteinyl-tRNA synthetase in addition to cysS but the present results indicate that cysS2 is actually the mshC gene encoding ATP-dependent Cys:GlcN-Ins ligase.

Adenosine Triphosphate↗

Comparative analysis of the base biases at the gene terminal portions in seven eukaryote genomes.

Adenine nucleotides have been found to appear preferentially in the regions after the initiation codons or before the termination codons of bacterial genes. Our previous experiments showed that AAA and AAT, the two most frequent second codons in Escherichia coli, significantly enhance translation efficiency. To determine whether such a characteristic feature of base frequencies exists in eukaryote genes, we performed a comparative analysis of the base biases at the gene terminal portions using the proteomes of seven eukaryotes. Here we show that the base appearance at the codon third positions of gene terminal regions is highly biased in eukaryote genomes, although the codon third positions are almost free from amino acid preference. The bias changes depending on its position in a gene, and is characteristic of each species. We also found that bias is most outstanding at the second codon, the codon after the initiation codon. NCN is preferred in every genome; in particular, GCG is strongly favored in human and plant genes. The presence of the bias implies that the base sequences at the second codon affect translation efficiency in eukaryotes as well as bacteria.

Animals↗