PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Overseer: a nucleotide sequence searching tool.

Overseer is a computer program that searches databases of nucleic acid sequences for objects of interest to the user. Such objects may consist of any number of simpler building blocks such as repeats, palindromes or stem-loops, strings of particular bases with or without mismatches, etc. Written in standard Pascal, this program runs under Unix and VMS and should also run under other operating systems. A simple interface allows the user to generate interactively a file containing a description of the target to be found. The searching program runs non-interactively, processing the information from the file and searching the sequences. The results are output to a file. Search capabilities are quite flexible and the code is designed to be modified. Since the framework of the program is simple, adding new modules to search for new target types as the need arises is possible.

Algorithms↗

New insights into the evolutionary history of type 1 rhodopsins.

Type 1 (archaeal) rhodopsins and related rhodopsin-like proteins had been described in a few halophile archaea, gamma-proteobacteria, a single cyanobacteria, some fungi, and a green alga. In exhaustive database searches, we detected rhodopsin-related sequences derived not only from additional fungal species but also from organisms belonging to three groups in which opsins had hitherto not been described: the alpha-proteobacterium Magnetospirillum magnetotacticum, the cryptomonad alga Guillardia theta, and the dinoflagellate Pyrocystis lunula. Putative plant and human type 1 rhodopsin sequences found in the databases are demonstrated to be contaminants of fungal origin. However, a highly diverged sequence supposedly from the plant Oryza sativa was found that is, together with the Pyrocystis sequence, quite similar to gamma-proteobacterial rhodopsins. These close relationships suggest that at least one horizontal gene transfer event involving rhodopsin genes occurred between prokaryotes and eukaryotes. Alternative hypotheses to explain the current phylogenetic range of type 1 rhodopsins are suggested. The broader phylogenetic range found is compatible with an ancient origin of type 1 rhodopsins, their patchy distribution being caused by losses in multiple lineages. However, the possibility of ancient horizontal transfer events between distant relatives cannot be dismissed.

Amino Acid Sequence↗

Mitochondrial genetic codes evolve to match amino acid requirements of proteins.

Mitochondria often use genetic codes different from the standard genetic code. Now that many mitochondrial genomes have been sequenced, these variant codes provide the first opportunity to examine empirically the processes that produce new genetic codes. The key question is: Are codon reassignments the sole result of mutation and genetic drift? Or are they the result of natural selection? Here we present an analysis of 24 phylogenetically independent codon reassignments in mitochondria. Although the mutation-drift hypothesis can explain reassignments from stop to an amino acid, we found that it cannot explain reassignments from one amino acid to another. In particular--and contrary to the predictions of the mutation-drift hypothesis--the codon involved in such a reassignment was not rare in the ancestral genome. Instead, such reassignments appear to take place while the codon is in use at an appreciable frequency. Moreover, the comparison of inferred amino acid usage in the ancestral genome with the neutral expectation shows that the amino acid gaining the codon was selectively favored over the amino acid losing the codon. These results are consistent with a simple model of weak selection on the amino acid composition of proteins in which codon reassignments are selected because they compensate for multiple slightly deleterious mutations throughout the mitochondrial genome. We propose that the selection pressure is for reduced protein synthesis cost: most reassignments give amino acids that are less expensive to synthesize. Taken together, our results strongly suggest that mitochondrial genetic codes evolve to match the amino acid requirements of proteins.

Amino Acids↗

Novel chaperonins in a prokaryote.

Group II chaperonins belong to the Hsp60 family occurring in archaea and eukaryotes. The archaeal chaperonins build the thermosome, which is similar to the eukaryotic CCT (chaperonin-containing TCP-1). Eukaryotes have eight subunits, and up until now, it was thought that archaea had between one and three subunits, depending on the species. We now report two novel subunits, termed Hsp60-4 and Hsp60-5, in the archaeon Methanosarcina acetivorans, which also has Hsp60-1, Hsp60-2, and Hsp60-3 with orthologs in Methanosarcinae. Hsp60-4 and Hsp60-5 occur only in M. acetivorans, which makes this organism unique in that it has the highest number of chaperonin subunits ever described for an archaeon. Evolutionary analysis suggests that either Hsp60-4 or Hsp60-5 paralogs have arisen by gene duplication with vastly increased accepted substitution rates or that they represent ancestral types found only in this species.

Amino Acid Sequence↗

Unusual usage of AGG and TTG codons in humans and their viruses.

Prior analysis on human protein-coding DNA sequences has identified local base composition as the primary predictor of synonymous codon usage. However, in many organisms, codon usage is influenced by natural selection, particularly for efficient expression of functional gene products. Because viruses are expected to evolve codon usage in the context of their host's molecular machinery, their genomes provide another window into the forces that guide their host's molecular evolution. Factor analysis was performed on codon usage of 16,654 genes annotated in Build 34 of the human genome, and the primary factor was correlated strongly with local base composition. However, two codons, AGG and TTG, rose in frequency as all other C- and G-ending codons decreased in frequency. These two codons were the only C- or G-ending codons with usages that negatively correlated with gene expression. Variation among viruses in codon usage also strongly reflects variation in base composition and, again, AGG and TTG decrease in frequency as all other C- and G-ending codons increase in frequency. It appears that usages of these two codons can not be explained by local compositional biases, implying a more direct role of natural selection on codon usage in humans.

Amino Acids↗

Gene expression levels influence amino acid usage and evolutionary rates in endosymbiotic bacteria.

Most endosymbiotic bacteria have extremely reduced genomes, accelerated evolutionary rates, and strong AT base compositional bias thought to reflect reduced efficacy of selection and increased mutational pressure. Here, we present a comparative study of evolutionary forces shaping five fully sequenced bacterial endosymbionts of insects. The results of this study were three-fold: (i) Stronger conservation of high expression genes at not just nonsynonymous, but also synonymous, sites. (ii) Variation in amino acid usage strongly correlates with GC content and expression level of genes. This pattern is largely explained by greater conservation of high expression genes, leading to their higher GC content. However, we also found indication of selection favoring GC-rich amino acids that contrasts with former studies. (iii) Although the specific nutritional requirements of the insect host are known to affect gene content of endosymbionts, we found no detectable influence on substitution rates, amino acid usage, or codon usage of bacterial genes involved in host nutrition.

AT Rich Sequence↗

Effects of a static magnetic field on cell growth and gene expression in Escherichia coli.

Escherichia coli cultures exposed to a 300mT static magnetic field (SMF) were studied in order to analyse possible induced changes in cellular growth and gene expression. Biomass was evaluated by visible-light spectrometry and gene expression analyses were carried out by use of RNA arbitrarily primed PCR. The bacterial strain XL-1Blue, cultivated in traditional and modified Luria-Bertani medium, was exposed to SMF generated by permanent neodymium magnetic disks. The results show alterations induced by SMF in terms of increased cell proliferation and changes in gene expression compared with control groups. Three cDNAs were found to be expressed only in the exposed cells, whereas one cDNA was more expressed in the controls. One clone, expressed only in the exposed cells, corresponds to a putative transposase. This is of particular interest in that it suggests that exposure to a magnetic field may stimulate transposition activity.

Amino Acid Sequence↗

Biochemical and molecular characterization of flavonoid 7-sulfotransferase from Arabidopsis thaliana.

Flavonoid compounds play important roles as flower pigments, stress metabolites formed in response to UV, during pollen germination and for polar auxin transport (Trends Plant Sci. 1 (1996) 377). Flavonoid sulfate esters are common in plants, especially the Asteraceae; however, due to the lack of information regarding the factors that regulate their accumulation, their exact role remains to be elucidated. The biosynthesis of flavonol sulfate esters is catalyzed by a number of position specific flavonol sulfotransferases (STs). An Arabidopsis thaliana database search has allowed us to identify and classify 18 putative ST coding sequences. We report here the cloning and characterization of the AtST3a member of this family that is expressed at early stages of seedling development and in the inflorescence stem and siliques of mature plants. The recombinant AtST3a protein exhibits strict specificity for position 7 of flavonoids. In contrast to previously characterized flavonol 7-ST from Flaveria bidentis that sulfonates only flavonol disulfates, AtST3a was found to accept as substrates a number of flavonols and flavone aglycones, as well as their monosulfate esters. The discovery of a flavonol ST from A. thaliana suggests that flavonol sulfates are more widely distributed than originally believed and this model plant could be used to study their biological significance.

Amino Acid Sequence↗

Widespread occurrence of serpin genes with multiple reactive centre-containing exon cassettes in insects and nematodes.

By applying homology-search and text-mining programs we have found that the Drosophila serine protease inhibitor (serpin) gene sp4 harbours four reactive centre-coding exons. The mutually exclusive use of these cassettes in combination with alternatively selectable exons at the 5'-end or in the 3'-untranslated region of the gene allows generation of more than ten different transcripts, all of which are expressed in Drosophila embryos. These transcripts may code for eight different Sp4 protein isoforms with different biological functions, which - dependent on the splice pattern - either may be secreted, reside in the endoplasmic reticulum, or may be located in the cytoplasm. An examination revealed the presence of two serpin genes, each coding for two or three likely alternative reactive centre exon cassettes, respectively, also in the Caenorhabditis elegans genome. The occurrence of such serpin genes in some groups of metazoa reflects a parsimonious way to enlarge the adaptive ability of these organisms to cope with a plethora of different serine and cysteine proteases.

Alternative Splicing↗

Genome-wide analysis of the GRAS gene family in rice and Arabidopsis.

Members of the GRAS gene family encode transcriptional regulators that have diverse functions in plant growth and development such as gibberellin signal transduction, root radial patterning, axillary meristem formation, phytochrome A signal transduction, and gametogenesis. Bioinformatic analysis identified 57 and 32 GRAS genes in rice and Arabidopsis, respectively. Here, we provide a complete overview of this gene family, describing the gene structure, gene expression, chromosome localization, protein motif organization, phylogenetic analysis, and comparative analysis between rice and Arabidopsis. Phylogenetic analysis divides the GRAS gene family into eight subfamilies, which have distinct conserved domains and functions. Both genome/segmental duplication and tandem duplication contributed to the expansion of the GRAS gene family in the rice and Arabidopsis genomes. The existence of GRAS-like genes in bryophytes suggests that GRAS is an ancient family of transcription factors, which arose before the appearance of land plants over 400 million years ago.

Amino Acid Motifs↗

Newly identified [correction of dentified] members of the TNF receptor superfamily (mTNFRH1 and mTNFRH2) inhibit T-cell proliferation.

By searching an EST database, we identified two TNF receptor superfamily members (named mTNFRH1 and mTNFRH2). Amino acid sequences are highly conserved between the two receptors (78% identity). The chromosomal loci of mTnfrh1 and mTnfrh2 genes are found in distal chromosome 7 in the mouse. mTNFRH1 and mTNFRH2 do not contain the cytoplasmic domain, indicating that they might function as decoy receptors. Furthermore, an alternatively spliced form of mTNFRH1 was found which contains neither the transmembrane domain nor the cytoplasmic domain, thus presumably existing as a soluble form. Northern blot analysis showed that mTnfrh1 mRNA was negligibly expressed in tissues, while mTnfrh2 mRNA was strongly expressed in spleen, lung, liver, kidney, and testis. When the extracellular domains of mTNFRH1 and mTNFRH2 were expressed in bacteria, their molecular weight of extracellular region was approximately 15 kDa. Both of the soluble forms were effective in inhibiting T-cell proliferation stimulated by anti-CD3 monoclonal antibody. Our data suggest that mTNFRH1 and mTNFRH2 may be implicated in exerting a modulatory role in the immune response.

Alternative Splicing↗

Divergence of Genbank and human tumor Bcl-2 sequences and implications for binding affinity to key apoptotic proteins.

Heterodimerization of antiapoptotic and pro-apoptotic Bcl-2 family of proteins provides an important mechanism for apoptosis regulation. Knowledge about key amino acids in the binding groove of native Bcl-2 contributing to this interaction will greatly facilitate the design of Bcl-2-specific inhibitors. There are two different Bcl-2 sequences, M13994 and M14745, in Genbank. Chimeric proteins Bcl-2(1) and Bcl-2(2) derived from the above sequences, although similar in structure, showed different binding affinities to Bak and Bad BH3 peptides (Petros et al., 2001). In this study, we show that the Bcl-2(1) sequence in normal and tumor human tissue samples differs from M13994 and M14745, and contains P59, T96, R110, S117 and G237. The actual sequence in the binding pocket matches the Bcl-2-Ig fusion sequence X06487, originally identified in a t(14:18) translocation of the Bcl-2 gene, associated with follicular lymphoma. The possible effects of the observed amino acid differences compared to M13994 and M14745 were investigated by combining structural data with fluorescence anisotropy. G110R substitution confers on Bcl-2(1) substantially increased binding affinity to Bak, Bad and Bax BH3 peptides, demonstrating that R110 is a key contributor to the BH3 binding affinity of Bcl-2. Although NMR structure did not predict R110 involvement in binding to these BH3 peptides, fluorescence anisotropy data clearly points to a critical role for this residue in binding to pro-apoptotic Bcl-2 family members.

Amino Acid Sequence↗

Natural selection drives recurrent formation of activating killer cell immunoglobulin-like receptor and Ly49 from inhibitory homologues.

Expression of killer cell Ig-like receptors (KIRs) diversifies human natural killer cell populations and T cell subpopulations. Whereas the major histocompatibility complex class I binding functions of inhibitory KIR are known, specificities for the activating receptors have resisted analysis. To understand better activating KIR and their relationship to inhibitory KIR, we took the approach of reconstructing their natural history and that of Ly49, the analogous system in rodents. A general principle is that inhibitory receptors are ancestral, the activating receptors having evolved from them by mutation. This evolutionary process of functional switch occurs independently in different species to yield activating KIR and Ly49 genes with similar signaling domains. Selecting such convergent evolution were the signaling adaptors, which are older and more conserved than any KIR or Ly49. After functional shift, further activating receptors form through recombination and gene duplication. Activating receptors are short lived and evolved recurrently, showing they are subject to conflicting selections, consistent with activating KIR's association with resistance to infection, reproductive success, and susceptibility to autoimmunity. Our analysis suggests a two-stage model in which activating KIR or Ly49 are initially subject to positive selection that rapidly increases their frequency, followed by negative selection that decreases their frequency and leads eventually to loss.

Amino Acid Sequence↗

RARTF: database and tools for complete sets of Arabidopsis transcription factors.

More than 5% of all genes in the Arabidopsis thaliana genome have been assumed to code for transcription factors. However, it has been difficult to accurately identify them. To construct proper sets of transcription factors, we used PSI-BLAST and InterProScan, and also checked several families manually. Especially to determine major Arabidopsis transcription factors (MYB, AP2/EREBP, bHLH, NAC, MADS, bZIP, WRKY), we compared the PSI-BLAST search results with those in recent reports. Finally, we identified 1968 proteins as transcription factors (7.4% of all Arabidopsis genes). We established a database named RARTF (RIKEN Arabidopsis Transcription Factor database, http://rarge.gsc.riken.jp/rartf/) based on the identified transcription factors. In RARTF, we provide information on the functional motif of transcription factors, full-length cDNAs, alternative pre-mRNA splicing events and Ac/Ds transposon-tagged mutants. We also provide expression profiles of 400 transcription factor genes in six experiments. We will report expression profiles of all transcription factor genes in various plant tissues under various stress and hormone conditions in the near future.

Amino Acid Sequence↗

The Ligand Gated Ion Channel Database.

The ligand gated ion channels (LGICs) are ionotropic receptors to neurotransmitters. Their physiological effect is carried out by the opening of an ionic channel upon binding of a particular neurotransmitter. These LGICs constitute superfamilies of receptors formed by homologous subunits. A database has been developed to handle the growing wealth of cloned subunits. This database contains nucleic acid sequences, protein sequences, as well as multiple sequence alignments and phylogenetic studies. This database is accessible via the worldwide web (http://www.pasteur.fr/units/neubiomol/LGIC.h tml), where it is continuously updated. A downloadable version is also available [currently v0.1 (98.06)].

Databases, Factual↗

Guar seed beta-mannan synthase is a member of the cellulose synthase super gene family.

Genes for the enzymes that make plant cell wall hemicellulosic polysaccharides remain to be identified. We report here the isolation of a complementary DNA (cDNA) clone encoding one such enzyme, mannan synthase (ManS), that makes the beta-1, 4-mannan backbone of galactomannan, a hemicellulosic storage polysaccharide in guar seed endosperm walls. The soybean somatic embryos expressing ManS cDNA contained high levels of ManS activities that localized to Golgi. Phylogenetically, ManS is closest to group A of the cellulose synthase-like (Csl) sequences from Arabidopsis and rice. Our results provide the biochemical proof for the involvement of the Csl genes in beta-glycan formation in plants.

Amino Acid Sequence↗

Different levels of variability in subtypes 1b and 4a of hepatitis C viruses.

We performed genetic and phenic analyses to evaluate nucleotide and amino-acid sequences of the amino-terminus of the E1 protein of HCV genotype 1b (extracted from databank) and 4a (characterised in this study). The non-synonymous (ka) mutation analysis demonstrated that the genome of genotype 1b was not saturated by variations, with a rate of transition/transversion (s/v) of 1.5, which is similar to the expected ratio (i.e., 2.0). The s/v ratio in genotype 4a isolates was lower (0.98), indicating saturation due long-term variability. Moreover, the genotype 1b sequences showed a higher number of ka mutations (s+v) (mean of 2.8 per sequence) than genotype 4a (mean of 1.5). The introduction of ka mutations resulted in a higher degree of amino acid variability in genotype 4a. In the genome of genotype 1b, each nucleotide mutation introduced new amino acids, with a Granthan distance of 3.35-42.5, whereas for genotype 4a the distances ranged from 48.8 to 102.1. The phenic analysis also indicated different and complex patterns of amino-acid substitution. Finally, diverse isoelectric points and hydrophobicity were predicted for the two genotypes, with a higher acidity for genotype 4a E1 proteins.

Amino Acid Substitution↗

Calculation of ligand-nucleic acid binding free energies with the generalized-born model in DOCK.

The calculation of ligand-nucleic acid binding free energies is investigated by including solvation effects computed with the generalized-Born model. Modifications of the solvation module in DOCK, including introduction of all-atom parameters and revision of coefficients in front of different terms, are shown to improve calculations involving nucleic acids. This computing scheme is capable of calculating binding energies, with reasonable accuracy, for a wide variety of DNA-ligand complexes, RNA-ligand complexes, and even for the formation of double-stranded DNA. This implementation of GB/SA is also shown to be capable of discriminating strong ligands from poor ligands for a series of RNA aptamers without sacrificing the high efficiency of the previous implementation. These results validate this approach to screening large databases against nucleic acid targets.

DNA↗