PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Function of 3' non-coding sequences and stop codon usage in expression of the chloroplast psaB gene in Chlamydomonas reinhardtii.

The rate of mRNA decay is an important step in the control of gene expression in prokaryotes, eukaryotes and cellular organelles. Factors that determine the rate of mRNA decay in chloroplasts are not well understood. Chloroplast mRNAs typically contain an inverted repeat sequence within the 3' untranslated region that can potentially fold into a stem-loop structure. These stem-loop structures have been suggested to stabilize the mRNA by preventing degradation by exonuclease activity, although such a function in vivo has not been clearly established. Secondary structures within the translation reading frame may also determine the inherent stability of an mRNA. To test the function of the inverted repeat structures in chloroplast mRNA stability mutants were constructed in the psaB gene that eliminated the 3' flanking sequences of psaB or extended the open reading frame into the 3' inverted repeat. The mutant psaB genes were introduced into the chloroplast genome of Chlamydomonas reinhardtii. Mutants lacking the 3' stem-loop exhibited a 75% reduction in the level of psaB mRNA. The accumulation of photosystem I complexes was also decreased by a corresponding amount indicating that the mRNA level is limiting to PsaB protein synthesis. Pulse-chase labeling of the mRNA showed that the decay rate of the psaB mRNA was significantly increased demonstrating that the stem-loop structure is required for psaB mRNA stability. When the translation reading frame was extended into the 3' inverted repeat the mRNA level was reduced to only 2% of wild-type indicating that ribosome interaction with stem-loop structures destabilizes chloroplast mRNAs. The non-photosynthetic phenotype of the mutant with an extended reading frame allowed us to test whether infrequently used stop codons (UAG and UGA) can terminate translation in vivo. Both UAG and UGA are able to effectively terminate PsaB synthesis although UGA is never used in any of the Chlamydomonas chloroplast genes that have been sequenced.

Amino Acid Sequence↗

Codon usage, genetic code and phylogeny of Dictyostelium discoideum mitochondrial DNA as deduced from a 7.3-kb region.

We have sequenced a region (7,376-bp) of the mitochondrial (mt) DNA (54 kb) of the cellular slime mold, Dictyostelium discoideum. From the DNA and amino-acid sequence comparisons with known sequences, genes for ATPase subunit 9 (ATP9), cytochrome b (CYTB), NADH dehydrogenase subunits 1, 3 and 6 (ND1, ND3 and ND6), small subunit rRNA (SSU rRNA) and seven tRNAs (Arg, Asn, Cys, Lys, f-Met, Met and Pro) have been identified. The sequenced region of the mtDNA has a high average A + T-content (70.8%). The A + T-content of protein-genes (73.6%) is considerably higher than that of RNA genes (61.3%). Even with the strong AT-bias, the genetic code employed is most probably the universal one. All seven tRNAs are able to form typical clover leaf structures. The molecular phylogenetic trees of CYTB and SSU rRNA suggest that D. discoideum is closer to green plants than to animals and fungi.

Amino Acid Sequence↗

Codon usage and bias in mitochondrial genomes of parasitic platyhelminthes.

Sequences of the complete protein-coding portions of the mitochondrial (mt) genome were analysed for 6 species of cestodes (including hydatid tapeworms and the pork tapeworm) and 5 species of trematodes (blood flukes and liver- and lung-flukes). A near-complete sequence was also available for an additional trematode (the blood fluke Schistosoma malayensis). All of these parasites belong to a large flatworm taxon named the Neodermata. Considerable variation was found in the base composition of the protein-coding genes among these neodermatans. This variation was reflected in statistically-significant differences in numbers of each inferred amino acid between many pairs of species. Both convergence and divergence in nucleotide, and hence amino acid, composition was noted among groups within the Neodermata. Considerable variation in skew (unequal representation of complementary bases on the same strand) was found among the species studied. A pattern is thus emerging of diversity in the mt genome in neodermatans that may cast light on evolution of mt genomes generally.

Amino Acid Sequence↗

Comparison of various algorithms for recognizing short coding sequences of human genes.

MOTIVATION: Since the early 1980s of the twentieth century, there has been great progress in the development of computational gene-finding algorithms. Some problems, however, have not yet been solved currently. Recognizing short genes in prokaryotes and short exons in eukaryotes is one of such problems. The paper is devoted to assessing various algorithms, including those currently available and the new ones proposed here, in order to find the best algorithm to solve the issue. RESULTS: The databases consisting of phase-specific coding and non-coding sequences of human genes with length of 192, 162, 129, 108, 87, 63 and 42 bp, respectively, have been established. Based on the databases and a standard benchmark, 19 algorithms were evaluated, which include the methods of Markov models with orders of 1 through 5, codon usage, hexamer usage, codon preference, amino acid usage, codon prototype, Fourier transform and 8 Z curve methods with various numbers of parameters. Consequently, the Z curve methods with 69 and 189 parameters are the best ones among them, based on the databases constructed here. In addition to the highest recognition accuracy confirmed by 10-fold cross-validation tests, the Z curve methods are much simpler computationally than the second best one, the fifth-order Markov chain model, in which 12 288 parameters are used. We hope that the Z curve methods presented in this paper would be beneficial to the further development of gene-finding algorithms. AVAILABILITY: The programs of various Z curve methods are available on request.

Algorithms↗

The genome of Campylobacter jejuni: codon and amino acid usage.

The genes from the genome of the AT-rich bacterium Campylobacter jejuni were analysed and characterised with respect to usage and amino acid usage. Codon usage is generally biased for all amino acids having synonymous codons, so that AT-rich synonyms are most frequently used. Markov chain analysis showed that codon bias and over- or underrepresentation of the corresponding tri-letter words are not related. Predicted secondary structure, lipophilicity, codon position within the gene, strand, and position on the (+)-strand were all shown to be determinants of codon usage, and these effects were in part directly explained by compositional phenomena. Codon context and the GC-content at the wobble position of the fourfold degenerate sites exert indirect effects on codon usage. The factors that affect codon usage seem to affect all amino acids, rather than selected amino acids. The usage of amino acids correlates well with the GC-content of genes, i.e. usage of amino acids encoded by GC-rich codons increases with GC-content and vice versa.

Amino Acids↗

Models of nearly neutral mutations with particular implications for nonrandom usage of synonymous codons.

The population dynamics of nearly neutral mutations are studied using a single-site and a multisite model. In the latter model, the nucleotides in a sequence are completely linked and the selection schemes employed are additive, multiplicative, and additive with a threshold. Although the third selection scheme is very different from the first two, the three schemes produce identical results for a wide range of parameter values. Thus the present study provides a general theory for the population dynamics of nearly neutral mutations because the results can also be used to draw inferences about other selection schemes such as stabilizing selection and synergistic selection. It is shown that the number of slightly deleterious mutations accumulated in a sequence can be considerably larger under the multisite model than under the single-site model, particularly if the sequence is long or if the mutation rate per site is high. The results show that even a very slight selective difference between synonymous codons can produce a strong bias in codon usage. Three alternative explanations for the strong bias in codon usage in bacterial and yeast genes are considered. The implications of the present results for molecular evolution are discussed.

Biological Evolution↗

In vivo evidence for non-universal usage of the codon CUG in Candida maltosa.

An alkane-assimilating yeast Candida maltosa had been studied in order to establish systems suitable for biotransformation of hydrophobic compounds. However, functional expression of heterologous genes tested for this purpose had not been successful in several cases. On the other hand, it had been reported that the codon CUG, a universal leucine codon, is read as serine in C. cylindracea. The same altered codon usage had also been suggested by in vitro experiments in some Candida yeasts which are phylogenetically closely related to C. maltosa. In this study we have shown that the failure in functional expression of a heterologous gene is due to the fact that the codon CUG is read as serine in C. maltosa. This conclusion was drawn from the following experimental results: (1) when a cytochrome P450 gene of C. maltosa containing a CTG codon was expressed in C. maltosa, the corresponding amino acid was found to be serine, and not leucine; (2) a tRNA gene with an almost identical structure to that of the tRNASerCAG gene of C. albicans could be isolated from the genome of C. maltosa; (3) the Saccharomyces cerevisiae URA3 gene, which has one CTG codon, could not complement the ura3 mutation of C. maltosa as itself, but when the CTG codon was changed to another leucine codon, CTC, the mutated gene could complement the ura3 mutation. The last result is the first example of succeeding in functional expression of a heterologous gene in Candida species having an altered codon usage by changing the CTG codon in the gene to another codon.

Amino Acid Sequence↗

Context-dependent codon bias and messenger RNA longevity in the yeast transcriptome.

Context-dependent codon bias and its relationship with messenger RNA (mRNA) longevity was examined in 4,648 mRNA transcripts of the Saccharomyces cerevisiae transcriptome for which mRNA half-lives have been empirically determined. Surprisingly, rare codon usage (codons used <13 times per 1,000 codons in the genome) increased with mRNA half-life. However, it is shown that this pattern was not due to preference for rare codon use within codon families containing both rare and nonrare codons. Rather, the pattern was due to an increase in the frequency of amino acids encoded solely by rare codons, and a decrease in the frequency of amino acids never encoded by rare codons, with mRNA half-life. When standardized by open reading frame length, the use of consecutive rare codons was also positively correlated with mRNA half-life. There was negative correlation between the usage of synonymous A|T dinucleotides spanning codon boundaries and mRNA half-life, despite the fact that the frequency of AT dinucleotide usage overall, and AT dinucleotide usage at other codon position contexts (e.g., 1-2, 2-3, or 3|1 total), was not correlated with mRNA half-life. The use of A|T dinucleotides at synonymous dicodon boundaries could potentially allow for more efficient 3'-5' degradation by endonucleolytic cleavage.

Codon↗

Unusual usage of AGG and TTG codons in humans and their viruses.

Prior analysis on human protein-coding DNA sequences has identified local base composition as the primary predictor of synonymous codon usage. However, in many organisms, codon usage is influenced by natural selection, particularly for efficient expression of functional gene products. Because viruses are expected to evolve codon usage in the context of their host's molecular machinery, their genomes provide another window into the forces that guide their host's molecular evolution. Factor analysis was performed on codon usage of 16,654 genes annotated in Build 34 of the human genome, and the primary factor was correlated strongly with local base composition. However, two codons, AGG and TTG, rose in frequency as all other C- and G-ending codons decreased in frequency. These two codons were the only C- or G-ending codons with usages that negatively correlated with gene expression. Variation among viruses in codon usage also strongly reflects variation in base composition and, again, AGG and TTG decrease in frequency as all other C- and G-ending codons increase in frequency. It appears that usages of these two codons can not be explained by local compositional biases, implying a more direct role of natural selection on codon usage in humans.

Amino Acids↗

Identification of genomic islands in six plant pathogens.

Genomic islands (GIs) play important roles in microbial evolution, which are acquired by horizontal gene transfer. In this paper, the GIs of six completely sequenced plant pathogens are identified using a windowless method based on Z curve representation of DNA sequences. Consequently, four, eight, four, one, two and four GIs are recognized with the length greater than 20-Kb in plant pathogens Agrobacterium tumefaciens str. C58, Rolstonia solanacearum GMI1000, Xanthomonas axonopodis pv. citri str. 306 (Xac), Xanthomonas campestris pv. campestris str. ATCC33913 (Xcc), Xylella fastidiosa 9a5c and Pseudomonas syringae pv. tomato str. DC3000, respectively. Most of these regions share a set of conserved features of GIs, including an abrupt change in GC content compared with that of the rest of the genome, the existence of integrase genes at the junction, the use of tRNA as the integration sites, the presence of genetic mobility genes, the difference of codon usage, codon preference and amino acid usage, etc. The identification of these GIs will benefit the research for the six important phytopathogens.

Agrobacterium tumefaciens↗

Molecular evolution of the major outer-membrane protein gene (oprF) of Pseudomonas.

The major outer-membrane protein of Pseudomonas, OprF, is multifunctional. It is a non-specific porin that plays a role in maintenance of cell shape, in growth in a low-osmolarity environment, and in adhesion to various supports or molecules. OprF has been studied extensively for its utility as a vaccine component, its role in antimicrobial drug resistance, and its porin function. The authors have previously shown important differences between the OprF and 16S rDNA phylogenies: Pseudomonas fluorescens isolates split into two quite separate clusters, probably according to their ecological niche. In this study, the evolutionary history of the oprF gene was investigated further. The study of G+C content at the third codon position, synonymous codon usage (codon adaptation index, CAI) and genomic context showed no evidence of horizontal transfer or gene duplication. Similarly, a robust likelihood test of incongruence showed no significant incongruence between the oprF phylogeny and the species phylogeny. In addition, the ratio of nonsynonymous mutations to synonymous mutations (K(a)/K(s)) is high between the different clusters, especially between the two clusters containing P. fluorescens isolates, highlighting important modifications in evolutionary constraints during the history of the oprF gene. Since OprF is known as a pleiotropic protein, modifications in evolutionary constraints could have resulted from variations in cryptic functions, correlated with the ecological fingerprint. Finally, relaxed constraints and/or episodic positive evolution, especially for some P. fluorescens strains, could have led to a phylogeny reconstruction artifact.

Bacterial Outer Membrane Proteins↗

The 'effective number of codons' used in a gene.

A simple measure is presented that quantifies how far the codon usage of a gene departs from equal usage of synonymous codons. This measure of synonymous codon usage bias, the 'effective number of codons used in a gene', Nc, can be easily calculated from codon usage data alone, and is independent of gene length and amino acid (aa) composition. Nc can take values from 20, in the case of extreme bias where one codon is exclusively used for each aa, to 61 when the use of alternative synonymous codons is equally likely. Nc thus provides an intuitively meaningful measure of the extent of codon preference in a gene. Codon usage patterns across genes can be investigated by the Nc-plot: a plot of Nc vs. G + C content at synonymous sites. Nc-plots are produced for Homo sapiens, Saccharomyces cerevisiae, Escherichia coli, Bacillus subtilis, Dictyostelium discoideum, and Drosophila melanogaster. A FORTRAN77 program written to calculate Nc is available on request.

Animals↗