PubMed HealthSearch

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Comprehensive analysis of synonymous codon usage bias and evolutionary dynamics in the chloroplast genomes of eight Coptis species.

Coptis is a medically important genus renowned for producing valuable isoquinoline alkaloids. Although its chloroplast genomes encode key components for photosynthesis and plastid gene expression, the evolutionary constraints acting on their coding sequences and synonymous codon usage remain poorly resolved. Here, we combined a transparent taxon-level sampling strategy with comparative analyses of chloroplast CDSs from eight Coptis taxa. We quantified nucleotide composition, relative synonymous codon usage, effective number of codons, neutrality and PR2 patterns, and correspondence analysis, and then integrated these results with a core-CDS distance analysis and gene-wise pairwise dN/dS estimates. The chloroplast genomes showed a conserved AT-rich composition, especially at the third codon position (GC3 approximately 30.3-30.8%), with a consistent GC1 > GC2 > GC3 trend. Thirty preferred codons were detected, 28 ending in A/T, and eleven optimal codons were shared across the genus. The core-CDS distance analysis recovered a close relationship between C. chinensis and C. chinensis var. brevisepala, whereas most coding genes showed dN/dS values below one, consistent with pervasive purifying constraint. Across 48 consistently filtered CDSs, GC3s was negatively associated with mean dN (Spearman rho = -0.404, P = 0.00439) and CAI was positively associated with mean dN (rho = 0.303, P = 0.0361), whereas the remaining associations were not significant (all P > = 0.0972). These results extend codon-usage analysis by linking synonymous-site composition to coding-sequence evolution within Coptis, while providing a hypothesis-generating resource for future plastid engineering studies.

Genome, Chloroplast

Codon usage is imposed by the gene location in the transcription unit.

A characteristic profile of the fluctuations of codon usage is observed in bacteriophages and mitochondria. By following the DNA in the direction of transcription, one moves slowly from a region where selective pressure favours codons ending with C to a region where the bias is in favour of codons ending with T; then, abruptly, one again enters a region of codons ending in C. The transcription end point takes place in the area of abrupt change in codon usage. By comparing Drosophila yakuba and mouse mitochondrial genomes, it is possible to show that the strategy of codon usage for a given gene depends on its location along the transcription unit and not on the encoded protein. The choice of codons ending in T or C allows large scale variations of DNA stability which could regulate the speed of propagation of the RNA polymerase.

Animals

Codon usage in the G+C-rich Streptomyces genome.

The codon usage (CU) patterns of 64 genes from the Gram+ prokaryotic genus Streptomyces were analysed. Despite the extremely high overall G+C content of the Streptomyces genome (estimated at 0.74), individual genes varied in G+C content from 0.610 to 0.797, and had third codon position G+C contents (GC3s) that varied from 0.764 to 0.983. The variation in GC3s explains a significant proportion of the variation in CU patterns. This is consistent with an evolutionary model of the Streptomyces genome where biased mutation pressure has led to a high average G+C content with random variation about the mean, although the variation observed is greater than that expected from a simple binomial model. The only gene in the sample that can be confidently predicted to be highly expressed, EF-Tu of Streptomyces coelicolor A3(2) (GC3s = 0.927), shows a preference for a third position C in several of the four codon families, and for CGY and GGY for Arg and Gly codons, respectively (Y = pyrimidine); similar CU patterns are found in highly expressed genes of the G+C-rich Micrococcus luteus genome. It thus appears that codon usage in Streptomyces is determined predominantly by mutation bias, with weak translational selection operating only in highly expressed genes. We discuss the possible consequences of the extreme codon bias of Streptomyces and consider how it may have evolved. A set of CU tables is provided for use with computer programs that locate protein-coding regions.

Base Composition

The rate of synonymous substitution in enterobacterial genes is inversely related to codon usage bias.

Genes sequences from Escherichia coli, Salmonella typhimurium, and other members of the Enterobacteriaceae show a negative correlation between the degree of synonymous-codon usage bias and the rate of nucleotide substitution at synonymous sites. In particular, very highly expressed genes have very biased codon usage and accumulate synonymous substitutions very slowly. In contrast, there is little correlation between the degree of codon bias and the rate of protein evolution. It is concluded that both the rate of synonymous substitution and the degree of codon usage bias largely reflect the intensity of selection at the translational level. Because of the high variability among genes in rates of synonymous substitution, separate molecular clocks of synonymous substitution might be required for different genes.

Biological Evolution

An analysis of codon usage in mammals: selection or mutation bias?

A new statistical test has been developed to detect selection on silent sites. This test compares the codon usage within a gene and thus does not require knowledge of which genes are under the greatest selection, that there exist common trends in codon usage across genes, or that genes have the same mutation pattern. It also controls for mutational biases that might be introduced by the adjacent bases. The test was applied to 62 mammalian sequences, and significant codon usage biases were detected in all three species examined (humans, rats, and mice). However, these biases appear not to be the consequence of selection, but of the first base pair in the codon influencing the mutation pattern at the third position.

Animals

Significance of nucleotide sequence alignments: a method for random sequence permutation that preserves dinucleotide and codon usage.

The similarity of two nucleotide sequences is often expressed in terms of evolutionary distance, a measure of the amount of change needed to transform one sequence into the other. Given two sequences with a small distance between them, can their similarity be explained by their base composition alone? The nucleotide order of these sequences contributes to their similarity if the distance is much smaller than their average permutation distance, which is obtained by calculating the distances for many random permutations of these sequences. To determine whether their similarity can be explained by their dinucleotide and codon usage, random sequences must be chosen from the set of permuted sequences that preserve dinucleotide and codon usage. The problem of choosing random dinucleotide and codon-preserving permutations can be expressed in the language of graph theory as the problem of generating random Eulerian walks on a directed multigraph. An efficient algorithm for generating such walks is described. This algorithm can be used to choose random sequence permutations that preserve (1) dinucleotide usage, (2) dinucleotide and trinucleotide usage, or (3) dinucleotide and codon usage. For example, the similarity of two 60-nucleotide DNA segments from the human beta-1 interferon gene (nucleotides 196-255 and 499-558) is not just the result of their nonrandom dinucleotide and codon usage.

Base Sequence

Nucleic acid composition, codon usage, and the rate of synonymous substitution in protein-coding genes.

Based on the rates of synonymous substitution in 42 protein-coding gene pairs from rat and human, a correlation is shown to exist between the frequency of the nucleotides in all positions of the codon and the synonymous substitution rate. The correlation coefficients were positive for A and T and negative for C and G. This means that AT-rich genes accumulate more synonymous substitutions than GC-rich genes. Biased patterns of mutation could not account for this phenomenon. Thus, the variation in synonymous substitution rates and the resulting unequal codon usage must be the consequence of selection against A and T in synonymous positions. Most of the variation in rates of synonymous substitution can be explained by the nucleotide composition in synonymous positions. Codon-anticodon interactions, dinucleotide frequencies, and contextual factors influence neither the rates of synonymous substitution nor codon usage. Interestingly, the nucleotide in the second position of codons (always a nonsynonymous position) was found to affect the rate of synonymous substitution. This finding links the rate of nonsynonymous substitution with the synonymous rate. Consequently, highly conservative proteins are expected to be encoded by genes that evolve slowly in terms of synonymous substitutions, and are consequently highly biased in their codon usage.

Animals

Codon usage in Homo sapiens: evidence for a coding pattern on the non-coding strand and evolutionary implications of dinucleotide discrimination.

This study reports the analysis of codon usage in 35 complete Homo sapiens genes. Both codon frequency and inter-codon interference exhibit patterns of evolutionary interest. There is a significant positive correlation between the frequency with which a given codon is used and the frequency with which its complement is used. Since the frequency of appearance of the complementary codon on the coding strand is equal to the frequency of appearance of the original codon on the non-coding strand, in the same phase, the non-coding strand is found to resemble the coding strand in triplet composition. The same effect has been observed in Escherichia coli. This preference for the use of certain complementary triplets as codons suggests that the evolution of the use of the genetic code depended to some extent upon the double-stranded nature of the coding material. In addition, the effect of discrimination against the use of two dinucleotides, CpG and UpA, is observed in codon usage and also in adjacent codon interference. Codons beginning with G, or A, are unlikely to be preceded by codons ending in C, or U, respectively. Consideration of codon assignment in the genetic code together with the observed CpG infrequency suggests that the evolution of the code may have been influenced by conditions in which the use of CpG dinucleotides was unfavorable. The infrequent use of UpA dinucleotides can be explained as the result of frameshift mutation during gene evolution.

Base Sequence

Structural organization and unusual codon usage in the DNA polymerase gene from herpes simplex virus type 1.

We have analyzed the protein and nucleic acid sequences of the DNA polymerase from herpes simplex virus type 1 (HSV-1) to provide insight into the expression and possible structure of this enzyme. Extensive similarity between the amino acid sequence and that of the Epstein Barr virus DNA polymerase is reported. We describe probable structural similarities between these proteins and the use of these similarities to define structural and functional domains within the polymerase. Analysis of base composition and codon usage reveals that several genes from HSV-1, including DNA polymerase, exhibit a strong preference for guanine or cytosine at the third codon position. This preference may result from the high guanine + cytosine content of the virus and produces a highly restricted codon usage, different from that of the host cell. Consequences of the unusual codon usage for viral expression include the potential for extensive mRNA secondary structure.

Amino Acid Sequence

Nucleotide sequence of the gene for the immunity protein to colicin A. Analysis of codon usage of immunity proteins as compared to colicins.

The nucleotide sequence of the structural gene for the immunity protein to colicin A (cai) has been established. This sequence consists of 534 base pairs. According to the predicted amino acid sequence, the polypeptide chain of this immunity protein comprises 178 amino acids and has a relative molecular mass of 20462. As expected from its localization in the inner membrane, large hydrophobic fragments are found along the polypeptide chain that also contains clusters of mostly positively charged residues. The cai like the ceiA genes encode proteins that are weakly expressed as compared to the corresponding colicins (A and E1). Codon usage reflects this difference. In contrast, the four genes for immunity to cloacin DF13 and to colicin E3 and for these bacteriocins, all of which are highly expressed and are organized in operon, display similar codon usage. These results are discussed with regards to the possible relationship between expressivity and codon usage.

Bacterial Proteins

Correlations between the compositional properties of human genes, codon usage, and amino acid composition of proteins.

We have analyzed the correlation that exists between the GC levels of third and first or second codon position for about 1400 human coding sequences. The linear relationship that was found indicates that the large differences in GC level of third codon positions of human genes are paralleled by smaller differences in GC levels of first and second codon positions. Whereas third codon position differences correspond to very large differences in codon usage within the human genome, the first and second codon position differences correspond to smaller, yet very remarkable, differences in the amino acid composition of encoded proteins. Because GC levels of codon positions are linearly correlated with the GC levels of the isochores harboring the corresponding genes, both codon usage and amino acid composition are different for proteins encoded by genes located in isochores of different GC levels. Furthermore, we have also shown that a linear relationship with a unit slope and a correlation coefficient of 0.77 exists between GC levels of introns and exons from the 238 human genes currently available for this analysis. Introns are, however, about 5% lower in GC, on average, than exons from the same genes.

Amino Acid Sequence

The effect of codon usage on the oligonucleotide composition of the E. coli genome and identification of over- and underrepresented sequences by Markov chain analysis.

As shown in the accompanying paper (5), the oligonucleotide composition of the E. coli genome is highly asymmetric for sequences up to 6 bp in length when ranked from highest to lowest abundance. We show here that this largely reflects codon usage because heavily used codons were found in the highly abundant oligomers whereas rarely used codons, with some exceptions, occurred in sequences in low abundance. Furthermore, linear regression analysis revealed a strong correlation between the frequencies of each trinucleotide and its usage as a codon. Dinucleotides are also not randomly distributed across each codon position and the dinucleotide composition of genes that are transcribed but not translated (rRNA and tRNA genes) was highly related to that seen in genes encoding polypeptides. However, 45 tetra-, 8 penta-, and 6 hexanucleotides were significantly over- or underabundant by Markov chain analysis and could not be accounted for by codon usage. Of these underrepresented sequences, many were palindromes, including the Dam methylation site.

Base Sequence

Codon replacement in the PGK1 gene of Saccharomyces cerevisiae: experimental approach to study the role of biased codon usage in gene expression.

The coding sequences of genes in the yeast Saccharomyces cerevisiae show a preference for 25 of the 61 possible coding triplets. The degree of this biased codon usage in each gene is positively correlated to its expression level. Highly expressed genes use these 25 major codons almost exclusively. As an experimental approach to studying biased codon usage and its possible role in modulating gene expression, systematic codon replacements were carried out in the highly expressed PGK1 gene. The expression of phosphoglycerate kinase (PGK) was studied both on a high-copy-number plasmid and as a single copy gene integrated into the chromosome. Replacing an increasing number (up to 39% of all codons) of major codons with synonymous minor ones at the 5' end of the coding sequence caused a dramatic decline of the expression level. The PGK protein levels dropped 10-fold. The steady-state mRNA levels also declined, but to a lesser extent (threefold). Our data indicate that this reduction in mRNA levels was due to destabilization caused by impaired translation elongation at the minor codons. By preventing translation of the PGK mRNAs by the introduction of a stop codon 3' and adjacent to the start codon, the steady-state mRNA levels decreased dramatically. We conclude that efficient mRNA translation is required for maintaining mRNA stability in S. cerevisiae. These findings have important implications for the study of the expression of heterologous genes in yeast cells.

Amino Acid Sequence

Codon usage and G + C content in Bradyrhizobium japonicum genes are not uniform.

To date, the sequences of 45 Bradyrhizobium japonicum genes are known. This provides sufficient information to determine their codon usage and G + C content. Surprisingly, B. japonicum nodulation and NifA-regulated genes were found to have a less biased codon usage and a lower G + C content than genes not belonging to these two groups. Thus, the coding regions of nodulation genes and NifA-regulated genes could hardly be identified in codon preference plots whereas this was not difficult with other genes. The codon frequency table of the highly biased genes was used in a codon preference plot to analyze the RSRj alpha 9 sequence which is an insertion sequence (IS)-like element. The plot helped identify a new open reading frame (ORF355) that escaped previous detection because of two sequencing errors. These were now corrected. The deduced gene product of ORF355 in RSRj alpha 9 showed extensive similarity to a putative protein encoded by an ORF in the T-DNA of Agrobacterium rhizogenes. The DNA sequences bordering both ORFs showed inverted repeats and potential target site duplications which supported the assumption that they were IS-like elements.

Amino Acid Sequence

Diagrammatization of codon usage in 339 human immunodeficiency virus proteins and its biological implication.

The occurrence frequencies of bases A (adenine), C (cytosine, G (guanine), and T (thymine) occurring in the 1st, 2nd, and 3rd codon positions in the codon usage table of viral genes for the 339 human immunodeficiency virus (HIV) proteins compiled recently have been calculated and diagrammatized. For comparison, the corresponding diagrammatic representations for the 2681 human proteins from the codon usage table for primate genes are also presented. The analyzed results based on these characteristic diagrams indicate that considerably similar features have been found between HIV and human proteins for the 1st and 2nd codon positions; i.e., they are all occupied predominantly by purine, especially base A. However, a significant difference in the 3rd codon position between HIV and human proteins has been observed; i.e., human proteins are of high C + G content and low A + G content in the 3rd codon position, whereas the case is just the opposite for HIV proteins. The biological implication of such a duality on the codon bias of HIV against human proteins is discussed. It is suggested that the 1st and 2nd codon positions can be termed as the structure-determining position, and the 3rd codon position termed as the species-determining position. The diagrammatic representation and analysis method described here possess a great potential for the study of molecular evolution from the viewpoint of the genetic code for which data have been accumulated rapidly and will continue to grow at a much faster pace.

Base Composition

The selection-mutation-drift theory of synonymous codon usage.

It is argued that the bias in synonymous codon usage observed in unicellular organisms is due to a balance between the forces of selection and mutation in a finite population, with greater bias in highly expressed genes reflecting stronger selection for efficiency of translation. A population genetic model is developed taking into account population size and selective differences between synonymous codons. A biochemical model is then developed to predict the magnitude of selective differences between synonymous codons in unicellular organisms in which growth rate (or possibly growth yield) can be equated with fitness. Selection can arise from differences in either the speed or the accuracy of translation. A model for the effect of speed of translation on fitness is considered in detail, a similar model for accuracy more briefly. The model is successful in predicting a difference in the degree of bias at the beginning than in the rest of the gene under some circumstances, as observed in Escherichia coli, but grossly overestimates the amount of bias expected. Possible reasons for this discrepancy are discussed.

Amino Acyl-tRNA Synthetases

Specific codon usage pattern and its implications on the secondary structure of silk fibroin mRNA.

We have identified two distinctive regions of the repetitive unit nucleotide sequence of fibroin mRNA of Bombyx mori. The codon usage for the major amino acids, glycine, alanine and serine is distinctly different in these two regions, indicating that it is determined by the fibroin mRNA or gene structure but not by the tRNA population. Comparative computer analyses of nucleotide substitutions in the unit sequence suggest that selection has operated on the codon usage to optimize the secondary structure characteristic of the fibroin mRNA.

Amino Acid Sequence