PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Codon Usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Large-scale, multi-genome analysis of alternate open reading frames in bacteria and archaea.

Analysis of over 300,000 annotated genes in 105 bacterial and archaeal genomes reveals an unexpectedly high frequency of large (>300 nucleotides) alternate open reading frames (ORFs). Especially notable is the very high frequency of alternate ORFs in frames +3 and -1 (where the annotated gene is defined as frame +1). The occurrence of alternate ORFs is correlated with genomic G+C content and is strongly influenced by synonymous codon usage bias. The frequency of alternate ORFs in frame -1 is also influenced by the occurrence of codons encoding leucine and serine in frame +1. Although some alternate ORFs have been shown to encode proteins, many others are probably not expressed because they lack appropriate signals for transcription and translation. These latter can be mis-annotated by automatic gene finding programs leading to errors in public databases. Especially prone to mis-annotation is frame -1, because it exhibits a potential codon usage and theoretical capacity to encode proteins with an amino acid composition most similar to real genes. Some alternate ORFs are conserved across bacterial or archaeal species, and can give rise to misannotated "conserved hypothetical" genes, while others are unique to a genome and are misidentified as "hypothetical orphan" genes, contributing significantly to the orphan gene paradox.

Algorithms↗

The ribosomal protein S8 from Thermus thermophilus VK1. Sequencing of the gene, overexpression of the protein in Escherichia coli and interaction with rRNA.

The gene of the ribosomal protein S8 from Thermus thermophilus VK1 has been isolated from a genomic library by hybridization of an oligonucleotide coding for the N-terminal amino acid sequence of the protein, amplified by PCR and sequenced. Nucleotide sequence reveals an open reading frame coding for a protein of 138 amino acid residues (M(r) 15,839). The codon usage shows that 94% of the codons possess G or C in the third position, and agrees with the preferential usage of codons of high G+C content in the bacteria of the genus Thermus. The amino acid sequence of the protein shows 48% identity with the protein from Escherichia coli. Ribosomal protein S8 from T. thermophilus has been expressed in E. coli under the control of the T7 promoter and purified to homogeneity by heat treatment of the extract followed by cation-exchange chromatography. Conditions were defined in which T. thermophilus protein S8 binds specifically an homologous 16S rRNA fragment containing the putative S8 binding site with an apparent association constant of 5 x 10(7) M-1. The overexpressed protein binds the rRNA with the same affinity as that extracted from T. thermophilus, indicating that the thermophilic protein is correctly folded in E. coli. The specificity of this binding is dependent on the ionic strength. The protein S8 from T. thermophilus recognizes the E. coli rRNA binding sites as efficiently as the S8 protein from E. coli. This result agrees with sequence comparisons of the S8 binding site on the small subunit rRNA from E. coli and from T. thermophilus, showing strong similarities in the regions involved in the interaction. It suggests that the structural features responsible for the recognition are conserved in the mesophilic and thermophilic eubacteria, despite structural peculiarities in the thermophilic partners conferring thermostability.

Amino Acid Sequence↗

[Use of codons in plant lectins].

Codon usage in the coding region of mature lectins has been examined for 11 plant species (8 leguminoseae, 1 euphorbiaceae, 2 gramineae). The different legume lectins exhibit nearly the same codon usage pattern whereas the choice for the silent position of codons is non-random.

Codon↗

Molecular characterization of the glyceraldehyde-3-phosphate dehydrogenase gene of Phaffia rhodozyma.

The glyceraldehyde-3-phosphate dehydrogenase (GPD; EC1.2.1.12)-encoding gene (gpd) was isolated from a genomic library of Phaffia rhodozyma CBS 6938. Unlike some other eukaryotic organisms the gpd gene is represented by a single copy in P. rhodozyma. The complete nucleotide sequence of the coding, as well as the flanking non-coding regions was determined. The nucleotide sequence of gpd predicted six introns and a polypeptide chain of 339 amino acids. The codon usage in the gpd gene of P. rhodozyma was highly biased and was significantly different from the codon usage in other yeasts. Phylogenetic analysis of different yeasts and filamentous asco- and basidiomycetes gpd sequences indicated that the gpd gene of P. rhodozyma forms a cluster with the corresponding genes of filamentous basidiomycetes.

Amino Acid Sequence↗

The Pseudomonas cepacia 249 chromosomal penicillinase is a member of the AmpC family of chromosomal beta-lactamases.

Pseudomonas cepacia 249 produces an inducible beta-lactamase with penicillinase activity. The nucleotide sequence of the penA gene, which encodes this beta-lactamase, was determined and found to include regions with a significant homology to the ampC-encoded beta-lactamases of members of the family Enterobacteriaceae and Pseudomonas aeruginosa. The predicted amino acid sequence of the PenA beta-lactamase contained 17 amino acids immediately preceding the putative active-site serine which were highly conserved among the enzymes of the AmpC family. Although the penA-coding sequence had a total GC content of 60%, the predicted codon usage was more characteristic of Escherichia coli ampC-encoded beta-lactamase, with 53% of the codons having G or C in the third position, in contrast to the values for the P. aeruginosa ampC (88.5%) or Pseudomonas cepacia (88 to 92%) metabolic genes. The inducible expression of penA can be regulated by the E. coli gene product AmpD. A putative P. cepacia AmpR homolog was associated with the positive regulation of both Enterobacter cloacae ampC and P. cepacia penA expression, as confirmed by gel retardation studies. The E. cloacae AmpR did not regulate penA expression. Thus, by homology studies, codon usage, and genetic analysis, the P. cepacia penA beta-lactamase appears to have been acquired from members of the family Enterobacteriaceae and belongs to the class C group of beta-lactamases.

Base Sequence↗

Correlated evolution of synonymous and nonsynonymous sites in Drosophila.

Recent work has shown that Drosophila melanogaster genes with fast-evolving nonsynonymous sites have lower codon usage bias. This pattern has been attributed to interference between positive selection at nonsynonymous sites and weak selection on codon usage. Here we have looked for this correlation in a much larger and less biased dataset, comprising 630 gene pairs from D. melanogaster and D. yakuba. We confirmed that there is a negative correlation between the rate of nonsynonymous substitutions (d(N)) and codon bias in D. melanogaster. We then tested the interference hypothesis and other alternative explanations, including one involving gene expression. We found that d(N) indeed correlates with the level of gene expression. Given that gene expression is a strong determinant of codon bias, the relationship between d(N) and codon bias might be a by-product of gene expression. However, our tests show that none of the hypotheses we consider seem to explain the data fully.

Amino Acid Substitution↗

Background selection in single genes may explain patterns of codon bias.

Background selection involves the reduction in effective population size caused by the removal of recurrent deleterious mutations from a population. Previous work has examined this process for large genomic regions. Here we focus on the level of a single gene or small group of genes and investigate how the effects of background selection caused by nonsynonymous mutations are influenced by the lengths of coding sequences, the number and length of introns, intergenic distances, neighboring genes, mutation rate, and recombination rate. We generate our predictions from estimates of the distribution of the fitness effects of nonsynonymous mutations, obtained from DNA sequence diversity data in Drosophila. Results for genes in regions with typical frequencies of crossing over in Drosophila melanogaster suggest that background selection may influence the effective population sizes of different regions of the same gene, consistent with observed differences in codon usage bias along genes. It may also help to cause the observed effects of gene length and introns on codon usage. Gene conversion plays a crucial role in determining the sizes of these effects. The model overpredicts the effects of background selection with large groups of nonrecombining genes, because it ignores Hill-Robertson interference among the mutations involved.

Animals↗

Ornithine decarboxylase and trypanothione reductase genes in Leishmania braziliensis guyanensis.

Ornithine decarboxylase and trypanothione reductase are the key enzymes in polyamine and trypanothione metabolism in kinetoplastids. Using a heterologous Trypanosoma brucei brucei probe for ornithine decarboxylase and a mixed synthetic probe of 29 oligonucleotides for trypanothione reductase, we have detected the putative genes for these enzymes by Southern blot hybridization using genomic DNA of Leishmania braziliensis guyanensis MHOM/SR/80/CUMC 1. The trypanothione reductase probe was constructed both from the conserved codon usage of the redox active site for other flavin oxidoreductases over a wide evolutionary scale, and the preferred codon usage for other genes in species of Leishmania.

Amino Acid Sequence↗

[Use of the hygromycin phosphotransferase gene as the dominant selective marker for Chlamydomonas reinhardtii transformation].

The hygromycin phosphotransferase gene (hpt) from E. coli under the control of the SV40 early promoter was used as a dominant selectable marker for transformation of Chlamydomonas reinhardtii. Cells were transformed by electroporation (pulse length, 2 ms, field strength, 1 kV/cm). The culture growth phase was a crucial parameter for transformation (optimal density approximately 10(6) cells/ml). It was possible to obtain approximately 10(3) Hyg-resistant colonies under these conditions. Foreign DNA integrated into the Chlamydomonas genome was maintained for at least 8 months but the Hyg-resistant phenotype of the transformed clones was unstable. The frequency of codon usage in the hpt gene was compared with the one in Chlamydomonas nuclear genes. It is supposed that highly biased codon usage in Chlamydomonas does not preclude expression. Advantages of this selection system for studying Chlamydomonas transformation by heterologous genes are discussed.

Animals↗

Mitochondrial genes collectively suggest the paraphyly of Crustacea with respect to Insecta.

Complete sequences of seven protein coding genes from Penaeus notialis mitochondrial DNA were compared in base composition and codon usage with homologous genes from Artemia franciscana and four insects. The crustacean genes are significantly less A + T-rich than their counterpart in insects and the pattern of codon usage (ratio of G + C-rich versus A + T-rich codon) is less biased. A phylogenetic analysis using amino acid sequences of the seven corresponding polypeptides supports a sister-taxon status for mollusks-annelid and arthropods. Furthermore, a distance matrix-based tree and two most-parsimonious trees both suggest that crustaceans are paraphyletic with respect to insects. This is also supported by the inclusion of Panulirus argus COII (complete) and COI and COIII (partial) sequence data. From analysis of single and combined genes to infer phylogenies, it is observed that obtained from single genes are not well supported in most topologies cases and notably differ from that of the tree based on all seven genes.

Animals↗

Contextual constraints on codon pair usage: structural and biological implications.

Complementary DNA sequence data of 278 protein coding genes from prokaryotic systems have been analysed at the level of near neighbour codon pairs. Our analysis points out that constraints exist even at the level of near neighbour codon pairs. These constraints are in addition to those which arise due to relative levels of tRNA. Codon pairs, which in the data base have different occurrence values from their expected values, neither have common secondary structure nor do have better stabilization due to high base stacking. Our study points out that there are strong interaction between constituent codons in these codon pairs. These strongly interacting codon pairs, we suggest, are involved in the formation of three dimensional structural elements of cDNA/mRNA and interact with ribosome and thus modulate translation.

Base Sequence↗

High expression of a synthetic gene encoding potato alpha-glucan phosphorylase in Aspergillus niger.

We describe the successful heterologous expression of the Solanum tuberosum alpha-glucan phosphorylase (GP) gene in Aspergillus niger. Special attention was paid to the influence of different codon usage and A+T content in the coding region on GP protein expression. Use of A. niger-preferred codon usage and lower A+T content in a synthetic gene (GP-syn) resulted in a significant improvement in the level of the GP mRNA and a dramatic increase in the quantity of GP protein produced such that it accounted for approximately 10% of the total soluble protein. We suggest that redesigning the primary DNA sequence encoding a desired protein product can be an extremely effective method for improving heterologous protein production in filamentous fungi.

Aspergillus niger↗

Evolution of tropomyosin functional domains: differential splicing and genomic constraints.

We have cloned and determined the nucleotide sequence of a complementary DNA (cDNA) encoded by a newly isolated human tropomyosin gene and expressed in liver. Using the least-square method of Fitch and Margoliash, we investigated the nucleotide divergences of this sequence and those published in the literature, which allowed us to clarify the classification and evolution of the tropomyosin genes expressed in vertebrates. Tropomyosin undergoes alternative splicing on three of its nine exons. Analysis of the exons not involved in differential splicing showed that the four human tropomyosin genes resulted from a duplication that probably occurred early, at the time of the amphibian radiation. The study of the sequences obtained from rat and chicken allowed a classification of these genes as one of the types identified for humans. The divergence of exons 6 and 9 indicates that functional pressure was exerted on these sequences, probably by an interaction with proteins in skeletal muscle and perhaps also in smooth muscle; such a constraint was not detected in the sequences obtained from nonmuscle cells. These results have led us to postulate the existence of a protein in smooth muscle that may be the counterpart of skeletal muscle troponin. We show that different kinds of functional pressure were exerted on a single gene, resulting in different evolutionary rates and different convergences in some regions of the same molecule. Codon usage analysis indicates that there is no strict relationship between tissue types (and hence the tRNA precursor pool) and codon usage. G + C content is characteristic of a gene and does not change significantly during evolution.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Horizontal transfer of accessory chromosomes in fungi - a regulated process for exchange of genetic material?

Horizontal transfer of entire chromosomes has been reported in several fungal pathogens, often significantly impacting the fitness of the recipient fungus. All documented instances of horizontal chromosome transfers (HCTs) showed a marked propensity for accessory chromosomes, consistently involving the transfer of an accessory chromosome while other chromosomes were seldom, if ever, co-transferred. The mechanisms underlying HCTs, as well as the factors regulating the specificity of HCTs for accessory chromosomes, remain unclear. In this perspective, we provide an overview of the observed propensity in reported cases of horizontal chromosome transfers. We hypothesize the existence of a signal that distinguishes mobile, i.e., horizontally transferred, accessory chromosomes from the rest of the donor genome. Recent findings in Metarhizium robertsii and Magnaporthe oryzae, suggest that a mobile accessory chromosome may contain putative histones and/or histone modifiers, which could generate such a signal. Based on this, we propose that mobile accessory chromosomes may encode the machinery required for their own horizontal transmission, implying that HCT could be a regulated process. Finally, we present evidence of substantial differences in codon usage bias between core and accessory chromosomes in 14 out of 19 analysed fungal species and strains. Such differences in codon usage bias could indicate past horizontal transfers of these accessory chromosomes. Interestingly, HCT was previously unknown for many of these species, suggesting that the horizontal transfer of accessory chromosomes may be more widespread than previously thought, and therefore an important factor in fungal genome evolution.

Gene Transfer, Horizontal↗

Support vector machine for classification of meiotic recombination hotspots and coldspots in Saccharomyces cerevisiae based on codon composition.

BACKGROUND: Meiotic double-strand breaks occur at relatively high frequencies in some genomic regions (hotspots) and relatively low frequencies in others (coldspots). Hotspots and coldspots are receiving increasing attention in research into the mechanism of meiotic recombination. However, predicting hotspots and coldspots from DNA sequence information is still a challenging task. RESULTS: We present a novel method for classification of hot and cold ORFs located in hotspots and coldspots respectively in Saccharomyces cerevisiae, using support vector machine (SVM), which relies on codon composition differences. This method has achieved a high classification accuracy of 85.0%. Since codon composition is a fusion of codon usage bias and amino acid composition signals, the ability of these two kinds of sequence attributes to discriminate hot ORFs from cold ORFs was also investigated separately. Our results indicate that neither codon usage bias nor amino acid composition taken separately performed as well as codon composition. Moreover, our SVM based method was applied to the full genome: We predicted the hot/cold ORFs from the yeast genome by using cutoffs of recombination rate. We found that the performance of our method for predicting cold ORFs is not as good as that for predicting hot ORFs. Besides, we also observed a considerable correlation between meiotic recombination rate and amino acid composition of certain residues, which probably reflects the structural and functional dissimilarity between the hot and cold groups. CONCLUSION: We have introduced a SVM-based novel method to discriminate hot ORFs from cold ones. Applying codon composition as sequence attributes, we have achieved a high classification accuracy, which suggests that codon composition has strong potential to be used as sequence attributes in the prediction of hot and cold ORFs.

Algorithms↗

High-level expression of staphylococcal nuclease R gene in Escherichia coli.

Staphylococcal nuclease R, an analogue of nuclease A, was overproduced under the transcriptional control of the bacteriophage lambda PRPL promoters regulated by temperature sensitive repressors. The expression level reached 200-300 mg l-1 and showed little host dependence in different strains. The investigations of the recombinant nuclease R have revealed that the amino terminal formyl methionine residue of the nuclease is precisely processed, the protein consists of 155 amino acid residues. The experiment shows that the pBV221-DH5 alpha is a quite suitable vector-host system for high-level expression and precise processing of heterologous genes in Escherichia coli. The comparative studies between the codons used in the staphylococcal nuclease R gene and the optimal codon usage in E. coli indicate that high level expression of heterologous genes in E. coli may not always require a high degree of codon usage bias.

Amino Acid Sequence↗

Does the 'non-coding' strand code?

The hypothesis that DNA strands complementary to the coding strand contain in phase coding sequences has been investigated. Statistical analysis of the 50 genes of bacteriophage T7 shows no significant correlation between patterns of codon usage on the coding and non-coding strands. In Bacillus and yeast genes the correlation observed is not different from that expected with random synonymous codon usage, while a high correlation seen in 52 E. coli genes can be explained in terms of an excess of RNY codons. A deficiency of UUA, CUA and UCA codons (complementary to termination) seems to be restricted to the E. coli genes, and may be due to low abundance of the relevant cognate tRNA species. Thus the analysis shows that the non-coding strand has the properties expected of a sequence complementary to a coding strand, with no indications that it encodes, or may have encoded, proteins.

Bacillus↗

An estimate on the effect of point mutation and natural selection on the rate of amino acid replacement in proteins.

We outline a method for estimating quantitatively the influence of point mutations and selection on the frequencies of codons and amino acids. We show how the mutation rate, i.e., the rate of amino acid replacement due to point mutation, can be affected by the codon usage as well as by the rates of the involved base exchanges. A comparison of the mutation rates calculated from reliable values of codon usage and base exchange probabilities with those that would be expected on the basis of chance reveals a notable suppression of replacements leading to tryptophan, glutamate, lysine, and methionine, and particularly of those leading to the termination codons. If selection constraints are neglected and only mutations are taken into account, the best agreement between expected and observed frequencies of both codons and amino acids is obtained for alpha = 1.13-1.15, where (Formula: see text). The "selection values" of codons and amino acids derived by our method show a pattern that partially deviates from others in the literature. For example, the selection pressure on methionine and cysteine turns out to be much more pronounced than expected if only the discrepancies between their observed and expected occurrences in proteins are considered. To estimate to what extent randomly occurring amino acid replacements are accepted by selection, we constructed an "acceptability matrix" from the well-established matrix of accepted point mutations. On the basis of this matrix "acceptability values" of the amino acids can be defined that correlate with their selection values. We also examine the significance of mutations and selection of amino acids with respect to their physicochemical properties and functions in proteins. The conservatism of amino acid replacements with respect to certain properties such as polarity can be brought about by the mutational process alone, whereas the conservatism with respect to other relevant properties--among them all measures of bulkiness--obviously is the result of additional selectional constraints on the evolution of protein structures.

Amino Acid Sequence↗