PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “codon usage”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

The two gamma-tubulin-encoding genes of the ciliate Euplotes crassus differ in their sequences, codon usage, transcription initiation sites and poly(A) addition sites.

We have isolated and sequenced two gamma-tubulin (gamma-Tub)-encoding macronuclear genes of the ciliate Euplotes crassus (Ec), as well as their corresponding cDNAs. Our results reveal that the two genes (gamma-tub 1 and gamma-tub 2) have introns in homologous positions, but differ in their sequences, codon usage, transcription initiation sites and poly(A) addition sites. They both consist of three exons, two introns and two short non-coding sequences on both ends, and they both code for polypeptides of 462 amino acids (aa). The two genes share 76% identity at the nucleotide (nt) level, 86% at the deduced aa level and show 61-92% aa homology to the gamma-Tubs of other organisms. The gamma-tub 2 gene contains two in-frame UGA codons which, like UGA codons in other Euplotes genes, probably code for cysteines. No UGA triplet was found in the gamma-tub 1 gene. Further studies on the cDNA ends indicate that gamma-tub 1 uses at least three transcription initiation sites and two poly(A) addition sites. In contrast, only one transcription initiation site and one poly(A) addition site were identified in gamma-tub 2.

Amino Acid Sequence↗

Lactococcus lactis glyceraldehyde-3-phosphate dehydrogenase gene, gap: further evidence for strongly biased codon usage in glycolytic pathway genes.

The gene gap, encoding glyceraldehyde-3-phosphate dehydrogenase (EC 1.2.1.12), was isolated from a genomic library of Lactococcus lactis LM0230 DNA. Plasmids containing the L. lactis gene were able to complement a gap mutant of Escherichia coli. The nucleotide sequence of gap predicted a polypeptide chain of 337 amino acids for the enzyme and a subunit molecular mass of 36,043. The codon usage in gap and four other glycolytic genes from L. lactis showed a high degree of bias, when compared with 84 other chromosomal genes. Northern blot analysis of total L. lactis RNA showed that gap hybridized strongly with a 1.3 kb transcript. The 5' end of the transcript was determined by primer extension analysis to be a C located 35 bp upstream from the gap start codon. These transcript analyses, and the orientation of the open reading frames in the DNA flanking gap, indicated that in L. lactis gap is expressed on a monocistronic transcript. Nucleotide sequencing indicated that the DNA adjacent to gap did not encode other glycolytic pathway enzymes. The DNA sequence flanking gap contained two open reading frames (ORF156 and ORF211) of unknown function. The 3' end of a clpA homologue was identified in the sequence upstream of ORF156. The location of gap on the L. lactis DL11 chromosome map was determined to be between map coordinates 0.530 and 0.660.

Amino Acid Sequence↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

Spectinomycin operon of Micrococcus luteus: evolutionary implications of organization and novel codon usage.

The complete DNA sequence of the Micrococcus luteus spectinomycin (spc) operon and its adjacent regions has been determined. The sequence has revealed the presence of genes that are homologous to those of the Escherichia coli ribosomal and related proteins, L14, L24, L5, S8, L6, L18, S5, L30, L15, and secretion protein Y (sec Y), and the gene for adenylate kinase (adk). The gene arrangement in the spc operon is essentially the same as that of E. coli except for the absence in the M. luteus spc operon of the genes for S14 and X protein that exist in the E. coli spc operon. SecY and adk seem to be composed of another operon (adk operon) with at least an open reading frame. The deduced amino acid sequences for these ribosomal proteins are well conserved among the two species (40-65% identity). Reflecting the high genomic guanine and cytosine (GC) content of M. luteus (74%), the codon usage of the genes is extremely biased toward use of G and C, about 94% of the codon third positions being G or C. Seven codons, AUA, AAA, AGA, UUA, GUA, CUA, and CAA, all of which have A at the codon third positions, are completely absent in the M. luteus genes examined. Out of 11 genes in the M. luteus spc and adk operons, 5 (10) use GUG (UGA) and 6 (1) use AUG (UAA) as an initiation (termination) codon.

Amino Acid Sequence↗

Codon-usage variants in the polymorphic (GGN)n trinucleotide repeat of the human androgen receptor gene.

The human androgen receptor gene (hAR) has a long, polymorphic trinucleotide (GGN; glycine)n repeat in the 3' portion of its first exon, with n = 10-31. Owing to technical difficulties that have precluded routine sequencing of this region, it is widely unknown that N represents T, G or C, and that the usual sense codon sequence of the GGN tract is (GGT)3GGG(GGT)2(GGC)4-25. Furthermore, on 4 of 61 X chromosomes, we observed that the internal GGT sequence was present three or four times instead of twice. Strikingly, each of the three alleles with an internal (GGT)3, and only these three, also had a (GGC)20 repeat. The size or composition of a (GGN)n repeat was not correlated with the length of the accompanying (CAG)nCAA repeat in the 5' portion of exon one. Hence, codon-usage variants of the GGN tract may be used to seek associations with particular diseases, as diagnostic aids in families with androgen insensitivity whose AR mutations have not yet been identified, or as internal controls for observations on intergenerational contractions or expansions of the (CAG)nCAA tract in a given hAR allele.

Alleles↗

The problem of counting sites in the estimation of the synonymous and nonsynonymous substitution rates: implications for the correlation between the synonymous substitution rate and codon usage bias.

Most methods for estimating the rate of synonymous and nonsynonymous substitution per site define a site as a mutational opportunity: the proportion of sites that are synonymous is equal to the proportion of mutations that would be synonymous under the model of evolution being considered. Here we demonstrate that this definition of a site can give misleading results and that a physical definition of site should be used in some circumstances. We illustrate our point by reexamining the relationship between codon usage bias and the synonymous substitution rate. It has recently been shown that the rate of synonymous substitution, calculated using the Goldman-Yang method, which encapsulates the mutational-opportunity definition of a site at a high level of sophistication, is either positively correlated or uncorrelated to synonymous codon bias in Drosophila. Using other methods, which account for synonymous codon bias but define a site physically, we show that there is a negative correlation between the synonymous substitution rate and codon bias and that the lack of a negative correlation using the Goldman-Yang method is due to the way in which the number of synonymous sites is counted. We also show that there is a positive correlation between the synonymous substitution rate and third position GC content in mammals, but that the relationship is considerably weaker than that obtained using the Goldman-Yang method. We argue that the Goldman-Yang method is misleading in this context and conclude that methods that rely on a mutational-opportunity definition of a site should be used with caution.

Animals↗

Structural features of multiple nifH-like sequences and very biased codon usage in nitrogenase genes of Clostridium pasteurianum.

The structural gene (nifH1) encoding the nitrogenase iron protein of Clostridium pasteurianum has been cloned and sequenced. It is located on a 4-kilobase EcoRI fragment (cloned into pBR325) that also contains a portion of nifD and another nifH-like sequence (nifH2). C. pasteurianum nifH1 encodes a polypeptide (273 amino acids) identical to that of the isolated iron protein, indicating that the smaller size of the C. pasteurianum iron protein does not result from posttranslational processing. The 5' flanking region of nifH1 or nifH2 does not contain the nif promoter sequences found in several gram-negative bacteria. Instead, a sequence resembling the Escherichia coli consensus promoter (TTGACA-N17-TATAAT) is present before C. pasteurianum nifH2, and a TATAAT sequence is present before C pasteurianum nifH1. Codon usage in nifH1, nifH2, and nifD (partial) is very biased. A preference for A or U in the third position of the codons is seen. nifH2 could encode a protein of 272 amino acid residues, which differs from the iron protein (nifH1 product) in 23 amino acid residues (8%). Another nifH-like sequence (nifH3) is located on a nonadjacent EcoRI fragment and has been partially sequenced. C. pasteurianum nifH2 and nifH3 may encode proteins having several amino acids that are conserved in other proteins but not in C. pasteurianum iron protein, suggesting a possible role for the multiple nifH-like sequences of C. pasteurianum in the evolution of nifH. Among the nine sequenced iron proteins, only the C. pasteurianum protein lacks a conserved lysine residue which is near the extended C terminus of the other iron proteins. The absence of this positive charge in the C. pasteurianum iron protein might affect the cross-reactivity of the protein in heterologous systems.

Amino Acid Sequence↗

Support vector machines for separation of mixed plant-pathogen EST collections based on codon usage.

MOTIVATION: Discovery of host and pathogen genes expressed at the plant-pathogen interface often requires the construction of mixed libraries that contain sequences from both genomes. Sequence identification requires high-throughput and reliable classification of genome origin. When using single-pass cDNA sequences difficulties arise from the short sequence length, the lack of sufficient taxonomically relevant sequence data in public databases and ambiguous sequence homology between plant and pathogen genes. RESULTS: A novel method is described, which is independent of the availability of homologous genes and relies on subtle differences in codon usage between plant and fungal genes. We used support vector machines (SVMs) to identify the probable origin of sequences. SVMs were compared to several other machine learning techniques and to a probabilistic algorithm (PF-IND) for expressed sequence tag (EST) classification also based on codon bias differences. Our software (Eclat) has achieved a classification accuracy of 93.1% on a test set of 3217 EST sequences from Hordeum vulgare and Blumeria graminis, which is a significant improvement compared to PF-IND (prediction accuracy of 81.2% on the same test set). EST sequences with at least 50 nt of coding sequence can be classified using Eclat with high confidence. Eclat allows training of classifiers for any host-pathogen combination for which there are sufficient classified training sequences. AVAILABILITY: Eclat is freely available on the Internet (http://mips.gsf.de/proj/est) or on request as a standalone version. CONTACT: friedel@informatik.uni-muenchen.de.

Algorithms↗

Classification of Arabidopsis thaliana gene sequences: clustering of coding sequences into two groups according to codon usage improves gene prediction.

While genomic sequences are accumulating, finding the location of the genes remains a major issue that can be solved only for about a half of them by homology searches. Prediction methods are thus required, but unfortunately are not fully satisfying. Most prediction methods implicitly assume a unique model for genes. This is an oversimplification as demonstrated by the possibility to group coding sequences into several classes in Escherichia coli and other genomes. As no classification existed for Arabidopsis thaliana, we classified genes according to the statistical features of their coding sequences. A clustering algorithm using a codon usage model was developed and applied to coding sequences from A. thaliana, E. coli, and a mixture of both. By using it, Arabidopsis sequences were clustered into two classes. The CU1 and CU2 classes differed essentially by the choice of pyrimidine bases at the codon silent sites: CU2 genes often use C whereas CU1 genes prefer T. This classification discriminated the Arabidopsis genes according to their expressiveness, highly expressed genes being clustered in CU2 and genes expected to have a lower expression, such as the regulatory genes, in CU1. The algorithm separated the sequences of the Escherichia-Arabidopsis mixed data set into five classes according to the species, except for one class. This mixed class contained 89 % Arabidopsis genes from CU1 and 11 % E. coli genes, mostly horizontally transferred. Interestingly, most genes encoding organelle-targeted proteins, except the photosynthetic and photoassimilatory ones, were clustered in CU1. By tailoring the GeneMark CDS prediction algorithm to the observed coding sequence classes, its quality of prediction was greatly improved. Similar improvement can be expected with other prediction systems.

Algorithms↗

Human hemoglobin expression in Escherichia coli: importance of optimal codon usage.

The overexpression of a nonfusion product of human beta-globin in Escherichia coli from its cDNA sequence has been accomplished for the first time. Expression of beta-globin from its native cDNA required the use of the strong bacteriophage T7 promoter. In this system, beta-globin accumulated to approximately 10% of total E. coli proteins. alpha-Globin was not expressed in the T7 system using the native cDNA. For the expression of alpha-globin, synthetic genes containing optimal E. coli codons were constructed. Neither synthetic alpha- nor beta-globin gene alone was expressed from the lac or tac promoter. Globin expression was achieved when the two synthetic alpha- and beta-globin genes were combined as an operon downstream of the lac promoter. The two proteins combined intracellularly with endogenous heme, which was concomitantly overproduced to yield tetrameric hemoglobin as roughly 5-10% of total E. coli protein. Cloning the alpha- and beta-globin cDNAs in a construct identical with the lac promoter did not yield globin production, establishing the requirement for optimal codon usage. The recombinant beta-globin from the T7 expression system was purified and reconstituted in vitro with heme and native alpha chains. N-terminal analyses showed that the beta-globin produced in the T7 system and the tetrameric hemoglobin produced from the synthetic genes contained an additional beta 1 methionine residue. Two additional mutants, beta 1 Val----Met and beta 1 Val----Ala were produced using the T7 system. Functional and structural properties of the purified hemoglobins will be discussed in the following papers.

Amino Acid Sequence↗

Nucleotide sequence of simian virus 40 DNA: structure of the middle segment of the HindII + III restriction fragment B (sixth part of the T antigen gene) and codon usage.

We report here the nucleotide sequence of the simian virus 40 DNA region that lies between the EcoRII restriction endonuclease cleavage sites at map positions 0.214 and 0.281. The sequence was determined by partial chemical degradation of terminally labeled DNA fragments according to the procedure of Maxam and Gilbert. This region represents 6.7% of the SV40 genome and is located in the middle of HindII + III restriction fragment B. It is expressed as part of the early 19-S messenger RNA, which codes for the large-T antigen protein. Only one open reading frame for translation can be deduced from the message strand of the DNA and this reading frame connects in phase with the one of both neighboring fragments. This publication is the last in a series of papers about the T-antigen gene, and several properties of this gene and its product are discussed. The non-randomness of codon usage is similar to that previously discussed for the late part of the genome. Moreover, it appears that the choice of a third letter can be determined by the nature of the following codon; some codons which start with a pyrimidine are almost never preceded by an adenosine and some ANN-type codons are almost never preceded by a guanosine.

Amino Acid Sequence↗

Expression and codon usage optimization of the erythroid-specific transcription factor cGATA-1 in baculoviral and bacterial systems.

Biochemical characterization of cGATA-1, a key transcription factor in the regulation of globin expression in chickens, has been precluded by the unavailability of appreciable amounts of the pure protein. Purification directly from embryonic red blood cells has been limited by the difficulty in obtaining large quantities of the starting material, and previous attempts at bacterial expression have consistently yielded truncated product. To solve these problems, we have taken two approaches to the expression of cGATA-1. First, we were able to produce efficient expression from baculovirus-infected insect cells. Second, by altering the codon usage in cDNA encoding the protein's carboxy-terminal region, we obtained good expression of full-length protein in Escherichia coli. These preparations should prove useful in biochemical and structural studies of the factor. Additionally, we describe a primer extension/PCR-based method which can be used to synthesize extended regions of DNA sequence for gene construction.

Animals↗

Annotation pattern of ESTs from Spodoptera frugiperda Sf9 cells and analysis of the ribosomal protein genes reveal insect-specific features and unexpectedly low codon usage bias.

MOTIVATION: A whole set of Expressed Sequence Tags (ESTs) from the Sf9 cell line of Spodoptera frugiperda is presented here for the first time. By this way we want to identify both conserved and specific genes of this pest species. We also expect from this analysis to find a class of protein sequences providing a tool to explore genomic features and phylogeny of Lepidoptera. RESULTS: The ESTs display both housekeeping as well as developmentally regulated genes, and a high percentage of sequences with unknown function. Among the identified ORFs, almost all ribosomal proteins (RPs) were found with high EST redundancy and hence sequence accuracy. The codon usage found among RP genes is in average surprisingly much less biased in Lepidoptera than in other organisms. Other Spodoptera genes also displayed a low bias, suggesting a general genome expression feature in this Lepidoptera. We also found that the L35A and L36 RP sequences, respectively, display 40 and 10 amino-acid insertions, both being present only in insects. Sequence analysis suggests that they are probably not subjected to a strong selective pressure and may be good phylogenetic markers for Lepidoptera. Most interestingly, the Lepidoptera sequences of 9 RP genes displayed a specific signature different from the canonical one. We conclude that the RP family allows valuable comparative genomics and phylogeny of Lepidoptera. AVAILABILITY: All EST sequence data are available from the private 'Spodo-Base' upon request.

Abstracting and Indexing↗

Amino acid cost and codon-usage biases in 6 prokaryotic genomes: a whole-genome analysis.

For most prokaryotic organisms, amino acid biosynthesis represents a significant portion of their overall energy budget. The difference in the cost of synthesis between amino acids can be striking, differing by as much as 7-fold. Two prokaryotic organisms, Escherichia coli and Bacillus subtilis, have been shown to preferentially utilize less costly amino acids in highly expressed genes, indicating that parsimony in amino acid selection may confer a selective advantage for prokaryotes. This study confirms those findings and extends them to 4 additional prokaryotic organisms: Chlamydia trachomatis, Chlamydophila pneumoniae AR39, Synechocystis sp. PCC 6803, and Thermus thermophilus HB27. Adherence to codon-usage biases for each of these 6 organisms is inversely correlated with a coding region's average amino acid biosynthetic cost in a fashion that is independent of chemoheterotrophic, photoautotrophic, or thermophilic lifestyle. The obligate parasites C. trachomatis and C. pneumoniae AR39 are incapable of synthesizing many of the 20 common amino acids. Removing auxotrophic amino acids from consideration in these organisms does not alter the overall trend of preferential use of energetically inexpensive amino acids in highly expressed genes.

Adaptation, Biological↗

Evolution of amino-acid sequences and codon usage on the Drosophila miranda neo-sex chromosomes.

We have studied patterns of DNA sequence variation and evolution for 22 genes located on the neo-X and neo-Y chromosomes of Drosophila miranda. As found previously, nucleotide site diversity is greatly reduced on the neo-Y chromosome, with a severely distorted frequency spectrum. There is also an accelerated rate of amino-acid sequence evolution on the neo-Y chromosome. Comparisons of nonsynonymous and silent variation and divergence suggest that amino-acid sequences on the neo-X chromosome are subject to purifying selection, whereas this is much weaker on the neo-Y. The same applies to synonymous variants affecting codon usage. There is also an indication of a recent relaxation of selection on synonymous mutations for genes on other chromosomes. Genes that are weakly expressed on the neo-Y chromosome appear to have a faster rate of accumulation of both nonsynonymous and unpreferred synonymous mutations than genes with high levels of expression, although the rate of accumulation when both types of mutation are pooled is higher for the neo-Y chromosome than for the neo-X chromosome even for highly expressed genes.

Amino Acid Sequence↗

Sequence, codon usage and cysteine periodicity of the SerH1 gene and in the encoded surface protein of Tetrahymena thermophila.

The temperature-regulated SerH1 gene coding for an immunodominant surface glycoprotein (i-Ag H1) of Tetrahymena thermophila has been sequenced. The gene is reproducibly rearranged during macronuclear development and steady state mRNA levels are present at < 36 degrees C. The deduced i-Ag H1 amino acid (aa) sequence is rich in Ser, Thr and Cys, and contains three periods each consisting of 85 aa punctuated by eight Cys with the general formula, CX6CX17CX2CX18CX2CX11CX2CX19 (where X = any aa). Such Cys periodicity is common to ciliate i-Ag. Codon usage in Tt, Paramecium primaurelia and P. tetraurelia i-Ag encoding genes is similar, with approx. 80% A+T in the 3' position which is in marked contrast to the approx. 54% 3' A+T in other ciliate genes.

Amino Acid Sequence↗

Relationship between codon usage and sequence-dependent curvature of genomes.

Static DNA curvature distributions of full-sequenced genomes and large DNA contigs from different organisms were calculated. Very distinctive differences among histogram profiles coming from archaebacteria, eubacteria, and eukaryotes were observed. Eubacterial profiles were, on average, more curved than were archaeal and eukaryotic profiles. A comparative analysis between real and randomized DNA sequences revealed that eubacterial genomes presented, overall, higher curvature values than random sequences. An opposite portrait was exhibited by archaeal and eukaryotic genomes. They displayed a lower frequency of curved regions than their corresponding randomized sequences. The contributions of coding and intergenic regions to the curvature profile were also analyzed. Intergenic regions, on average, were found to be more curved than the overall genomic sequences, especially in prokaryotic organisms. Nevertheless, because of their small size with respect to coding regions, the contribution of intergenic sequences to the overall curvature profile tended to be minor. A clear relationship between codon usage and DNA curvature was demonstrated, and a proposal of the possible coevolution of both systems is discussed. Finally, we present a procedure to quantify the deviation of a curvature profile from randomness through a formal statistical analysis.

Archaea↗

The codon usage of the nisZ operon in Lactococcus lactis N8 suggests a non-lactococcal origin of the conjugative nisin-sucrose transposon.

An 11.6 kb area downstream from the structural gene of nisin Z in the conjugative nisin-sucrose transposon of Lactococcus lactis subsp. lactis N8 was cloned and sequenced. Analysis of the sequence revealed eight open reading frames, nisZBTClPRK, followed by a putative rho-independent terminator (delta G degrees = -4.7 kcal/mol). The C-terminal hydrophilic domain of the NisK protein is homologous to the C-termini of several histidine kinases of bacterial two-component regulator systems, such as SpaK from Bacillus subtilis and KdpD and RcsC of Escherichia coli. The nisin Z biosynthetic genes were highly similar with the genes of the nisin A operons having, however, a 0-3% difference in the amino acid sequences of the individual proteins. The codon usage of eleven genes within the same conjugative transposon was calculated and found to be strikingly different from that of other lactococcal genes. This, together with the low GC-content (32%) compared to the 38% (G+C) of the lactococcal chromosome in general strongly suggests a non-lactococcal origin of this transposon.

Amino Acid Sequence↗