PubMed HealthSearch

SEARCH · PubMed Health

Results for “synonymous codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Distribution and evolution of sequence characteristics in the E. coli genome.

The mean (G + C) composition (51.0%) and standard deviation (+/- 3.8%) of published DNA sequences accounting for 10% of the E. coli genome is in excellent agreement with the principal overall distribution determined by high resolution melting. While differences in base and neighbor characteristics are small and uniform throughout all regions of the genome, it is found that the (G + C) content of sequences varies in segmented fashion within boundaries corresponding to coding (53% G + C) and noncoding (46% G + C) regions; with variances in the latter being six-fold greater than in coding regions. The variance in different regions shows a strong negative dependence on (G + C) content of the region, reflecting the condition that A-T and G-C base pairs are preferred neighbors of A-T and C-G pairs, respectively; with the bias increasing with decreasing (G + C) content. Neighbor analysis indicates the most extreme positive biases occur in AA, TT, GC and CG throughout all regions, but particularly in noncoding regions. Extraordinary numbers of oligomeric strings of (A)n, etc., are the further consequence of this bias. These and other characteristics point to the existence of inherent biases in neighbor frequencies levied during replication or repair, and which reflect, in turn, neighbor influences during mutation. The bias in codon usage noted by Grantham and others is seen here as due, in part, to the adaptation of coding sequences to this microenvironment through selection among synonymous codons so as to preserve inherent neighbor biases.

Base Composition

Transgene sequence codon optimization and composition determines replication competence of self-amplifying RNA.

Self-amplifying RNA (saRNA) is an emerging RNA therapeutic modality that can facilitate higher magnitude and more durable protein expression at substantially lower doses than nonreplicating mRNA. Unlike conventional messenger RNA (mRNA), alphavirus-derived saRNA must support a replicase-driven RNA amplification step in addition to translation, raising the possibility that transgene coding sequences impose sequence-level constraints on replication. Here, saRNA replication was found to be dependent on the codon composition of the transgene; multiple therapeutic transgenes were replication defective despite an intact Venezuelan Equine Encephalitis Virus (VEEV)-derived saRNA backbone. Replication defects were rescued by synonymous codon re-optimization of the same transgenes, indicating that nucleotide-level features of the coding sequence, rather than the encoded protein, govern replication competence. Comparative compositional analyses identified a distinct signature associated with productive replication, characterized by elevated GC (>53%) and GC3 (>63%) content, higher codon adaptation to human (>0.75), and reduced UpA (<43/kb) and UpU (<41/kb) dinucleotide density. Moreover, deliberate compositional perturbation of an otherwise replication-competent transgene shifted these features and abolished replication, supporting a causal and combinatorial role for sequence composition in defining saRNA replication outcome. These findings define an underappreciated constraint in saRNA therapeutics and motivate saRNA-specific payload design frameworks that incorporate alphavirus-associated compositional biases during transgene sequence optimization.

Codon

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence

Codon usage in regulatory genes in Escherichia coli does not reflect selection for 'rare' codons.

It has often been suggested that differential usage of codons recognized by rare tRNA species, i.e. "rare codons", represents an evolutionary strategy to modulate gene expression. In particular, regulatory genes are reported to have an extraordinarily high frequency of rare codons. From E. coli we have compiled codon usage data for highly expressed genes, moderately/lowly expressed genes, and regulatory genes. We have identified a clear and general trend in codon usage bias, from the very high bias seen in very highly expressed genes and attributed to selection, to a rather low bias in other genes which seems to be more influenced by mutation than by selection. There is no clear tendency for an increased frequency of rare codons in the regulatory genes, compared to a large group of other moderately/lowly expressed genes with low codon bias. From this, as well as a consideration of evolutionary rates of regulatory genes, and of experimental data on translation rates, we conclude that the pattern of synonymous codon usage in regulatory genes reflects primarily the relaxation of natural selection.

Base Sequence

Codon usage in plant genes.

We have examined codon bias in 207 plant gene sequences collected from Genbank and the literature. When this sample was further divided into 53 monocot and 154 dicot genes, the pattern of relative use of synonymous codons was shown to differ between these taxonomic groups, primarily in the use of G + C in the degenerate third base. Maize and soybean codon bias were examined separately and followed the monocot and dicot codon usage patterns respectively. Codon preference in ribulose 1,5 bisphosphate and chlorophyll a/b binding protein, two of the most abundant proteins in leaves was investigated. These highly expressed are more restricted in their codon usage than plant genes in general.

Amino Acid Sequence

Processes of genome evolution reflected by base frequency differences among Serratia marcescens genes.

The G + C content of silent sites in codons varies greatly among Serratia marcescens genes; the value in any one gene seems to reflect a balance between mutation pressure towards high G + C content and natural selection constraining choice among synonymous codons. Interestingly, non-coding sequences have substantially lower G + C content than silent sites thought to be under little selective constraint.

Base Composition

Two types of linkage between codon usage and gene-expression levels.

The relation between codon usage and gene-expression levels is an intensively investigated and discussed topic in the field of molecular evolution. We statistically analyzed 25 Escherichia coli gene sequences by a new classification of synonymous codons and found that (i) there are two distinct types of linkage between codon usage and gene-expression levels in E. coli, and (ii) one of the two kinds of codon preferences (the codon preference concerned with interaction of GC/AT choice at three codon positions) is observed significantly in weakly expressed genes.

Base Sequence

Codon bias variation in Staphylococcus aureus.

BACKGROUND: Staphylococcus aureus causes a multiplicity of human diseases acquired in community and healthcare settings alike around the globe. While most studies focus on coding changes to assess genome evolution and study genetic adaptation, interrogation of silent mutations in the form of synonymous codon usage bias is less well-studied. As such, understanding of patterns in codon bias at the gene and genome levels, and how codon bias impacts protein expression in S. aureus remains incomplete. METHODS: The codon bias of 2,565 protein encoding genes from NCTC 8325 was queried against all publicly available closed S. aureus genomes. Using public BioSample data, genomes were sorted by disease state, submitting institution, and collection site. Codon bias was assessed at the level of gene and genome using the codon adaptation index (CAI), calculated using 30S and 50S ribosomal genes. Gene set enrichment analysis was applied to determine associations between physiological functions, CAI gene scores, and interquartile ranges. CAI scores were also compared to an in vitro S. aureus proteomics database to correlate codon bias and protein expression. RESULTS: CAI scores varied within and between isolates at the gene and genome levels. Genes with ribosome-associated functions were most enriched among high CAI genes, and had low CAI interquartile ranges (IQR), suggesting selective pressure to maintain high expression of these genes across all S. aureus isolates. Genome sequences submitted by Aga Khan University Hospital, Nairobi, Kenya were most different from others. For the LAC USA 300 strain, CAI and protein expression were moderately positively correlated (cor&#x2009;=&#x2009;0.534, p&#x2009;<&#x2009;2.2e-16). CONCLUSIONS: Codon bias in S. aureus was shown to vary between gene, and to be a source of genetic variation between isolates; CAI and in vitro protein expression were positively correlated.

Staphylococcus aureus

Natural selection versus primitive gene structure as determinant of codon usage.

Different codons are not utilized equally in known gene sequences. One of the important biases of codon usage is observed in the form of an enrichment of RNY codons, especially within RNN codon families. Such biases could represent the residue of a primitive repeating-RNY gene structure, or the outcome of natural selection, or both. Analyses based on the rates of silent substitutions, the frequencies of base doublets, and synonymous codon ratios for Escherichia coli, yeast, Drosophila and Xenopus proteins have been performed. The results rule out any significant support for a primitive repeating-RNY or repeating-RRY gene structure, and establish the important role of natural selection in determining the choice of codons. With strong intervention by natural selection, the relationship between primitive gene structure and codon usage necessarily becomes minimal.

Animals

An analysis of the codon usage of Pasteurella haemolytica A1.

Analysis of approximately 17 kbp of nucleotide sequences from three different regions of the genome of Pasteurella haemolytica A1 showed that the mol% G+C of P. haemolytica A1 DNA is 38.5%. When only the coding sequences (approx. 10 kbp) were analysed, a similar value of 38.8% was obtained. A comparison of the relative synonymous codon usage values of the cloned genes showed that P. haemolytica A1 has a very different codon usage pattern from that of Escherichia coli.

Base Composition

Amino acid composition is correlated with protein abundance in Escherichia coli: can this be due to optimization of translational efficiency?

Amino acid occurrence frequencies were found for four groups of Escherichia coli proteins with different abundance levels in the cell. These frequencies decrease with increasing protein abundance for amino acids whose codons are translated by tRNAs present at low concentrations (e.g., Cys, Trp, Ser, etc.); the opposite tendency was observed for amino acids translated by abundant tRNAs (Lys, Val, etc.). The efficiency (rate and accuracy) of codon translation is expected to be proportional to the concentration of the cognate tRNA. Therefore, the observed constraints on amino acid composition may be explained as resulting from evolutionary pressure optimizing the translational efficiency of a gene (the same pressure is responsible for the nonrandom choice of synonymous codons).

Amino Acids

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae

Some aspects of the organization and evolution of the genetic code.

In this paper, I define a measure of the relative position of each amino acid in the genetic code by means of a 21-dimensional vector describing its potential for mutation, in a single step, to each of the other amino acids, or to a chain termination codon. This measure allows us to make a systematic investigation of the type and number of the physicochemical properties of the amino acids that were involved in evolution. The polar character and size of amino acids are identified in this analysis as properties that played a leading role in the evolutionary history of the genetic code. The application of cluster analysis and discriminant analysis reveals the characteristics of the structural organization of the genetic code. Finally, I suggest the existence of a relationship between the molecular weight of the amino acids and the number of synonymous codons.

Amino Acids

Conservation of the mammalian RNA polymerase II largest-subunit C-terminal domain.

We have isolated and sequenced a portion of the gene encoding the carboxy-terminal domain (CTD) of the largest subunit of RNA polymerase II from three mammals. These mammalian sequences include one rodent and two primate CTDs. Comparisons of the new sequences to mouse and Chinese hamster show a high degree of conservation among the mammalian CTDs. Due to synonymous codon usage, the nucleotide differences between hamster, rat, ape, and human result in no amino acid changes. The amino acid sequence for the mouse CTD appears to have one different amino acid when compared to the other four sequences. Therefore, except for the one variation in mouse, all of the known mammalian CTDs have identical amino acid sequences. This is in marked contrast to the situation among more divergent species. The present study suggests that there is a strong evolutionary pressure to maintain the primary structure of the mammalian CTD.

Animals

Nucleotide sequences from the colicin E8 operon: homology with plasmid ColE2-P9.

The primary structures of the immunity (Imm) and lysis (Lys) proteins, and the C-terminal 205 amino acid residues of colicin E8 were deduced from nucleotide sequencing of the 1,265 bp ClaI-PvuI DNA fragment of plasmid ColE8-J. The gene order is col-imm-lys confirming previous genetic data. A comparison of the colicin E8 peptide sequence with the available colicin E2-P9 sequence shows an identical receptor-binding domain but 20 amino acid replacements and a clustering of synonymous codon usage in the nuclease-active region. Sequence homology of the two colicins indicates that they are descended from a common ancestral gene and that colicin E8, like colicin E2, may also function as a DNA endonuclease. The native ColE8 imm (resident copy) is 258 bp long and is predicted to encode an acidic protein of 9,604 mol. wt. The six amino acid replacements between the resident imm and the previously reported non-resident copy of the ColE8 imm ([E8 imm]) found in the ribonuclease-producing ColE3-CA38 plasmid offer an explanation for the incomplete protection conferred by [E8 Imm] to exogenously added colicin E8. Except for one nucleotide and amino acid change in the putative signal peptide sequence, the ColE8 lys structure is identical to that present in ColE2-P9 and ColE3-CA38.

Amino Acid Sequence

Two genes encoding gas vacuole proteins in Halobacterium halobium.

The archaebacterium Halobacterium halobium contains two related gas vacuole protein-encoding genes (vac). One of these genes encodes a protein of 76 amino acids and resides on the major plasmid. The second gene is located on the chromosome in a (G + C)-rich DNA fraction and encodes a slightly larger but highly homologous protein consisting of 79 amino acids. The plasmid encoded vac gene is transcribed constitutively throughout the growth cycle while the chromosomal vac gene is expressed during the stationary phase of growth. Comparison of the nucleotide sequences of the two genes indicates differences in the putative promoter regions as well as 35 single base-pair exchanges within the coding regions of the two genes. The majority of the nucleotide exchanges in the coding region occur in the third position of a codon triplet generating the codon synonym. The only differences between the two encoded proteins are the exchange of 2 amino acids (positions 8 and 29) and a deletion of 3 amino acids near the carboxy-terminus of the plasmid encoded vac protein. The genomic DNAs from other halobacterial isolates (Halobacterium sp. SB3, GN101 and YC819-9) were found to contain only a chromosomal vac gene copy. There is a high conservation of the chromosomal vac gene and the genomic region surrounding it among the halobacterial strains investigated.

Amino Acid Sequence

Prokaryotic genetic code.

The prokaryotic genetic code has been influenced by directional mutation pressure (GC/AT pressure) that has been exerted on the entire genome. This pressure affects the synonymous codon choice, the amino acid composition of proteins and tRNA anticodons. Unassigned codons would have been produced in bacteria with extremely high GC or AT genomes by deleting certain codons and the corresponding tRNAs. A high AT pressure together with genomic economization led to a change in assignment of the UGA codon, from stop to tryptophan, in Mycoplasma.

Anticodon

Evolution in bacteria: evidence for a universal substitution rate in cellular genomes.

This paper constructs a temporal scale for bacterial evolution by tying ecological events that took place at known times in the geological past to specific branch points in the genealogical tree relating the 16S ribosomal RNAs of eubacteria, mitochondria, and chloroplasts. One thus obtains a relationship between time and bacterial RNA divergence which can be used to estimate times of divergence between other branches in the bacterial tree. According to this approach, Salmonella typhimurium and Escherichia coli diverged between 120 and 160 million years (Myr) ago, a date which fits with evidence that the chief habitats occupied now by these two enteric species became available that long ago. The median extent of divergence between S. typhimurium and E. coli at synonymous sites for 21 kilobases of protein-coding DNA is 100%. This implies a silent substitution rate of 0.7-0.8%/Myr--a rate remarkably similar to that observed in the nuclear genes of mammals, invertebrates, and flowering plants. Similarities in the substitution rates of eucaryotes and procaryotes are not limited to silent substitutions in protein-coding regions. The average substitution rate for 16S rRNA in eubacteria is about 1%/50 Myr, similar to the average rate for 18S rRNA in vertebrates and flowering plants. Likewise, we estimate a mean rate of roughly 1%/25 Myr for 5S rRNA in both eubacteria and eucaryotes. For a few protein-coding genes of these enteric bacteria, the extent of silent substitution since the divergence of S. typhimurium and E. coli is much lower than 100%, owing to extreme bias in the usage of synonymous codons. Furthermore, in these bacteria, rates of amino acid replacement were about 20 times lower, on average, than the silent rate. By contrast, for the mammalian genes studied to date, the average replacement rate is only four to five times lower than the rate of silent substitution.

Bacteria