PubMed HealthSearch

SEARCH · PubMed Health

Results for “synonymous codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

An analysis of the codon usage of Pasteurella haemolytica A1.

Analysis of approximately 17 kbp of nucleotide sequences from three different regions of the genome of Pasteurella haemolytica A1 showed that the mol% G+C of P. haemolytica A1 DNA is 38.5%. When only the coding sequences (approx. 10 kbp) were analysed, a similar value of 38.8% was obtained. A comparison of the relative synonymous codon usage values of the cloned genes showed that P. haemolytica A1 has a very different codon usage pattern from that of Escherichia coli.

Base Composition

Amino acid composition is correlated with protein abundance in Escherichia coli: can this be due to optimization of translational efficiency?

Amino acid occurrence frequencies were found for four groups of Escherichia coli proteins with different abundance levels in the cell. These frequencies decrease with increasing protein abundance for amino acids whose codons are translated by tRNAs present at low concentrations (e.g., Cys, Trp, Ser, etc.); the opposite tendency was observed for amino acids translated by abundant tRNAs (Lys, Val, etc.). The efficiency (rate and accuracy) of codon translation is expected to be proportional to the concentration of the cognate tRNA. Therefore, the observed constraints on amino acid composition may be explained as resulting from evolutionary pressure optimizing the translational efficiency of a gene (the same pressure is responsible for the nonrandom choice of synonymous codons).

Amino Acids

Genetic code redundancy and the evolutionary stability of protein secondary structure.

The genetic code has an inherent bias towards some amino acids because of the variable number of synonymous codons per amino acid. The extent to which these biases are expressed in protein secondary structure is described through the analysis of the overall amino acid compositions of the alpha-helix, beta-sheet, beta-turn and random coil segments elucidated by X-ray crystallography. Given the concept of neutral mutation in proteins, the allocation of synonyms in the genetic code appears to protect secondary structures from amino acid changes and discourages the appearance of chemically complex residues. The level of protection is similar for each structural form, despite their clear preferences for certain amino acids. The organization of the code is therefore relevant to the preservation of conformation seen in the evolution of many protein families.

Amino Acids

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae

Some aspects of the organization and evolution of the genetic code.

In this paper, I define a measure of the relative position of each amino acid in the genetic code by means of a 21-dimensional vector describing its potential for mutation, in a single step, to each of the other amino acids, or to a chain termination codon. This measure allows us to make a systematic investigation of the type and number of the physicochemical properties of the amino acids that were involved in evolution. The polar character and size of amino acids are identified in this analysis as properties that played a leading role in the evolutionary history of the genetic code. The application of cluster analysis and discriminant analysis reveals the characteristics of the structural organization of the genetic code. Finally, I suggest the existence of a relationship between the molecular weight of the amino acids and the number of synonymous codons.

Amino Acids

Conservation of the mammalian RNA polymerase II largest-subunit C-terminal domain.

We have isolated and sequenced a portion of the gene encoding the carboxy-terminal domain (CTD) of the largest subunit of RNA polymerase II from three mammals. These mammalian sequences include one rodent and two primate CTDs. Comparisons of the new sequences to mouse and Chinese hamster show a high degree of conservation among the mammalian CTDs. Due to synonymous codon usage, the nucleotide differences between hamster, rat, ape, and human result in no amino acid changes. The amino acid sequence for the mouse CTD appears to have one different amino acid when compared to the other four sequences. Therefore, except for the one variation in mouse, all of the known mammalian CTDs have identical amino acid sequences. This is in marked contrast to the situation among more divergent species. The present study suggests that there is a strong evolutionary pressure to maintain the primary structure of the mammalian CTD.

Animals

Nucleotide sequences from the colicin E8 operon: homology with plasmid ColE2-P9.

The primary structures of the immunity (Imm) and lysis (Lys) proteins, and the C-terminal 205 amino acid residues of colicin E8 were deduced from nucleotide sequencing of the 1,265 bp ClaI-PvuI DNA fragment of plasmid ColE8-J. The gene order is col-imm-lys confirming previous genetic data. A comparison of the colicin E8 peptide sequence with the available colicin E2-P9 sequence shows an identical receptor-binding domain but 20 amino acid replacements and a clustering of synonymous codon usage in the nuclease-active region. Sequence homology of the two colicins indicates that they are descended from a common ancestral gene and that colicin E8, like colicin E2, may also function as a DNA endonuclease. The native ColE8 imm (resident copy) is 258 bp long and is predicted to encode an acidic protein of 9,604 mol. wt. The six amino acid replacements between the resident imm and the previously reported non-resident copy of the ColE8 imm ([E8 imm]) found in the ribonuclease-producing ColE3-CA38 plasmid offer an explanation for the incomplete protection conferred by [E8 Imm] to exogenously added colicin E8. Except for one nucleotide and amino acid change in the putative signal peptide sequence, the ColE8 lys structure is identical to that present in ColE2-P9 and ColE3-CA38.

Amino Acid Sequence

Two genes encoding gas vacuole proteins in Halobacterium halobium.

The archaebacterium Halobacterium halobium contains two related gas vacuole protein-encoding genes (vac). One of these genes encodes a protein of 76 amino acids and resides on the major plasmid. The second gene is located on the chromosome in a (G + C)-rich DNA fraction and encodes a slightly larger but highly homologous protein consisting of 79 amino acids. The plasmid encoded vac gene is transcribed constitutively throughout the growth cycle while the chromosomal vac gene is expressed during the stationary phase of growth. Comparison of the nucleotide sequences of the two genes indicates differences in the putative promoter regions as well as 35 single base-pair exchanges within the coding regions of the two genes. The majority of the nucleotide exchanges in the coding region occur in the third position of a codon triplet generating the codon synonym. The only differences between the two encoded proteins are the exchange of 2 amino acids (positions 8 and 29) and a deletion of 3 amino acids near the carboxy-terminus of the plasmid encoded vac protein. The genomic DNAs from other halobacterial isolates (Halobacterium sp. SB3, GN101 and YC819-9) were found to contain only a chromosomal vac gene copy. There is a high conservation of the chromosomal vac gene and the genomic region surrounding it among the halobacterial strains investigated.

Amino Acid Sequence

Prokaryotic genetic code.

The prokaryotic genetic code has been influenced by directional mutation pressure (GC/AT pressure) that has been exerted on the entire genome. This pressure affects the synonymous codon choice, the amino acid composition of proteins and tRNA anticodons. Unassigned codons would have been produced in bacteria with extremely high GC or AT genomes by deleting certain codons and the corresponding tRNAs. A high AT pressure together with genomic economization led to a change in assignment of the UGA codon, from stop to tryptophan, in Mycoplasma.

Anticodon

Pattern of nucleotide substitution and the extent of purifying selection in retroviruses.

The patterns of point mutation and nucleotide substitution are inferred from nucleotide differences in three coding and two noncoding regions of retroviral genomes. Evidence is presented in favor of the view that the majority of mutations accumulate at the reverse transcription stage. Purifying selection is apparently very weak at the amino acid level, and almost nonexistent between synonymous codons. The pattern of purifying selection obeys the rules previously established in vertebrates [Gojobori T, Li W-H, Graur D (1982) J Mol Evol 18:360-369]; i.e., the magnitude of purifying selection at the amino acid level is negatively correlated with Grantham's [Grantham R (1974) Science 185: 862-864] chemical distances between the amino acids interchanged. We refute Modiano et al.'s [Modiano G, Battistuzzi G, Motulsky AG (1981) Proc Natl Acad Sci USA 78:1110-1114] hypothesis, according to which the pattern of mutation is preadapted to buffer against deleterious mutations. On the contrary, the pattern of mutation reduces the level of conservativeness from that imposed on the amino acid substitution pattern by the structure of the genetic code. The extraordinarily high rate of nucleotide substitution in retroviruses in comparison with that in other organisms is apparently caused by an extremely high rate of mutation coupled with a lack of stringent purifying selection at both the codon and the amino acid levels.

Animals

Evolution in bacteria: evidence for a universal substitution rate in cellular genomes.

This paper constructs a temporal scale for bacterial evolution by tying ecological events that took place at known times in the geological past to specific branch points in the genealogical tree relating the 16S ribosomal RNAs of eubacteria, mitochondria, and chloroplasts. One thus obtains a relationship between time and bacterial RNA divergence which can be used to estimate times of divergence between other branches in the bacterial tree. According to this approach, Salmonella typhimurium and Escherichia coli diverged between 120 and 160 million years (Myr) ago, a date which fits with evidence that the chief habitats occupied now by these two enteric species became available that long ago. The median extent of divergence between S. typhimurium and E. coli at synonymous sites for 21 kilobases of protein-coding DNA is 100%. This implies a silent substitution rate of 0.7-0.8%/Myr--a rate remarkably similar to that observed in the nuclear genes of mammals, invertebrates, and flowering plants. Similarities in the substitution rates of eucaryotes and procaryotes are not limited to silent substitutions in protein-coding regions. The average substitution rate for 16S rRNA in eubacteria is about 1%/50 Myr, similar to the average rate for 18S rRNA in vertebrates and flowering plants. Likewise, we estimate a mean rate of roughly 1%/25 Myr for 5S rRNA in both eubacteria and eucaryotes. For a few protein-coding genes of these enteric bacteria, the extent of silent substitution since the divergence of S. typhimurium and E. coli is much lower than 100%, owing to extreme bias in the usage of synonymous codons. Furthermore, in these bacteria, rates of amino acid replacement were about 20 times lower, on average, than the silent rate. By contrast, for the mammalian genes studied to date, the average replacement rate is only four to five times lower than the rate of silent substitution.

Bacteria

Incipient mitochondrial evolution in yeasts. II. The complete sequence of the gene coding for cytochrome b in Saccharomyces douglasii reveals the presence of both new and conserved introns and discloses major differences in the fixation of mutations in evolution.

We have determined the complete sequence of the mitochondrial gene coding for cytochrome b in Saccharomyces douglasii. The gene is 6310 base-pairs long and is interrupted by four introns. The first one (1311 base-pairs) belongs to the group ID of secondary structure, contains a fragment open reading frame with a characteristic GIY ... YIG motif, is absent from Saccharomyces cerevisiae and is inserted in the same site in which introns 1 and 2 are inserted in Neurospora crassa and Podospora anserina, respectively. The next three S. douglasii introns are homologous to the first three introns of S. cerevisiae, are inserted at the same positions and display various degrees of similarity ranging from an almost complete identity (intron 2 and 4) to a moderate one (intron 3). We have compared secondary structures of intron RNAs, and nucleotide and amino acid sequences of cytochrome b exons and intron open reading frames in the two Saccharomyces species. The rules that govern fixation of mutations in exon and intron open reading frames are different: the relative proportion of mutations occurring in synonymous codons is low in some introns and high in exons. The overall frequency of mutations in cytochrome b exons is much smaller than in nuclear genes of yeasts, contrary to what has been found in vertebrates, where mitochondrial mutations are more frequent. The divergence of the cytochrome b gene is modular: various parts of the gene have changed with a different mode and tempo of evolution.

Amino Acid Sequence

The significance of redundancy in the genetic code.

The genetic code has an inherent bias towards some amino acids because of the variable number of synonymous codons per amino acid. In proteins generally, this bias is expressed in the relative proportions of the twenty amino acids. It is suggested that even though neutral mutation may be responsible for the expression of this bias, the latter could be providing a positive advantage by directing mutation to introduce chemically simpler and more immutable amino acids where selective criteria have become relaxed.

Amino Acid Sequence

Periodicities and tandem repeats in a Balbiani ring gene.

The Balbiani ring (BR) DNAs show prominent periodicities of restriction enzyme sites. Studies using a cloned fragment of the BRc gene strongly suggest that these periodicities reflect the existence of tandemly repetitive sequences within BR DNA. Tandem repeats measuring 54-58 bp have been demonstrated by partial sequence analysis of the BRc clone; the restriction site periodicities suggest the existence of additional 175 (= 3 X 58) and 1050 (= 6 X 175) bp repeat units. The short, medium and long repeats (58, 175 and 1050 bp, respectively) show sequence homology. Constrained unequal crossing over (resulting from misalignment of repeat arrays, usually by one repeat) is proposed as the mechanism for evolution of short, medium and long repeats from each other, in a manner analogous to evolution of satellite DNA sequences. Paradoxically, the dominant restriction site periodicities appear to be more conservative than might be expected on the basis of the overall sequence divergence between the sequenced repeats. This may be a consequence of functionally important, long-range amino acid or oligopeptide periodicities (for example, Asp x Ser or Glu x Ser corresponding to Hinf I sites) in the BRc protein product, in conjunction with preferential use of certain synonymous codons.

Animals

Circumsporozoite gene of a Plasmodium falciparum strain from Thailand.

The nucleotide and deduced amino acid sequences of the CS gene of a Plasmodium falciparum strain from Thailand (T4) are presented. Comparison with the nucleotide sequences of two other P. falciparum CS genes, 7G8 from Brazil and Wellcome from West Africa, shows that: the coding regions outside the repeats of T4 and 7G8 are co-extensive and lack 30 nucleotides present in the Wellcome strain 5' to the repeats; in this region, T4 also differs at 3 nucleotide positions from the 7G8 and the Wellcome strains; in the region 3' to the repeats, T4 differs at two positions from 7G8 and at two other positions from the Wellcome strain--remarkably, all of these differences result in amino acid substitutions; the structure of the tandem repeats in the CS gene of T4 is, 5' to 3', [NANP-NVDP] X 3, [NANP] X 38, which is different from that of the two other strains. Due to the use of synonymous codons, the repetition of the sequence is more precise at the amino acid level than at the nucleotide level. These features contrast with those observed in the CS genes of other plasmodial species.

Amino Acid Sequence

Modification of mRNA secondary structure and alteration of the expression of human interferon alpha 1 in Escherichia coli.

A plasmid (pNL015) was constructed to contain a human interferon alpha 1 (IFN-alpha 1) gene under the transcriptional control of the Escherichia coli lipoprotein promoter. The E. coli cells harboring this plasmid produce 2.8 x 10(4) units/ml of IFN. Secondary structure analysis of the transcripts produced by pNL015 showed that the coding region could base pair with the Shine-Dalgarno (SD) region with a delta G = -3.9kcal/mol. A new plasmid pNL008 was constructed by modifying pNL015 with an 11-bp deletion and a 2-bp insertion in the coding region, so that the SD region is not involved in the secondary structure. E. coli cells harboring pNL008 produce ten times more IFN activity than cells harboring pNL015. A series of experiments were carried out to show that the specific activities of IFN, differential rates of IFN transcription, protein degradation or mRNA degradation could not account for the difference observed in expression. A rigorous test on this model of translational inhibition was conducted by the construction of pNL017 with a single bp substitution which did not change the amino acid sequence of the IFN (synonymous codon substitution) but which increased the calculated energy of interaction with the SD sequence to delta G = -10.8 kcal/mol. The E. coli cells harboring pNL017 produced no detectable IFN activity.

Base Sequence

The consensus sequence of ice nucleation proteins from Erwinia herbicola, Pseudomonas fluorescens and Pseudomonas syringae.

The consensus sequence of three bacterial ice nucleation proteins was determined by extrapolation from the nucleotide (nt) sequences of three ice nucleation-encoding genes, iceE (presented here), inaW and inaZ. The three proteins possess considerable similarity, so that a preferred amino acid is shown in most positions of the consensus. The corresponding genes show considerable divergence in the third nt positions of synonymous codons, suggesting that the proteins' conserved features have been maintained by selection. Therefore, the consensus sequence is likely to represent the components of primary structure most important to the ice nucleation function.

Amino Acid Sequence

Four synonymous genes encode calmodulin in the teleost fish, medaka (Oryzias latipes): conservation of the multigene one-protein principle.

We cloned four distinct calmodulin (CaM)-encoding cDNAs from a small teleost fish, medaka (Oryzias latipes). The deduced amino acid (aa) sequences were exactly the same in these four genes and identical to the aa sequence of mammalian CaM, because of synonymous codon usages. The four cDNAs from medaka, termed CaM-A, -B, -C and -D, corresponded to mRNAs of 1.8, 1.4, 2.5 and 1.8 kb, respectively, in Northern blot analysis. Our results demonstrated that the 'multigene one-protein' principle of CaM synthesis is applicable to medaka, as well as to mammals whose CaM is encoded by at least three different genes.

Amino Acid Sequence