PubMed HealthSearch

SEARCH · PubMed Health

Results for “synonymous codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Genetic code redundancy and the evolutionary stability of protein secondary structure.

The genetic code has an inherent bias towards some amino acids because of the variable number of synonymous codons per amino acid. The extent to which these biases are expressed in protein secondary structure is described through the analysis of the overall amino acid compositions of the alpha-helix, beta-sheet, beta-turn and random coil segments elucidated by X-ray crystallography. Given the concept of neutral mutation in proteins, the allocation of synonyms in the genetic code appears to protect secondary structures from amino acid changes and discourages the appearance of chemically complex residues. The level of protection is similar for each structural form, despite their clear preferences for certain amino acids. The organization of the code is therefore relevant to the preservation of conformation seen in the evolution of many protein families.

Amino Acids

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae

Some aspects of the organization and evolution of the genetic code.

In this paper, I define a measure of the relative position of each amino acid in the genetic code by means of a 21-dimensional vector describing its potential for mutation, in a single step, to each of the other amino acids, or to a chain termination codon. This measure allows us to make a systematic investigation of the type and number of the physicochemical properties of the amino acids that were involved in evolution. The polar character and size of amino acids are identified in this analysis as properties that played a leading role in the evolutionary history of the genetic code. The application of cluster analysis and discriminant analysis reveals the characteristics of the structural organization of the genetic code. Finally, I suggest the existence of a relationship between the molecular weight of the amino acids and the number of synonymous codons.

Amino Acids

Conservation of the mammalian RNA polymerase II largest-subunit C-terminal domain.

We have isolated and sequenced a portion of the gene encoding the carboxy-terminal domain (CTD) of the largest subunit of RNA polymerase II from three mammals. These mammalian sequences include one rodent and two primate CTDs. Comparisons of the new sequences to mouse and Chinese hamster show a high degree of conservation among the mammalian CTDs. Due to synonymous codon usage, the nucleotide differences between hamster, rat, ape, and human result in no amino acid changes. The amino acid sequence for the mouse CTD appears to have one different amino acid when compared to the other four sequences. Therefore, except for the one variation in mouse, all of the known mammalian CTDs have identical amino acid sequences. This is in marked contrast to the situation among more divergent species. The present study suggests that there is a strong evolutionary pressure to maintain the primary structure of the mammalian CTD.

Animals

Nucleotide sequences from the colicin E8 operon: homology with plasmid ColE2-P9.

The primary structures of the immunity (Imm) and lysis (Lys) proteins, and the C-terminal 205 amino acid residues of colicin E8 were deduced from nucleotide sequencing of the 1,265 bp ClaI-PvuI DNA fragment of plasmid ColE8-J. The gene order is col-imm-lys confirming previous genetic data. A comparison of the colicin E8 peptide sequence with the available colicin E2-P9 sequence shows an identical receptor-binding domain but 20 amino acid replacements and a clustering of synonymous codon usage in the nuclease-active region. Sequence homology of the two colicins indicates that they are descended from a common ancestral gene and that colicin E8, like colicin E2, may also function as a DNA endonuclease. The native ColE8 imm (resident copy) is 258 bp long and is predicted to encode an acidic protein of 9,604 mol. wt. The six amino acid replacements between the resident imm and the previously reported non-resident copy of the ColE8 imm ([E8 imm]) found in the ribonuclease-producing ColE3-CA38 plasmid offer an explanation for the incomplete protection conferred by [E8 Imm] to exogenously added colicin E8. Except for one nucleotide and amino acid change in the putative signal peptide sequence, the ColE8 lys structure is identical to that present in ColE2-P9 and ColE3-CA38.

Amino Acid Sequence

Two genes encoding gas vacuole proteins in Halobacterium halobium.

The archaebacterium Halobacterium halobium contains two related gas vacuole protein-encoding genes (vac). One of these genes encodes a protein of 76 amino acids and resides on the major plasmid. The second gene is located on the chromosome in a (G + C)-rich DNA fraction and encodes a slightly larger but highly homologous protein consisting of 79 amino acids. The plasmid encoded vac gene is transcribed constitutively throughout the growth cycle while the chromosomal vac gene is expressed during the stationary phase of growth. Comparison of the nucleotide sequences of the two genes indicates differences in the putative promoter regions as well as 35 single base-pair exchanges within the coding regions of the two genes. The majority of the nucleotide exchanges in the coding region occur in the third position of a codon triplet generating the codon synonym. The only differences between the two encoded proteins are the exchange of 2 amino acids (positions 8 and 29) and a deletion of 3 amino acids near the carboxy-terminus of the plasmid encoded vac protein. The genomic DNAs from other halobacterial isolates (Halobacterium sp. SB3, GN101 and YC819-9) were found to contain only a chromosomal vac gene copy. There is a high conservation of the chromosomal vac gene and the genomic region surrounding it among the halobacterial strains investigated.

Amino Acid Sequence

Prokaryotic genetic code.

The prokaryotic genetic code has been influenced by directional mutation pressure (GC/AT pressure) that has been exerted on the entire genome. This pressure affects the synonymous codon choice, the amino acid composition of proteins and tRNA anticodons. Unassigned codons would have been produced in bacteria with extremely high GC or AT genomes by deleting certain codons and the corresponding tRNAs. A high AT pressure together with genomic economization led to a change in assignment of the UGA codon, from stop to tryptophan, in Mycoplasma.

Anticodon

Evolution in bacteria: evidence for a universal substitution rate in cellular genomes.

This paper constructs a temporal scale for bacterial evolution by tying ecological events that took place at known times in the geological past to specific branch points in the genealogical tree relating the 16S ribosomal RNAs of eubacteria, mitochondria, and chloroplasts. One thus obtains a relationship between time and bacterial RNA divergence which can be used to estimate times of divergence between other branches in the bacterial tree. According to this approach, Salmonella typhimurium and Escherichia coli diverged between 120 and 160 million years (Myr) ago, a date which fits with evidence that the chief habitats occupied now by these two enteric species became available that long ago. The median extent of divergence between S. typhimurium and E. coli at synonymous sites for 21 kilobases of protein-coding DNA is 100%. This implies a silent substitution rate of 0.7-0.8%/Myr--a rate remarkably similar to that observed in the nuclear genes of mammals, invertebrates, and flowering plants. Similarities in the substitution rates of eucaryotes and procaryotes are not limited to silent substitutions in protein-coding regions. The average substitution rate for 16S rRNA in eubacteria is about 1%/50 Myr, similar to the average rate for 18S rRNA in vertebrates and flowering plants. Likewise, we estimate a mean rate of roughly 1%/25 Myr for 5S rRNA in both eubacteria and eucaryotes. For a few protein-coding genes of these enteric bacteria, the extent of silent substitution since the divergence of S. typhimurium and E. coli is much lower than 100%, owing to extreme bias in the usage of synonymous codons. Furthermore, in these bacteria, rates of amino acid replacement were about 20 times lower, on average, than the silent rate. By contrast, for the mammalian genes studied to date, the average replacement rate is only four to five times lower than the rate of silent substitution.

Bacteria

Incipient mitochondrial evolution in yeasts. II. The complete sequence of the gene coding for cytochrome b in Saccharomyces douglasii reveals the presence of both new and conserved introns and discloses major differences in the fixation of mutations in evolution.

We have determined the complete sequence of the mitochondrial gene coding for cytochrome b in Saccharomyces douglasii. The gene is 6310 base-pairs long and is interrupted by four introns. The first one (1311 base-pairs) belongs to the group ID of secondary structure, contains a fragment open reading frame with a characteristic GIY ... YIG motif, is absent from Saccharomyces cerevisiae and is inserted in the same site in which introns 1 and 2 are inserted in Neurospora crassa and Podospora anserina, respectively. The next three S. douglasii introns are homologous to the first three introns of S. cerevisiae, are inserted at the same positions and display various degrees of similarity ranging from an almost complete identity (intron 2 and 4) to a moderate one (intron 3). We have compared secondary structures of intron RNAs, and nucleotide and amino acid sequences of cytochrome b exons and intron open reading frames in the two Saccharomyces species. The rules that govern fixation of mutations in exon and intron open reading frames are different: the relative proportion of mutations occurring in synonymous codons is low in some introns and high in exons. The overall frequency of mutations in cytochrome b exons is much smaller than in nuclear genes of yeasts, contrary to what has been found in vertebrates, where mitochondrial mutations are more frequent. The divergence of the cytochrome b gene is modular: various parts of the gene have changed with a different mode and tempo of evolution.

Amino Acid Sequence

Circumsporozoite gene of a Plasmodium falciparum strain from Thailand.

The nucleotide and deduced amino acid sequences of the CS gene of a Plasmodium falciparum strain from Thailand (T4) are presented. Comparison with the nucleotide sequences of two other P. falciparum CS genes, 7G8 from Brazil and Wellcome from West Africa, shows that: the coding regions outside the repeats of T4 and 7G8 are co-extensive and lack 30 nucleotides present in the Wellcome strain 5' to the repeats; in this region, T4 also differs at 3 nucleotide positions from the 7G8 and the Wellcome strains; in the region 3' to the repeats, T4 differs at two positions from 7G8 and at two other positions from the Wellcome strain--remarkably, all of these differences result in amino acid substitutions; the structure of the tandem repeats in the CS gene of T4 is, 5' to 3', [NANP-NVDP] X 3, [NANP] X 38, which is different from that of the two other strains. Due to the use of synonymous codons, the repetition of the sequence is more precise at the amino acid level than at the nucleotide level. These features contrast with those observed in the CS genes of other plasmodial species.

Amino Acid Sequence

Modification of mRNA secondary structure and alteration of the expression of human interferon alpha 1 in Escherichia coli.

A plasmid (pNL015) was constructed to contain a human interferon alpha 1 (IFN-alpha 1) gene under the transcriptional control of the Escherichia coli lipoprotein promoter. The E. coli cells harboring this plasmid produce 2.8 x 10(4) units/ml of IFN. Secondary structure analysis of the transcripts produced by pNL015 showed that the coding region could base pair with the Shine-Dalgarno (SD) region with a delta G = -3.9kcal/mol. A new plasmid pNL008 was constructed by modifying pNL015 with an 11-bp deletion and a 2-bp insertion in the coding region, so that the SD region is not involved in the secondary structure. E. coli cells harboring pNL008 produce ten times more IFN activity than cells harboring pNL015. A series of experiments were carried out to show that the specific activities of IFN, differential rates of IFN transcription, protein degradation or mRNA degradation could not account for the difference observed in expression. A rigorous test on this model of translational inhibition was conducted by the construction of pNL017 with a single bp substitution which did not change the amino acid sequence of the IFN (synonymous codon substitution) but which increased the calculated energy of interaction with the SD sequence to delta G = -10.8 kcal/mol. The E. coli cells harboring pNL017 produced no detectable IFN activity.

Base Sequence

The consensus sequence of ice nucleation proteins from Erwinia herbicola, Pseudomonas fluorescens and Pseudomonas syringae.

The consensus sequence of three bacterial ice nucleation proteins was determined by extrapolation from the nucleotide (nt) sequences of three ice nucleation-encoding genes, iceE (presented here), inaW and inaZ. The three proteins possess considerable similarity, so that a preferred amino acid is shown in most positions of the consensus. The corresponding genes show considerable divergence in the third nt positions of synonymous codons, suggesting that the proteins' conserved features have been maintained by selection. Therefore, the consensus sequence is likely to represent the components of primary structure most important to the ice nucleation function.

Amino Acid Sequence

Four synonymous genes encode calmodulin in the teleost fish, medaka (Oryzias latipes): conservation of the multigene one-protein principle.

We cloned four distinct calmodulin (CaM)-encoding cDNAs from a small teleost fish, medaka (Oryzias latipes). The deduced amino acid (aa) sequences were exactly the same in these four genes and identical to the aa sequence of mammalian CaM, because of synonymous codon usages. The four cDNAs from medaka, termed CaM-A, -B, -C and -D, corresponded to mRNAs of 1.8, 1.4, 2.5 and 1.8 kb, respectively, in Northern blot analysis. Our results demonstrated that the 'multigene one-protein' principle of CaM synthesis is applicable to medaka, as well as to mammals whose CaM is encoded by at least three different genes.

Amino Acid Sequence

cDNA cloning and sequence determination of pig gastric (H+ + K+)-ATPase.

Complementary DNA to pig gastric mRNA encoding (H+ + K+)-ATPase was cloned, and its amino acid sequence was deduced from the nucleotide sequence. The enzyme contained 1034 amino acid residues (Mr. 114,285) including the initiation methionine. The sequence of pig (H+ + K+)-ATPase was highly homologous with that of the corresponding enzyme from rat, but had high degree of synonymous codon changes. Potential sites of phosphorylation by cAMP-dependent protein kinase and N-linked glycosylation sites were identified. The amino terminal region contained a lysine-rich sequence similar to that of the alpha subunit of (Na+ + K+)-ATPase, although a cluster of glycine residues was inserted into the sequence of the (H+ + K+)-ATPase. As the pig enzyme is advantageous for biochemical studies, the information of the primary structure is useful for further detailed studies.

Adenosine Triphosphatases

Prime numbers and the amino acid code: analogy in coding properties.

Natural numbers are characterized as being odd or even, prime or non-prime. If the quaternary information units of (DNA or RNA) nucleotide bases are assigned as 0 (for A), 1 (C), 2 (U or T) and 3 (G), then a unique set of amino acid numbers can be obtained by comparing the properties of numbers and coding properties. These numbers are: 0 for "stop" signals, 1 for Trp, 2 for Ile and 3 for Met. For other codons, synonymous quartets follow exclusively the P1 number series (prime numbers of the form 4n + 1); doublets mostly follow the P3 series (primes with quaternary remainder 3). A "one-to-one correspondence" between these numbers and the genetic code is established by considering their combinatorial specificities.

Amino Acid Sequence

Nucleotide sequence divergence in the -chain-structural genes of tryptophan synthetase from Escherichia coli, Salmonella typhimurium, and Aerobacter aerogenes.

Two different estimates were obtained for the extent of nucleotide sequence divergence in the structural genes of the tryptophan synthetase alpha-chains of Escherichia coli, Salmonella typhimurium, and Aerobacter aerogenes. One estimate was based on comparisons of the amino acid sequences of the respective alpha chains. The other was derived from measurements of the thermal stability of RNA-DNA hybrids formed with phage DNA carrying the alpha-chain structural gene of E. coli and labeled messenger RNA from the three bacterial species. Comparison of the two estimates suggests that during the course of evolution synonymous codon changes have accumulated in the alpha-chain-structural genes.

Amino Acid Sequence

Molecular evolution of human and rabbit beta-globin mRNAs.

The primary structures of human and rabbit beta-globin mRNAs are compared. Using as a standard the extent of nucleotide substitutions inferred from the hypervariable amino acid residues of fibrinopeptides A and B, which are thought to change largely by neutral evolution, we show that not all silent mutations in globin mRNA are neutral. The divergence of the sequences is limited in part by the selective usage of synonymous codons. The divergent nucleotides tend to be distributed nonrandomly: in the coding region silent substitutions are most rare in segments that are also deficient in substitutions leading to replacements.

Animals

Transient mutators: a semiquantitative analysis of the influence of translation and transcription errors on mutation rates.

A population of bacteria growing in a nonlimiting medium includes mutator bacteria and transient mutators defined as wild-type bacteria which, due to occasional transcription or translation errors, display a mutator phenotype. A semiquantitative theoretical analysis of the steady-state composition of an Escherichia coli population suggests that true strong genotypic mutators produce about 3 x 10(-3) of the single mutations arising in the population, while transient mutators produce at least 10% of the single mutations and more than 95% of the simultaneous double mutations. Numbers of mismatch repair proteins inherited by the offspring, proportions of lethal mutations and mortality rates are among the main parameters that influence the steady-state composition of the population. These results have implications for the experimental manipulation of mutation rates and the evolutionary fixation of frequent but nearly neutral mutations (e.g., synonymous codon substitutions).

Bacteria