PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “synonymous codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Molecular evolution of human and rabbit beta-globin mRNAs.

The primary structures of human and rabbit beta-globin mRNAs are compared. Using as a standard the extent of nucleotide substitutions inferred from the hypervariable amino acid residues of fibrinopeptides A and B, which are thought to change largely by neutral evolution, we show that not all silent mutations in globin mRNA are neutral. The divergence of the sequences is limited in part by the selective usage of synonymous codons. The divergent nucleotides tend to be distributed nonrandomly: in the coding region silent substitutions are most rare in segments that are also deficient in substitutions leading to replacements.

Animals↗

Rapid evolution of sex-related genes in Chlamydomonas.

Biological speciation ultimately results in prezygotic isolation-the inability of incipient species to mate with one another-but little is understood about the selection pressures and genetic changes that generate this outcome. The genus Chlamydomonas comprises numerous species of unicellular green algae, including numerous geographic isolates of the species C. reinhardtii. This diverse collection has allowed us to analyze the evolution of two sex-related genes: the mid gene of C. reinhardtii, which determines whether a gamete is mating-type plus or minus, and the fus1 gene, which dictates a cell surface glycoprotein utilized by C. reinhardtii plus gametes to recognize minus gametes. Low stringency Southern analyses failed to detect any fus1 homologs in other Chlamydomonas species and detected only one mid homolog, documenting that both genes have diverged extensively during the evolution of the lineage. The one mid homolog was found in C. incerta, the species in culture that is most closely related to C. reinhardtii. Its mid gene carries numerous nonsynonymous and synonymous codon changes compared with the C. reinhardtii mid gene. In contrast, very high sequence conservation of both the mid and fus1 sequences is found in natural isolates of C. reinhardtii, indicating that the genes are not free to drift within a species but do diverge dramatically between species. Striking divergence of sex determination and mate recognition genes also has been encountered in a number of other eukaryotic phyla, suggesting that unique, and as yet unidentified, selection pressures act on these classes of genes during the speciation process.

Amino Acid Sequence↗

Problems in protein biosynthesis.

Outline of the steps in protein synthesis. Nature of the genetic code. The use of synthetic oligo- and polynucleotides in deciphering the code. Structure of the code: relatedness of synonym codons. The wobble hypothesis. Chain initiation and N-formyl-methionine. Chain termination and nonsense codons. Mistakes in translation: ambiguity in vitro. Suppressor mutations resulting in ambiguity. Limitations in the universality of the code. Attempts to determine the particular codons used by a species. Mechanisms of suppression, caused by (a) abnormal aminoacyl-tRNA, (b) ribosomal malfunction. Effect of streptomycin. The problem of "reading" a nucleic acid template. Different ribosomal mutants and DNA polymerase mutants might cause different mistakes. The possibility of involvement of allosteric proteins in template reading.

Genetic Code↗

Causal analysis of CpG suppression in the Mycoplasma genome.

Some bacterial genomes are known to have low CpG dinucleotide frequencies. While their causes are not clearly understood, the frequency of CpG is suppressed significantly in the genome of Mycoplasma genitalium, but not in that of Mycoplasma pneumoniae. We compared orthologous gene pairs of the two closely related species to analyze CpG substitution patterns between these two genomes. We also divided genome sequences into three regions: protein-coding, noncoding, and RNA-coding, and obtained the CpG frequencies for each region for each organism. It was found that the observed/expected ratio of CpG dinucleotides is low in both the protein-coding and noncoding regions; while that ratio is in the normal range in the RNA-coding region. Our results indicate that CpG suppression of the Mycoplasma genome is not caused by (1) biased usage amino acid; (2) biased usage of synonymous codon; or (3) methylation effects by the CpG methyltransferase in the genomes of their hosts. Instead, we consider it likely that a certain global pressure, such as genome-wide pressure for the advantages of DNA stability or replication, has the effect of decreasing CpG over the entire genome, which, in turn, resulted in the biased codon usage.

Base Composition↗

FISH: a guide to protein-coding DNA sequences in the GenBank database.

FISH (Fast Index Search for Homologous coding sequences) consists of a database and associated software and is intended to function as a directory of protein-coding gene sequences. The FISH index contains descriptions of 22,361 DNA sequences from release 69.0 of the GenBank genetic sequence database. Complete coding sequences are represented numerically with counts of nucleotides and synonymous codons, and with GenBank LOCUS names and short descriptions. The software permits the database to be queried by GenBank LOCUS name, sequence length (expressed as total number of codons), or by comparison with a DNA sequence. In the latter case, the numerical descriptions are compared with simple distance measures in place of actual DNA sequences. The FISH package can be used to rapidly assemble lists of similar coding sequences, without regard to functional annotation or sequence alignments. Typical search times are well under a minute on widely available IBM-compatible microcomputers.

Algorithms↗

Transient mutators: a semiquantitative analysis of the influence of translation and transcription errors on mutation rates.

A population of bacteria growing in a nonlimiting medium includes mutator bacteria and transient mutators defined as wild-type bacteria which, due to occasional transcription or translation errors, display a mutator phenotype. A semiquantitative theoretical analysis of the steady-state composition of an Escherichia coli population suggests that true strong genotypic mutators produce about 3 x 10(-3) of the single mutations arising in the population, while transient mutators produce at least 10% of the single mutations and more than 95% of the simultaneous double mutations. Numbers of mismatch repair proteins inherited by the offspring, proportions of lethal mutations and mortality rates are among the main parameters that influence the steady-state composition of the population. These results have implications for the experimental manipulation of mutation rates and the evolutionary fixation of frequent but nearly neutral mutations (e.g., synonymous codon substitutions).

Bacteria↗

Inferring the fitness effects of DNA mutations from polymorphism and divergence data: statistical power to detect directional selection under stationarity and free recombination.

The fitness effects of classes of DNA mutations can be inferred from patterns of nucleotide variation. A number of studies have attributed differences in levels of polymorphism and divergence between silent and replacement mutations to the action of natural selection. Here, I investigate the statistical power to detect directional selection through contrasts of DNA variation among functional categories of mutations. A variety of statistical approaches are applied to DNA data simulated under Sawyer and Hartl's Poisson random field model. Under assumptions of free recombination and stationarity, comparisons that include both the frequency distributions of mutations segregating within populations and the numbers of mutations fixed between populations have substantial power to detect even very weak selection. Frequency distribution and divergence tests are applied to silent and replacement mutations among five alleles of each of eight Drosophila simulans genes. Putatively "preferred" silent mutations segregate at higher frequencies and are more often fixed between species than "unpreferred" silent changes, suggesting fitness differences among synonymous codons. Amino acid changes tend to be either rare polymorphisms or fixed differences, consistent with a combination of deleterious and adaptive protein evolution. In these data, a substantial fraction of both silent and replacement DNA mutations appear to affect fitness.

Adaptation, Biological↗

Ancient allelism at the cytosolic chaperonin-alpha-encoding gene of the zebrafish.

The T-complex protein 1, TCP1, gene codes for the CCT-alpha subunit of the group II chaperonins. The gene was first described in the house mouse, in which it is closely linked to the T locus at a distance of approximately 11 cM from the Mhc. In the zebrafish, Danio rerio, in which the T homolog is linked to the class I Mhc loci, the TCP1 locus segregates independently of both the T and the Mhc loci. Despite its conservation between species, the zebrafish TCP1 locus is highly polymorphic. In a sample of 15 individuals and the screening of a cDNA library, 12 different alleles were found, and some of the allelic pairs were found to differ by up to nine nucleotides in a 275-bp-long stretch of sequence. The substitutions occur in both translated and untranslated regions, but in the former they occur predominantly at synonymous codon sites. Phylogenetically, the alleles fall into two groups distinguished also by the presence or absence of a 10-bp insertion/deletion in the 3' untranslated region. The two groups may have diverged as long as 3.5 mya, and the polymorphic differences may have accumulated by genetic drift in geographically isolated populations.

Alleles↗

The nucleotide sequence of myosin light chain (L-2A) mRNA from embryonic chicken cardiac muscle tissue.

The nucleotide sequence of a cDNA clone (pML10) for chicken cardiac myosin light chain is described. The cDNA insert contains 613 nucleotides representing the entire coding sequence, with the exception of nine NH2-terminal amino acids, and the full 3'-non-coding region of 146 nucleotides. The missing 5' terminus of the mRNA, not represented in the clone pML10, was obtained by extension of the cDNA using a 43 nucleotide long internal EcoR1 fragment as a primer. The non-coding region contains several direct and inverted repeated sequences and the polyadenylation signal sequence AATAAA. The coding portion exhibits non-random usage of synonymous codons with a strong bias for codons ending in G and C.

Amino Acid Sequence↗

The secondary structure of mRNAs from Escherichia coli: its possible role in increasing the accuracy of translation.

A secondary structure model was proposed for mRNAs during translation (in a polysome) where the secondary structure is described by a set of small unbranched hairpins. Computer simulation experiments reveal that the number of hairpins is much greater (P less than 10(-6) in highly expressed mRNAs from E. coli as compared with the random sequences coding for the same amino acid sequence, i.e. certain synonymous codons are used in definite mRNA positions to increase the number of hairpins. No constraints on the amino acid sequence, which would affect the secondary structure of mRNAs, were found. The codons UGU, UGC (Cys), GCC (Ala), ACA, ACG (Thr), CCU, CCC (Pro), etc. translated by minor tRNAs were found to occur significantly more frequently in the position 5' to the hairpins than the other codons translated by major tRNAs (P less than 5.10(-6). This correlation leads to the hypothesis that the process of hairpin unfolding can increase the time of translocation from the A to P ribosome site of the codon 5' to the hairpin, thus decreasing the probability of translational error (the latter would likely occur more frequently in the codons translated by minor tRNAs).

Base Sequence↗

Levels of tRNAs in bacterial cells as affected by amino acid usage in proteins.

Transfer RNAs of Mycoplasma capricolum were separated by two-dimensional polyacrylamide gel electrophoresis, and the relative abundance of each of the 28 known tRNA species was measured. There existed a correlation between the relative amount of isoacceptor tRNAs and the frequency in choosing synonymous codons that could be translated by the isoacceptors. Furthermore, it was observed that the total amount of tRNAs for a particular amino acid was paralleled by the composition of the amino acid in ribosomal proteins. A similar relationship was obtained from reexamination of the previous data on Escherichia coli tRNAs, suggesting that the amount of tRNAs for an amino acid is affected by the usage of the amino acid in proteins.

Amino Acid Sequence↗

Regional base composition variation along yeast chromosome III: evolution of chromosome primary structure.

The recent determination of the complete sequence of chromosome III from the yeast Saccharomyces cerevisiae allows, for the first time, the investigation of the long range primary structure of a eukaryotic chromosome. We have found that, against a background G+C level of about 35%, there are two regions (one in each chromosome arm) in which G+C values rise to over 50%. This effect is seen in silent sites within genes, but not in noncoding intergenic sequences. The variation in G+C content is not related to differential selection of synonymous codons, and probably reflects mutational biases. That the intergenic regions do not exhibit the same phenomenon is particularly interesting, and suggests that they are under substantial constraint. The yeast chromosome may be a model of the structure of the human genome, since there is evidence that it is also a mosaic of long regions of different base compositions, reflected in wide variation of G+C content at silent sites among genes. Two possible causes of this regional effect, replication timing, and recombination frequency, are discussed.

Animals↗

An Integrated Sequence-Structure Database incorporating matching mRNA sequence, amino acid sequence and protein three-dimensional structure data.

We have constructed a non-homologous database, termed the Integrated Sequence-Structure Database (ISSD) which comprises the coding sequences of genes, amino acid sequences of the corresponding proteins, their secondary structure and straight phi,psi angles assignments, and polypeptide backbone coordinates. Each protein entry in the database holds the alignment of nucleotide sequence, amino acid sequence and the PDB three-dimensional structure data. The nucleotide and amino acid sequences for each entry are selected on the basis of exact matches of the source organism and cell environment. The current version 1.0 of ISSD is available on the WWW at http://www.protein.bio.msu.su/issd/ and includes 107 non-homologous mammalian proteins, of which 80 are human proteins. The database has been used by us for the analysis of synonymous codon usage patterns in mRNA sequences showing their correlation with the three-dimensional structure features in the encoded proteins. Possible ISSD applications include optimisation of protein expression, improvement of the protein structure prediction accuracy, and analysis of evolutionary aspects of the nucleotide sequence-protein structure relationship.

Algorithms↗

ISSD Version 2.0: taxonomic range extended.

Two more organisms from different taxonomic groups were added to a new version of the Integrated Sequence-Structure Database (ISSD). ISSD serves as an integrated source of sequence and structure information for the analysis of correlations between mRNA synonymous codon usage and three-dimensional structure of the encoded proteins. ISSD now holds 88 non-homologous Escherichia coli proteins and 25 yeast Saccharomyces cerevisiae proteins in addition to the expanded set of mammalian proteins, which includes 166 proteins (107 in ISSD Version 1.0). Comparison of ISSD sequences with organism-specific codon usage data derived from CUTG database shows that it is a representative subset of the GenBank coding sequences data. Preliminary results of the statistical analysis confirm that sequence-structure correlations observed by us earlier are also present in the upgraded ISSD (Version 2.0), including bacterial and yeast proteins. The ISSD Version 2.0 release includes an improved Web-based data search and retrieval system and is accessible via URL http://www.protein.bio.msu.su/issd/. ISSD can be also accessed at ExPASy, URL http://www.expasy.ch/swissmod/swiss-model.htm l

Animals↗

Secondary structure of MS2 phage RNA and bias in code word usage.

Based on the secondary structural model of MS2 RNA, it is shown that, in base-pairing regions of the RNA, there is a bias in the use of synonymous codons which favours C and/or G over U and/or A in the third codon positions, and that in non-pairing regions, there is an opposite bias which favours U and/or A over C and/or G. This nature is interpreted as a result of selective constraint which stabilises the secondary structure of the single-stranded RNA genome of the MS2 phage.

Base Sequence↗

Orotate phosphoribosyltransferase from Thermus thermophilus: overexpression in Escherichia coli, purification and characterization.

Orotate phosphoribosyltransferase (OPRTase, EC2.4.2.10) plays a role in de novo synthesis of pyrimidine nucleotide and transfers orotate to 5-phosphoribosyl-1-pyrophosphate (PRPP) to form orotidine-5'-monophosphate (OMP). To obtain heat-stable OPRTase and to elucidate the mechanism of heat stability, this enzyme from Thermus thermophilus was expressed in Escherichia coli and purified. The pyrE gene of T. thermophilus which encodes OPRTase, contains an open reading frame of 549 base pairs with 69% G+C content. Since this gene expressed itself inefficiently in E. coli, the 5' and 3' ends of the coding regions were replaced with synonymous codons which contain more A+T and corresponds to major codons for E. coli. Introduction of the modified gene fragments into a plasmid having a tac promoter resulted in production of a polypeptide of molecular weight (M(r)) 20,000 in the presence of isopropyl-beta-D-thiogalactopyranoside (IPTG) in E. coli. This protein represented as much as 16% of the bacterial total protein and showed the OPRTase activity. Three purification steps, consisting of heat treatment at 65 degrees C, 40% ammonium sulfate fractionation, and KCl gradient elution from DEAE-Sephadex A-50, resulted in highly purified single polypeptide. The optimum activity of the purified OPRTase was observed at 150 mM KCl, pH 9.0, 75-80 degrees C, and in the presence of 100 microM PRPP. The activation energy of this enzyme reaction was 20.3 kJ/mol. The Km of this enzyme for orotate as a substrate was 75 microM and the maximum specific activity was 300 units/mg protein under the optimum conditions. The purified OPRTase was stable for 20 min at 85 degrees C.

Amino Acid Sequence↗

Elevated rates of nonsynonymous substitution in island birds.

Slightly deleterious mutations are expected to fix at relatively higher rates in small populations than in large populations. Support for this prediction of the nearly-neutral theory of molecular evolution comes from many cases in which lineages inferred to differ in long-term average population size have different rates of nonsynonymous substitution. However, in most of these cases, the lineages differ in many other ways as well, leaving open the possibility that some factor other than population size might have caused the difference in substitution rates. We compared synonymous and nonsynonymous substitutions in the mitochondrial cyt b and ND2 genes of nine closely related island and mainland lineages of ducks and doves. We assumed that island taxa had smaller average population sizes than those of their mainland sister taxa for most of the time since they were established. In all nine cases, more nonsynonymous substitutions occurred on the island branch, but synonymous substitutions showed no significant bias. As in previous comparisons of this kind, the lineages with smaller populations might differ in other respects that tend to increase rates of nonsynonymous substitution, but here such differences are expected to be slight owing to the relatively recent origins of the island taxa. An examination of changes to apparently "preferred" and "unpreferred" synonymous codons revealed no consistent difference between island and mainland lineages.

Amino Acid Sequence↗

Determinants of substitution rates in mammalian genes: expression pattern affects selection intensity but not mutation rate.

To determine whether gene expression patterns affect mutation rates and/or selection intensity in mammalian genes, we studied the relationships between substitution rates and tissue distribution of gene expression. For this purpose, we analyzed 2,400 human/rodent and 834 mouse/rat orthologous genes, and we measured (using expressed sequence tag data) their expression patterns in 19 tissues from three development states. We show that substitution rates at nonsynonymous sites are strongly negatively correlated with tissue distribution breadth: almost threefold lower in ubiquitous than in tissue-specific genes. Nonsynonymous substitution rates also vary considerably according to the tissues: the average rate is twofold lower in brain-, muscle-, retina- and neuron-specific genes than in lymphocyte-, lung-, and liver-specific genes. Interestingly, 5' and 3' untranslated regions (UTRs) show exactly the same trend. These results demonstrate that the expression pattern is an essential factor in determining the selective pressure on functional sites in both coding and noncoding regions. Conversely, silent substitution rates do not vary with expression pattern, even in ubiquitously expressed genes. This latter result thus suggests that synonymous codon usage is not constrained by selection in mammals. Furthermore, this result also indicates that there is no reduction of mutation rates in genes expressed in the germ line, contrary to what had been hypothesized based on the fact that transcribed DNA is more efficiently repaired than nontranscribed DNA.

Animals↗