PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “synonymous codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Molecular evolution of human and rabbit beta-globin mRNAs.

The primary structures of human and rabbit beta-globin mRNAs are compared. Using as a standard the extent of nucleotide substitutions inferred from the hypervariable amino acid residues of fibrinopeptides A and B, which are thought to change largely by neutral evolution, we show that not all silent mutations in globin mRNA are neutral. The divergence of the sequences is limited in part by the selective usage of synonymous codons. The divergent nucleotides tend to be distributed nonrandomly: in the coding region silent substitutions are most rare in segments that are also deficient in substitutions leading to replacements.

Animals↗

Rapid evolution of sex-related genes in Chlamydomonas.

Biological speciation ultimately results in prezygotic isolation-the inability of incipient species to mate with one another-but little is understood about the selection pressures and genetic changes that generate this outcome. The genus Chlamydomonas comprises numerous species of unicellular green algae, including numerous geographic isolates of the species C. reinhardtii. This diverse collection has allowed us to analyze the evolution of two sex-related genes: the mid gene of C. reinhardtii, which determines whether a gamete is mating-type plus or minus, and the fus1 gene, which dictates a cell surface glycoprotein utilized by C. reinhardtii plus gametes to recognize minus gametes. Low stringency Southern analyses failed to detect any fus1 homologs in other Chlamydomonas species and detected only one mid homolog, documenting that both genes have diverged extensively during the evolution of the lineage. The one mid homolog was found in C. incerta, the species in culture that is most closely related to C. reinhardtii. Its mid gene carries numerous nonsynonymous and synonymous codon changes compared with the C. reinhardtii mid gene. In contrast, very high sequence conservation of both the mid and fus1 sequences is found in natural isolates of C. reinhardtii, indicating that the genes are not free to drift within a species but do diverge dramatically between species. Striking divergence of sex determination and mate recognition genes also has been encountered in a number of other eukaryotic phyla, suggesting that unique, and as yet unidentified, selection pressures act on these classes of genes during the speciation process.

Amino Acid Sequence↗

Problems in protein biosynthesis.

Outline of the steps in protein synthesis. Nature of the genetic code. The use of synthetic oligo- and polynucleotides in deciphering the code. Structure of the code: relatedness of synonym codons. The wobble hypothesis. Chain initiation and N-formyl-methionine. Chain termination and nonsense codons. Mistakes in translation: ambiguity in vitro. Suppressor mutations resulting in ambiguity. Limitations in the universality of the code. Attempts to determine the particular codons used by a species. Mechanisms of suppression, caused by (a) abnormal aminoacyl-tRNA, (b) ribosomal malfunction. Effect of streptomycin. The problem of "reading" a nucleic acid template. Different ribosomal mutants and DNA polymerase mutants might cause different mistakes. The possibility of involvement of allosteric proteins in template reading.

Genetic Code↗

Synonymous genetic polymorphisms within Brazilian human immunodeficiency virus Type 1 subtypes may influence mutational routes to drug resistance.

BACKGROUND: Most published data on antiretroviral-drug resistance is generated from in vitro or in vivo studies of subtype B virus. However, this subtype is associated with <10% of HIV infections worldwide, and it is essential to explore subtype-specific determinants of drug resistance. One potential cause of the differences between subtypes is the synonymous codon usage at key resistance positions. METHODS: We investigated the nucleotide sequences at drug resistance-related sites, for all major Brazilian subtypes (B, C, and F1) of human immunodeficiency virus type 1 (HIV-1) group M. RESULTS: We identified a change at positions 151 and 210 of the reverse-transcriptase region in subtype F1, such that the emergence of these key nucleoside/nucleotide analogue resistance mutations required an extra nucleotide change in subtype F1, compared with subtypes B and C. The clinical significance of position 210 was confirmed within a large Brazilian database, in which we identified a lower prevalence of the L210W mutation in subtype F1 virus, compared with subtype B virus, in patients matched for thymidine-analogue experience. An inverse relationship between the L210W and K70R mutations was also observed. CONCLUSIONS: The findings of the present study illustrate an important mechanism by which a subtype may determine genetic routes to resistance, with implications for treatment strategies for populations infected with HIV-1 subtype F.

Base Sequence↗

Causal analysis of CpG suppression in the Mycoplasma genome.

Some bacterial genomes are known to have low CpG dinucleotide frequencies. While their causes are not clearly understood, the frequency of CpG is suppressed significantly in the genome of Mycoplasma genitalium, but not in that of Mycoplasma pneumoniae. We compared orthologous gene pairs of the two closely related species to analyze CpG substitution patterns between these two genomes. We also divided genome sequences into three regions: protein-coding, noncoding, and RNA-coding, and obtained the CpG frequencies for each region for each organism. It was found that the observed/expected ratio of CpG dinucleotides is low in both the protein-coding and noncoding regions; while that ratio is in the normal range in the RNA-coding region. Our results indicate that CpG suppression of the Mycoplasma genome is not caused by (1) biased usage amino acid; (2) biased usage of synonymous codon; or (3) methylation effects by the CpG methyltransferase in the genomes of their hosts. Instead, we consider it likely that a certain global pressure, such as genome-wide pressure for the advantages of DNA stability or replication, has the effect of decreasing CpG over the entire genome, which, in turn, resulted in the biased codon usage.

Base Composition↗

Large-scale, multi-genome analysis of alternate open reading frames in bacteria and archaea.

Analysis of over 300,000 annotated genes in 105 bacterial and archaeal genomes reveals an unexpectedly high frequency of large (>300 nucleotides) alternate open reading frames (ORFs). Especially notable is the very high frequency of alternate ORFs in frames +3 and -1 (where the annotated gene is defined as frame +1). The occurrence of alternate ORFs is correlated with genomic G+C content and is strongly influenced by synonymous codon usage bias. The frequency of alternate ORFs in frame -1 is also influenced by the occurrence of codons encoding leucine and serine in frame +1. Although some alternate ORFs have been shown to encode proteins, many others are probably not expressed because they lack appropriate signals for transcription and translation. These latter can be mis-annotated by automatic gene finding programs leading to errors in public databases. Especially prone to mis-annotation is frame -1, because it exhibits a potential codon usage and theoretical capacity to encode proteins with an amino acid composition most similar to real genes. Some alternate ORFs are conserved across bacterial or archaeal species, and can give rise to misannotated "conserved hypothetical" genes, while others are unique to a genome and are misidentified as "hypothetical orphan" genes, contributing significantly to the orphan gene paradox.

Algorithms↗

FISH: a guide to protein-coding DNA sequences in the GenBank database.

FISH (Fast Index Search for Homologous coding sequences) consists of a database and associated software and is intended to function as a directory of protein-coding gene sequences. The FISH index contains descriptions of 22,361 DNA sequences from release 69.0 of the GenBank genetic sequence database. Complete coding sequences are represented numerically with counts of nucleotides and synonymous codons, and with GenBank LOCUS names and short descriptions. The software permits the database to be queried by GenBank LOCUS name, sequence length (expressed as total number of codons), or by comparison with a DNA sequence. In the latter case, the numerical descriptions are compared with simple distance measures in place of actual DNA sequences. The FISH package can be used to rapidly assemble lists of similar coding sequences, without regard to functional annotation or sequence alignments. Typical search times are well under a minute on widely available IBM-compatible microcomputers.

Algorithms↗

Silent DNA: speaking RNA language?

The sequence of silent DNA in the human genome (intergenic spacers, introns and synonymous codon positions of protein-coding genes) was found here to have the higher thermostability of corresponding RNA/RNA and RNA/DNA duplexes as compared with randomized sequence. This difference increased with elevation of GC content. The revealed effect was not due to correlation of RNA/RNA and RNA/DNA thermostabilities with thermostability of the DNA/DNA duplex, which, on the contrary, was lower than in the randomized sequence and lagged behind the elevation of GC content. The same picture was observed in the genomes of other warm-blooded vertebrates but not in the lower organisms. This finding suggests that RNA-RNA and RNA-DNA interactions could be involved in the putative function of silent DNA.

Animals↗

Transient mutators: a semiquantitative analysis of the influence of translation and transcription errors on mutation rates.

A population of bacteria growing in a nonlimiting medium includes mutator bacteria and transient mutators defined as wild-type bacteria which, due to occasional transcription or translation errors, display a mutator phenotype. A semiquantitative theoretical analysis of the steady-state composition of an Escherichia coli population suggests that true strong genotypic mutators produce about 3 x 10(-3) of the single mutations arising in the population, while transient mutators produce at least 10% of the single mutations and more than 95% of the simultaneous double mutations. Numbers of mismatch repair proteins inherited by the offspring, proportions of lethal mutations and mortality rates are among the main parameters that influence the steady-state composition of the population. These results have implications for the experimental manipulation of mutation rates and the evolutionary fixation of frequent but nearly neutral mutations (e.g., synonymous codon substitutions).

Bacteria↗

Inferring the fitness effects of DNA mutations from polymorphism and divergence data: statistical power to detect directional selection under stationarity and free recombination.

The fitness effects of classes of DNA mutations can be inferred from patterns of nucleotide variation. A number of studies have attributed differences in levels of polymorphism and divergence between silent and replacement mutations to the action of natural selection. Here, I investigate the statistical power to detect directional selection through contrasts of DNA variation among functional categories of mutations. A variety of statistical approaches are applied to DNA data simulated under Sawyer and Hartl's Poisson random field model. Under assumptions of free recombination and stationarity, comparisons that include both the frequency distributions of mutations segregating within populations and the numbers of mutations fixed between populations have substantial power to detect even very weak selection. Frequency distribution and divergence tests are applied to silent and replacement mutations among five alleles of each of eight Drosophila simulans genes. Putatively "preferred" silent mutations segregate at higher frequencies and are more often fixed between species than "unpreferred" silent changes, suggesting fitness differences among synonymous codons. Amino acid changes tend to be either rare polymorphisms or fixed differences, consistent with a combination of deleterious and adaptive protein evolution. In these data, a substantial fraction of both silent and replacement DNA mutations appear to affect fitness.

Adaptation, Biological↗

Ancient allelism at the cytosolic chaperonin-alpha-encoding gene of the zebrafish.

The T-complex protein 1, TCP1, gene codes for the CCT-alpha subunit of the group II chaperonins. The gene was first described in the house mouse, in which it is closely linked to the T locus at a distance of approximately 11 cM from the Mhc. In the zebrafish, Danio rerio, in which the T homolog is linked to the class I Mhc loci, the TCP1 locus segregates independently of both the T and the Mhc loci. Despite its conservation between species, the zebrafish TCP1 locus is highly polymorphic. In a sample of 15 individuals and the screening of a cDNA library, 12 different alleles were found, and some of the allelic pairs were found to differ by up to nine nucleotides in a 275-bp-long stretch of sequence. The substitutions occur in both translated and untranslated regions, but in the former they occur predominantly at synonymous codon sites. Phylogenetically, the alleles fall into two groups distinguished also by the presence or absence of a 10-bp insertion/deletion in the 3' untranslated region. The two groups may have diverged as long as 3.5 mya, and the polymorphic differences may have accumulated by genetic drift in geographically isolated populations.

Alleles↗

Translational selection and yeast proteome evolution.

The primary structures of peptides may be adapted for efficient synthesis as well as proper function. Here, the Saccharomyces cerevisiae genome sequence, DNA microarray expression data, tRNA gene numbers, and functional categorizations of proteins are employed to determine whether the amino acid composition of peptides reflects natural selection to optimize the speed and accuracy of translation. Strong relationships between synonymous codon usage bias and estimates of transcript abundance suggest that DNA array data serve as adequate predictors of translation rates. Amino acid usage also shows striking relationships with expression levels. Stronger correlations between tRNA concentrations and amino acid abundances among highly expressed proteins than among less abundant proteins support adaptation of both tRNA abundances and amino acid usage to enhance the speed and accuracy of protein synthesis. Natural selection for efficient synthesis appears to also favor shorter proteins as a function of their expression levels. Comparisons restricted to proteins within functional classes are employed to control for differences in amino acid composition and protein size that reflect differences in the functional requirements of proteins expressed at different levels.

Adaptation, Physiological↗

Gene "volatility" is most unlikely to reveal adaptation.

It has recently been claimed that adaptive molecular evolution can be detected within single genome sequences by use of gene "volatility" scores. However, the approach used was entirely based on the assumption that synonymous codon usage is normally shaped by selection for low volatility; this is most unlikely to be true. Furthermore, even if that assumption could be justified, the method would clearly lack power, detecting only genes where a very large number of nonsynonymous substitutions had occurred. Volatility scores are susceptible to other influences. The unusually high volatilities of the Mycobacterium tuberculosis and Plasmodium falciparum genes that were identified as putatively having undergone adaptive changes were largely the result of internally repetitive structures, in which unusual codon usage was caused by the mechanisms that generated this repetition rather than by adaptive changes.

Adaptation, Physiological↗

A dissection of volatility in yeast.

It has been suggested that volatility, the proportion of mutations which change an amino acid, can be used to infer the level of natural selection acting upon a gene. This conjecture is supported by a correlation between volatility and the rate of nonsynonymous substitution (dN), or the ratio of nonsynonymous and synonymous substitution rates, in a variety of organisms. These organisms include yeast, in which the correlations are quite strong. Here we show that these correlations are a by-product of a correlation between synonymous codon bias toward translationally optimal codons and dN. Although this analysis suggests that volatility is not a good measure of the selection, we suggest that it might be possible to infer something about the level of natural selection, from a single genome sequence, using translational codon bias.

Base Sequence↗

The rate of adaptive evolution in enteric bacteria.

Here we estimate the rate of adaptive substitution in a set of 410 genes that are present in 6 Escherichia coli and 6 Salmonella enterica genomes. We estimate that more than 50% of amino acid substitutions in this set of genes have been fixed by positive selection between the E. coli and S. enterica lineages. We also show that the proportion of adaptive substitutions is uncorrelated with the rate of amino acid substitution or gene function but that it may be correlated with levels of synonymous codon usage bias.

Adaptation, Biological↗

The nucleotide sequence of myosin light chain (L-2A) mRNA from embryonic chicken cardiac muscle tissue.

The nucleotide sequence of a cDNA clone (pML10) for chicken cardiac myosin light chain is described. The cDNA insert contains 613 nucleotides representing the entire coding sequence, with the exception of nine NH2-terminal amino acids, and the full 3'-non-coding region of 146 nucleotides. The missing 5' terminus of the mRNA, not represented in the clone pML10, was obtained by extension of the cDNA using a 43 nucleotide long internal EcoR1 fragment as a primer. The non-coding region contains several direct and inverted repeated sequences and the polyadenylation signal sequence AATAAA. The coding portion exhibits non-random usage of synonymous codons with a strong bias for codons ending in G and C.

Amino Acid Sequence↗

The secondary structure of mRNAs from Escherichia coli: its possible role in increasing the accuracy of translation.

A secondary structure model was proposed for mRNAs during translation (in a polysome) where the secondary structure is described by a set of small unbranched hairpins. Computer simulation experiments reveal that the number of hairpins is much greater (P less than 10(-6) in highly expressed mRNAs from E. coli as compared with the random sequences coding for the same amino acid sequence, i.e. certain synonymous codons are used in definite mRNA positions to increase the number of hairpins. No constraints on the amino acid sequence, which would affect the secondary structure of mRNAs, were found. The codons UGU, UGC (Cys), GCC (Ala), ACA, ACG (Thr), CCU, CCC (Pro), etc. translated by minor tRNAs were found to occur significantly more frequently in the position 5' to the hairpins than the other codons translated by major tRNAs (P less than 5.10(-6). This correlation leads to the hypothesis that the process of hairpin unfolding can increase the time of translocation from the A to P ribosome site of the codon 5' to the hairpin, thus decreasing the probability of translational error (the latter would likely occur more frequently in the codons translated by minor tRNAs).

Base Sequence↗

Levels of tRNAs in bacterial cells as affected by amino acid usage in proteins.

Transfer RNAs of Mycoplasma capricolum were separated by two-dimensional polyacrylamide gel electrophoresis, and the relative abundance of each of the 28 known tRNA species was measured. There existed a correlation between the relative amount of isoacceptor tRNAs and the frequency in choosing synonymous codons that could be translated by the isoacceptors. Furthermore, it was observed that the total amount of tRNAs for a particular amino acid was paralleled by the composition of the amino acid in ribosomal proteins. A similar relationship was obtained from reexamination of the previous data on Escherichia coli tRNAs, suggesting that the amount of tRNAs for an amino acid is affected by the usage of the amino acid in proteins.

Amino Acid Sequence↗