PubMed Health⌕ Search

PubMed · 16679346

Accounting for background nucleotide composition when measuring codon usage bias: brilliant idea, difficult in practice.

Abstract

The effective number of codons used in a gene is a commonly used measure of codon usage. It varies between 20 and 61 (standard genetic code) and indicates to which degree the entire genetic code is used. It is a drawback of this method that it does not take background composition into account. This led Novembre to introduce a variant called Nc' (Novembre JA. 2002. Accounting for background nucleotide composition when measuring codon usage bias. Mol Biol Evol 19:1390-4). In this letter, its properties are under the loupe, with special emphasis on phenomena relating to codon homozygosity. A theoretical misunderstanding regarding this estimator is explained in detail, notably Nc varies between 0 and 61 instead of 20 and 61 (with the standard genetic code). Practical examples from the genome of Pseudomonas aeruginosa are given which demonstrate that the problem is not just theoretical.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Anders Fuglsang. 2006-05-05. Accounting for background nucleotide composition when measuring codon usage bias: brilliant idea, difficult in practice.. https://doi.org/10.1093/molbev%2Fmsl009

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Application of CE for determination of DNA base composition.

DNA base composition expressed as mol% of guanine plus cytosine (% GC) or GC content is a key parameter of bacterial taxonomy and genomic analyses. Direct chemical determination methods such as HPLC as well as indirect methods based on physical properties of deoxyribonucleic acid (DNA), melting point (T(m)), and buoyant density (B(d)) have been conventionally applied to determine the GC content. However, these methods require relatively large amounts of sample DNA, time, and labor. We have developed a protocol to determine the GC content by fine separation of nucleosides with CZE. Genomic DNAs with known GC content from 23 bacterial strains were determined by CE at the optimized conditions of 27 degrees C, 20 kV in 50 mM of NaHCO(3) (pH 9.0) and 70 mM SDS added. Nucleosides from <1 microg of DNA hydrolyzed with nuclease-P1 and bacterial alkaline phosphatase were separated in a 75 microm wide and 80 cm long silica capillary. The nucleoside peak areas were determined at 254 nm in less than 12 min. The CE-based determination of GC content requires only small amounts of DNA, and thus should be applicable to environmental genomics (metagenomics), as >90% of environmental micro-organisms are nonculturable and produce only small amounts of genomic DNA.

Base Composition↗

Characterizing viral microRNAs and its application on identifying new microRNAs in viruses.

MicroRNAs (miRNAs) are a newly identified class of non-protein-coding small RNAs, which play important roles in multiple biological and metabolic processes at the post-transcriptional level by directly cleaving targeted mRNAs or inhibiting translation. The lengths of viral miRNA precursors vary from 60 to 119 with an average of 79 nucleotides, which was smaller than observed for plant or animal miRNAs. Viral miRNAs are less conserved than plant and animal miRNAs, suggesting that viral miRNAs may evolve rapidly. Uracil nucleotide was highly dominant in the first position of 5' mature miRNAs. Viral miRNAs had high minimal folding free energy index (MFEI, 0.9 +/- 0.1). Based on these features and the well-known characteristics of miRNAs, 20 new potential miRNAs were identified in viruses by using expressed sequence tag (EST) analysis and genomic sequence survey (GSS) analysis. A better understanding of viral miRNA functions will be useful to design new approaches for treating viruses, especially those viruses that can induce human, animal, and plant diseases.

Base Composition↗