PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “nucleotide composition bias”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Evolutionary basis of codon usage and nucleotide composition bias in vertebrate DNA viruses.

Understanding the extent and causes of biases in codon usage and nucleotide composition is essential to the study of viral evolution, particularly the interplay between viruses and host cells or immune responses. To understand the common features and differences among viruses we analyzed the genomic characteristics of a representative collection of all sequenced vertebrate-infecting DNA viruses. This revealed that patterns of codon usage bias are strongly correlated with overall genomic GC content, suggesting that genome-wide mutational pressure, rather than natural selection for specific coding triplets, is the main determinant of codon usage. Further, we observed a striking difference in CpG content between DNA viruses with large and small genomes. While the majority of large genome viruses show the expected frequency of CpG, most small genome viruses had CpG contents far below expected values. The exceptions to this generalization, the large gammaherpesviruses and iridoviruses and the small dependoviruses, have sufficiently different life-cycle characteristics that they may help reveal some of the factors shaping the evolution of CpG usage in viruses.

Animals↗

Nucleotide composition bias and CpG dinucleotide content in the genomes of HIV and HTLV 1/2.

Nucleotide compositions of the HIV subfamily and HTLV 1/2 genomes are strongly biased in a remarkably opposite way; HIV is adenine-rich and cytosine-poor while HTLV 1/2 is cytosine-rich and adenine-poor. In addition, the CpG dinucleotides are underrepresented in HIV but abundant in HTLV 1/2. By these two properties the genomes of HIV and HTLV 1/2 mimic an (A + T)-rich and (G + C)-rich segment of the host genome, respectively. These dramatic differences between the two human retroviruses might have evolved to direct integration of the retroviral genomes into specific segments of the human chromosomes.

Adenine↗

Effects of nucleotide composition bias on the success of the parsimony criterion in phylogenetic inference.

Convergence in nucleotide composition (CNC) in unrelated lineages is a factor potentially affecting the performance of most phylogeny reconstruction methods. Such convergence has deleterious effects because unrelated lineages show similarities due to similar nucleotide compositions and not shared histories. While some methods (such as the LogDet/paralinear distance measure) avoid this pitfall, the amount of convergence in nucleotide composition necessary to deceive other phylogenetic methods has never been quantified. We examined analytically the relationship between convergence in nucleotide composition and the consistency of parsimony as a phylogenetic estimator for four taxa. Our results show that rather extreme amounts of convergence are necessary before parsimony begins to prefer the incorrect tree. Ancillary observations are that (for unweighted Fitch parsimony) transition/transversion bias contributes to the impact of CNC and, for a given amount of CNC and fixed branch lengths, data sets exhibiting substantial site-to-site rate heterogeneity present fewer difficulties than data sets in which rates are homogeneous. We conclude by reexamining a data set originally used to illustrate the problems caused by CNC. Using simulations, we show that in this case the convergence in nucleotide composition alone is insufficient to cause any commonly used methods to fail, and accounting for other evolutionary factors (such as site-to-site rate heterogeneity) can give a correct inference without accounting for CNC.

Animals↗

Exploitation of the low fidelity of human immunodeficiency virus type 1 (HIV-1) reverse transcriptase and the nucleotide composition bias in the HIV-1 genome to alter the drug resistance development of HIV.

The RNA genome of the lentivirus human immunodeficiency virus type 1 (HIV-1) is significantly richer in adenine nucleotides than the statistically equal distribution of the four different nucleotides that is expected. This compositional bias may be due to the guanine-to-adenine (G-->A) nucleotide hypermutation of the HIV genome, which has been explained by dCTP pool imbalances during reverse transcription. The adenine nucleotide bias together with the poor fidelity of HIV-1 reverse transcriptase markedly enhances the genetic variation of HIV and may be responsible for the rapid emergence of drug-resistant HIV-1 strains. We have now attempted to counteract the normal mutational pattern of HIV-1 in response to anti-HIV-1 drugs by altering the endogenous deoxynucleoside triphosphate pool ratios with antimetabolites in virus-infected cell cultures. We showed that administration of these antimetabolic compounds resulted in an altered drug resistance pattern due to the reversal of the predominant mutational flow of HIV (G-->A) to an adenine-to-guanine (A-->G) nucleotide pattern in the intact HIV-1-infected lymphocyte cultures. Forcing the virus to change its inherent nucleotide bias may lead to better control of viral drug resistance development.

Anti-HIV Agents↗

Biased nucleotide composition of the genome of HERV-K related endogenous retroviruses and its evolutionary implications.

The human genome contains a large number of sequences that belong to the HERV-K family of human endogenous retroviruses. Most of these elements are likely remnants of ancient infections by ancestral exogenous retroviruses. To obtain further insight into the evolutionary history and molecular mechanisms responsible for the diversity of the human HERV-K elements, we analyzed several aspects of their genome structure. The nucleotide composition of the HERV-K genome was found to be highly biased and asymmetric, with an abundance of the A nucleotide in the viral (+) strand. A similar trend has been reported for the genomes of several exogenous retroviruses, with different nucleotides as the preferred building block. Other genome characteristics that were reported previously for actively replicating retroviruses are also apparent for the endogenous HERV-K virus. In particular, we observed suppression of the dinucleotide CpG, which represents potential methylation sites, and a strong preference for synonymous substitutions within the open reading frame of the reverse transcriptase (RT) enzyme. Furthermore, the mutational spectrum of the HERV-K RT enzyme was evaluated by nucleotide sequence comparison of 34 available elements. Interestingly, this analysis revealed a striking similarity with the mutational pattern of the HIV-1 RT enzyme, with a preference for G-to-A and C-to-T transitions. It is proposed that the mutational bias of the HERV-K RT enzyme played a role in the shaping of this retroviral genome, which was actively replicating more than 30 million years ago. This effect can still be observed in the contemporary endogenous HERV-K elements.

Base Composition↗

Shared nucleotide composition biases among species and their impact on phylogenetic reconstructions of the Drosophilidae.

Compositional changes are a major feature of genome evolution. Overlooking nucleotide composition differences among sequences can seriously mislead phylogenetic reconstructions. Large compositional variation exists among the members of the family Drosophilidae. Until now, however, base composition differences have been largely neglected in the formulations of the nucleotide substitution process used to reconstruct the phylogeny of this important group of species. The present study adopts a maximum-likelihood framework of phylogenetic inference in order to analyze five nuclear gene regions and shows that (1) the pattern of compositional variation in the Drosophilidae does not match the phylogeny of the species; (2) accounting for the heterogeneous GC content with Galtier and Gouy's nucleotide substitution model leads to a tree that differs in significant aspects from the tree inferred when the nucleotide composition differences are ignored, even though both phylogenetic hypotheses attain strong nodal support in the bootstrap analyses; and (3) the LogDet distance correction cannot completely overcome the distorting effects of the compositional variation that exists among the species of the Drosophilidae. Our analyses confidently place the Chymomyza genus as an outgroup closer than the genus Scaptodrosophila to the Drosophila genus and conclusively support the monophyly of the Sophophora subgenus.

Alcohol Dehydrogenase↗

Retroviral oligonucleotide distributions correlate with biased nucleotide compositions of retrovirus sequences, suggesting a duplicative stepwise molecular evolution.

A computer-assisted analysis was made of 24 complete nucleotide sequences selected from the vertebrate retroviruses to represent the ten viral groups. The conclusions of this analysis extend and strengthen the previously made hypothesis on the Moloney murine leukemia virus: The evolution of the nucleotide sequence appears to have occurred mainly through at least three overlapping levels of duplication: (1) The distributions of overrepresented (3-6)-mers are consistent with the universal rule of a trend toward TG/CT excess and with the persistence of a certain degree of symmetry between the two strands of DNA. This suggests one or several original tandemly repeated sequences and some inverted duplications. (2) The existence of two general core consensuses at the level of these (3-6)-mers supports the hypothesis of a common evolutionary origin of vertebrate retroviruses. Consensuses more specific to certain sequences are compatible with phylogenetic trees established independently. The consensuses could correspond to intermediary evolutionary stages. (3) Most of the (3-6)-mers with a significantly higher than average frequency appear to be internally repeated (with monomeric or oligomeric internal iterations) and seem to be at least partly the cause of the bias observed by other researchers at the level of retroviral nucleotide composition. They suggest a third evolutionary stage by slippage-like stepwise local duplications.

Animals↗

Nucleotide composition bias affects amino acid content in proteins coded by animal mitochondria.

We show that in animal mitochondria homologous genes that differ in guanine plus cytosine (G + C) content code for proteins differing in amino acid content in a manner that relates to the G + C content of the codons. DNA sequences were analyzed using square plots, a new method that combines graphical visualization and statistical analysis of compositional differences in both DNA and protein. Square plots divide codons into four groups based on first and second position A + T (adenine plus thymine) and G + C content and indicate differences in amino acid content when comparing sequences that differ in G + C content. When sequences are compared using these plots, the amino acid content is shown to correlate with the nucleotide bias of the genes. This amino acid effect is shown in all protein-coding genes in the mitochondrial genome, including cox I, cox II, and cyt b, mitochondrial genes which are commonly used for phylogenetic studies. Furthermore, nucleotide content differences are shown to affect the content of all amino acids with A + T- and G + C-rich codons. We speculate that phylogenetic analysis of genes so affected may tend erroneously to indicate relatedness (or lack thereof) based only on amino acid content.

Amino Acids↗

Strand-specific nucleotide composition bias in echinoderm and vertebrate mitochondrial genomes.

The gene organization of starfish mitochondrial DNA is identical with that of the sea urchin counterpart except for a reported inversion of an approximately 4.6-kb segment containing two structural genes for NADH dehydrogenase subunits 1 and 2 (ND 1 and ND 2). When the codon usage of each structural gene in starfish, sea urchin, and vertebrate mitochondrial DNAs is examined, it is striking that codons ending in T and G are preferentially used more for heavy strand-encoded genes, including starfish ND 1 and ND 2, than for light strand-encoded genes, including sea urchin ND 1 and ND 2. On the contrary, codons ending in A and C are preferentially used for the light strand-encoded genes rather than for the heavy strand-encoded ones. Moreover, G-U base pairs are more frequently found in the possible secondary structures of heavy strand-encoded tRNAs than in those of light strand-encoded tRNAs. These observations suggest the existence of a certain constraint operating on mitochondrial genomes from various animal phyla, which results in the accumulation of G and T on one strand, and A and C on the other.

Amino Acid Sequence↗

Bias in nucleotide composition of antisense oligonucleotides.

In this study we investigated if specific sequence motifs occur with a higher frequency in antisense oligonucleotides than can be expected on the basis of the mRNA composition to get an impression of the importance of these motifs for antisense effects. Computer analysis of 206 antisense oligonucleotides extracted from the literature and from sequence databases, all targeted against human mRNA, was performed. We compared the sequence composition of these oligonucleotides with the average of 100 equally large and randomly selected sequences from sequence databases and of their target mRNA. We found that the frequency of sequence motifs containing GG, CCC, CC, GAC, and CG is significantly higher and TT and that TCC is significantly lower in antisense oligonucleotide sequences than in the randomly selected mRNA sequences. We conclude that there is a bias in the nucleotide composition of antisense oligonucleotides. Some of these biased sequence motifs have been reported to induce nonantisense effects mediated by protein binding. Further analysis of the biologic function of these motifs is necessary to investigate if they should be avoided or incorporated into future designs of therapeutic effective oligonucleotides.

Base Composition↗

Compositional bias may affect both DNA-based and protein-based phylogenetic reconstructions.

It is now well-established that compositional bias in DNA sequences can adversely affect phylogenetic analysis based on those sequences. Phylogenetic analyses based on protein sequences are generally considered to be more reliable than those derived from the corresponding DNA sequences because it is believed that the use of encoded protein sequences circumvents the problems caused by nucleotide compositional biases in the DNA sequences. There exists, however, a correlation between AT/GC bias at the nucleotide level and content of AT- and GC-rich codons and their corresponding amino acids. Consequently, protein sequences can also be affected secondarily by nucleotide compositional bias. Here, we report that DNA bias not only may affect phylogenetic analysis based on DNA sequences, but also drives a protein bias which may affect analyses based on protein sequences. We present a striking example where common phylogenetic tools fail to recover the correct tree from complete animal mitochondrial protein-coding sequences. The data set is very extensive, containing several thousand sites per sequence, and the incorrect phylogenetic trees are statistically very well supported. Additionally, neither the use of the LogDet/paralinear transform nor removal of positions in the protein alignment with AT- or GC-rich codons allowed recovery of the correct tree. Two taxa with a large compositional bias continually group together in these analyses, despite a lack of close biological relatedness. We conclude that even protein-based phylogenetic trees may be misleading, and we advise caution in phylogenetic reconstruction using protein sequences, especially those that are compositionally biased.

Animals↗

Codon and amino acid usage in retroviral genomes is consistent with virus-specific nucleotide pressure.

Retroviral RNA genomes are known to have a biased nucleotide composition. For instance, the plus-strand RNA of human immunodeficiency virus (HIV) is A-rich, and the genome of human T cell leukemia virus (HTLV) is C-rich, and other retroviruses have a U-rich or G-rich genome. The biased composition of these genomes is most likely caused by directional mutational pressure of the respective reverse transcriptase enzymes. Using a set of retroviral genomes with a distinct nucleotide composition, we performed skew analyses of the nucleotide bias along the complete viral genome. Distinct nucleotide signatures were apparent, and these typical patterns were generally conserved across the viral genome. Furthermore, it is demonstrated that this typical nucleotide bias, combined with a profound discrimination against the CpG dinucleotide sequence, strongly influences the codon usage of the retroviruses in a direct manner, and their amino acid usage in an indirect manner. The fact that both codon usage and amino acid usage are so closely entwined with the genome composition has important practical implications. For instance, the typical trends in nucleotide usage could influence the molecular phylogenetic reconstruction of the family Retroviridae.

Amino Acids↗

Discovering patterns in Plasmodium falciparum genomic DNA.

A method has been developed for discovering patterns in DNA sequences. Loosely based on the well-known Lempel Ziv model for text compression, the model detects repeated sequences in DNA. The repeats can be forward or inverted, and they need not be exact. The method is particularly useful for detecting distantly related sequences, and for finding patterns in sequences of biased nucleotide composition, where spurious patterns are often observed because the bias leads to coincidental nucleotide matches. We show here the utility of the method by applying it to genomic sequences of Plasmodium falciparum. A single scan of chromosomes 2 and 3 of P. falciparum, using our method and no other a priori information about the sequences, reveals regions of low complexity in both telomeric and central regions, long repeats in the subtelomeric regions, and shorter repeat areas in dense coding regions. Application of the method to a recently sequenced contig of chromosome 10 that has a particularly biased base composition detects a long internal repeat more readily than does the conventional dot matrix plot. Space requirements are linear, so the method can be used on large sequences. The observed repeat patterns may be related to large-scale chromosomal organization and control of gene expression. The method has general application in detecting patterns of potential interest in newly sequenced genomic material.

Algorithms↗

Selection for highly biased amino acid frequency in the TolA cell envelope protein of Proteobacteria.

The bacterial cell envelope protein TolA functions to maintain the integrity of the cell membrane. This protein contains high levels of alanine and lysine that are used in the formation of alpha helices, which are required for normal protein function. The neutral model of molecular evolution predicts that amino acid composition and nucleotide composition are driven by the underlying GC content, as a result of mutation bias. However, this study shows that selection has acted to maintain high levels of alanine and lysine in the TolA protein of Proteobacteria, which in turn has biased nucleotide composition in the corresponding tolA gene.

Amino Acids↗

The construction of amino acid substitution matrices for the comparison of proteins with non-standard compositions.

MOTIVATION: Amino acid substitution matrices play a central role in protein alignment methods. Standard log-odds matrices, such as those of the PAM and BLOSUM series, are constructed from large sets of protein alignments having implicit background amino acid frequencies. However, these matrices frequently are used to compare proteins with markedly different amino acid compositions, such as transmembrane proteins or proteins from organisms with strongly biased nucleotide compositions. It has been argued elsewhere that standard matrices are not ideal for such comparisons and, furthermore, a rationale has been presented for transforming a standard matrix for use in a non-standard compositional context. RESULTS: This paper presents the mathematical details underlying the compositional adjustment of amino acid or DNA substitution matrices.

Algorithms↗

Codon usage patterns in cytochrome oxidase I across multiple insect orders.

Synonymous codon usage bias is determined by a combination of mutational biases, selection at the level of translation, and genetic drift. In a study of mtDNA in insects, we analyzed patterns of codon usage across a phylogeny of 88 insect species spanning 12 orders. We employed a likelihood-based method for estimating levels of codon bias and determining major codon preference that removes the possible effects of genome nucleotide composition bias. Three questions are addressed: (1) How variable are codon bias levels across the phylogeny? (2) How variable are major codon preferences? and (3) Are there phylogenetic constraints on codon bias or preference? There is high variation in the level of codon bias values among the 88 taxa, but few readily apparent phylogenetic patterns. Bias level shifts within the lepidopteran genus Papilio are most likely a result of population size effects. Shifts in major codon preference occur across the tree in all of the amino acids in which there was bias of some level. The vast majority of changes involves double-preference models, however, and shifts between single preferred codons within orders occur only 11 times. These shifts among codons in double-preference models are phylogenetically conservative.

Animals↗

Patterns of nucleotide substitution in mitochondrial protein coding genes of vertebrates.

Maximum likelihood methods were used to study the differences in substitution rates among the four nucleotides and among different nucleotide sites in mitochondrial protein-coding genes of vertebrates. In the 1st + 2nd codon position data, the frequency of nucleotide G is negatively correlated with evolutionary rates of genes, substitution rates vary substantially among sites, and the transition/transversion rate bias (R) is two to five times larger than that expected at random. Generally, largest transition biases and greatest differences in substitution rates among sites are found in the highly conserved genes. The 3rd positions in placental mammal genes exhibit strong nucleotide composition biases and the transitional rates exceed transversional rates by one to two orders of magnitude. Tamura-Nei and Hasegawa-Kishino-Yano models with gamma distributed variable rates among sites (gamma parameter, alpha) adequately describe the nucleotide substitution process in 1st+2nd position data. In these data, ignoring differences in substitution rates among sites leads to largest biases while estimating substitution rates. Kimura's two-parameter model with variable-rates among sites performs satisfactorily in likelihood estimation of R, alpha, and overall amount of evolution for 1st+2nd position data. It can also be used to estimate pairwise distances with appropriate values of alpha for a majority of genes.

Amino Acid Sequence↗

Mitochondrial gene rearrangements and partial genome duplications detected by multigene asymmetric compositional bias analysis.

Asymmetric compositional and mutation bias between the two strands occurs in mitochondrial genomes, and an asymmetric mechanism of mtDNA replication is a potential source of this bias. Some evidence indicates that during replication the heavy strand is subject to a gradient of time spent in a single-stranded state (D (ssH)) and a gradient of mutational damage. The nucleotide composition bias among genes varies with D (ssH). Consequently, partial genome duplications (PGD) will alter the skew for genes located downstream of the duplication, relatively to nascent light strand synthesis, and in the same way, gene rearrangements (GRr) will affect genes by changing their skews. We examined cases where there had been PGD or GRr and determined whether this left a trace in the form of unusual patterns of base composition. We compared the skew of genes differently located on the mtDNA genome of previously published whole mtDNA genomes from amphibians, a group that shows considerable levels of both GRr and PGD. After observing a significant correlation between AT and GC skew with D (ssH) at fourfold redundant sites, we ran our analysis and detected 31.3% of the species with GRr and/or PGD. By comparing the nucleotide composition at fourfold redundant sites in normal and "abnormal" species, we found that A/C variation occurs and is associated with GRr/PGD. These results show that by analyzing the nucleotide skews of only three genes, it may be possible to predict some mitochondrial GRr and/or PGD without knowing the complete mtDNA genome sequence.

Amphibians↗