PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “synonymous codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

DNA and the neutral theory.

The neutral theory claims that the great majority of evolutionary changes at the molecular (DNA) level are caused not by Darwinian selection but by random fixation of selectively neutral or nearly neutral mutants. The theory also asserts that the majority of protein and DNA polymorphisms are selectively neutral and that they are maintained in the species by mutational input balanced by random extinction. In conjunction with diffusion models (the stochastic theory) of gene frequencies in finite populations, it treats these phenomena in quantitative terms based on actual observations. Although the theory has been strongly criticized by the 'selectionists', supporting evidence has accumulated over the years. Particularly, the recent outburst of DNA sequence data lends strong support to the theory both with respect to evolutionary base substitutions and DNA polymorphism, including rapid evolutionary base substitutions in pseudogenes. In addition, the observed pattern of synonymous codon choice can now be readily explained in the framework of this theory. I review these recent findings in the light of the neutral theory.

Animals↗

DNA sequence evolution: the sounds of silence.

Silent sites (positions that can undergo synonymous substitutions) in protein-coding genes can illuminate two evolutionary processes. First, despite being silent, they may be subject to natural selection. Among eukaryotes this is exemplified by yeast, where synonymous codon usage patterns are shaped by selection for particular codons that are more efficiently and/or accurately translated by the most abundant tRNAs; codon usage across the genome, and the abundance of different tRNA species, are highly co-adapted. Second, in the absence of selection, silent sites reveal underlying mutational patterns. Codon usage varies enormously among human genes, and yet silent sites do not appear to be influenced by natural selection, suggesting that mutation patterns vary among regions of the genome. At first, the yeast and human genomes were thought to reflect a dichotomy between unicellular and multicellular organisms. However, it now appears that natural selection shapes codon usage in some multicellular species (e.g. Drosophila and Caenorhabditis), and that regional variations in mutation biases occur in yeast. Silent sites (in serine codons) also provide evidence for mutational events changing adjacent nucleotides simultaneously.

Animals↗

Evolutionary relationships within a subgroup of HERV-K-related human endogenous retroviruses.

The prototype endogenous retrovirus HERV-K10 was identified in the human genome by its homology to the exogenous mouse mammary tumour virus. By analysis of a short 244 bp segment of the reverse transcriptase (RT) gene of other HERV-K10-like sequences, it has become clear that these elements represent an extended family consisting of multiple groups (the HML-1 to HML-6 subgroups). Some of these elements are transcriptionally active and contain an intact open reading frame for the RT protein, raising the possibility that this family is still expanding through retrotransposition. To better define the relationship of these endogenous retroviruses, we identified ten new members of the HML-2 subgroup. PCR was used to amplify reverse-transcribed RNA of a 595 bp region of the RT gene in a variety of human cell samples, including normal and leukaemic bone marrow and peripheral blood, placenta cells and a transformed T cell line. We provide an extensive phylogenetic analysis of the relationships for this cluster of HERV-K-related endogenous retroviral elements. Nucleotide diversity values for nonsynonymous versus synonymous codon positions indicate that moderately strong selection is or was operating on these retroviral RT gene segments. The evolution of this class of endogenous retroelements is discussed.

Base Sequence↗

Cloning, nucleotide sequence and expression in Streptomyces lividans and Escherichia coli of pabB from Lactococcus lactis subsp. lactis NCDO 496.

A gene (pabB) encoding the aminase activity of p-aminobenzoate (PABA) synthase in Lactococcus lactis subsp. lactis was cloned in pIJ41 and expressed in Streptomyces lividans strains defective in PABA biosynthesis. Expression of the gene was associated with a 1.2 kb deletion between the aph promoter and the cloning site in pIJ41. Subcloning in pBR322 and expression in Escherichia coli AB3295 of the cloned L. lactis DNA fragment localized the pabB-complementing gene in a 1.9 kb segment. The nucleotide sequence of this segment contained a 1410 bp open reading frame encoding a 470-amino-acid polypeptide of 50937 Da. The deduced amino acid sequence showed substantial similarity to those reported for PabB and TrpE from several organisms. Synonymous codon usage reflected the low G + C content in the genomic DNA of L. lactis subsp. lactis, and therefore differed markedly from the preferred usage in the S. lividans host. The cloned heterologous pabB DNA was expressed in amounts that allowed accumulation of excreted PABA in cultures of S. lividans transformants.

Amino Acid Sequence↗

Sequence analysis of the cloned mRNA coding for glyceraldehyde-3-phosphate dehydrogenase from chicken heart muscle.

Using a cloned cDNA (pGAP30) the nucleotide sequence for chicken glyceraldehyde-3-phosphate dehydrogenase mRNA has been determined. The cDNA insert contains 1051 nucleotides representing the amino acid coding sequence, with the exception of 49 NH2-terminal amino acids, and includes the entire 3'-noncoding region. Sequence information on the missing 5' terminus of the mRNA, not represented in the clone pGAP30, was obtained by extension of the cDNA using an 85-nucleotide-long internal fragment as a primer. Thus the sequence of 310 amino acids of chicken glyceraldehyde-3-phosphate dehydrogenase representing 93% of the complete primary structure could be derived. The coding portion exhibits non-random utilization of synonymous codons with a strong bias for codons with G or C at the third position. The non-coding region contains several octanucleotides which are repeated and shows a potentially stable stem-and-loop structure located towards the end of the mRNA. Hypothetical functional implications of the putative secondary structure are discussed.

Amino Acid Sequence↗

Massive overproduction of dihydrofolate reductase in bacteria as a response to the use of trimethoprim.

Among several observations of greatly increased levels of chromosomal dihydrofolate reductase as a cause of resistance to high concentrations of the antifolate drug trimethoprim, in clinically isolated bacteria, one is described here of a strain of Escherichia coli overproducing dihydrofolate reductase several hundredfold. The chromosomally located resistance gene of this strain was isolated, inserted into a plasmid vector, and analyzed for its nucleotide sequence. The structural gene for the overproduced dihydrofolate reductase was found to be identical to that of E. coli K12, with nine exceptions, of which seven resulted in synonymous codon usage. Two transversions resulted in a substitution of Gly or Trp at amino acid position 30, and of Gln for Glu at position 154. Six of the nine base changes resulted in codons more frequently used. The Gly substitution which leads to a less commonly used codon, was thought to relate to the observed threefold increase in Ki for trimethoprim. Furthermore, a C----T transition was found in the -35 region of the promoter, increasing its homology with the E. coli consensus promoter sequence. In the ribosome-binding area of the resistant strain, finally, seven base changes were observed, two of which resulted in a five-base sequence of complementarity with the 3'-end of ribosomal 16S RNA. The distance between the -10 site of the promoter and the start codon for translation was finally increased one base pair by the insertion of an A at position +9 in the resistant strain. These genetic changes towards more efficient transcriptional and translational start sequences and towards increased mRNA expressivity are interpreted to reflect an evolutionary adaptation to the presence of antifolates.

Base Sequence↗

Analysis of pFQ31, a 8551-bp cryptic plasmid from the symbiotic nitrogen-fixing actinomycete Frankia.

The actinomycete Frankia has never been transformed genetically. To favour the development of Frankia cloning vectors, we have fully sequenced the Frankia alni pFQ31 cryptic plasmid and performed analyses to characterise its coding and non-coding regions. This plasmid is 8551 bp-long and contains 72% G+C. Computer-assisted analyses identified 18 open reading frames (ORFs). These ORFs show a synonymous codon usage different from the one of Frankia chromosomal genes, suggesting an evolutionary bias linked to the nature of the replicon or a horizontal transfer. Three ORFs were found to encode genes likely to be involved in plasmid replication and stability: parFA (partition protein), ptrFA (transcriptional repressor of the GntR family) and repFA (initiation of replication). DNA signatures of a replication origin were identified in the ptrFA-repFA intergenic region. These structural motifs are similar to those observed among origins of iteron-containing plasmids replicating via a θ mode.

Actinomycetales↗

Chromosomal location and evolutionary rate variation in enterobacterial genes.

The basal rate of DNA sequence evolution in enterobacteria, as seen in the extent of divergence between Escherichia coli and Salmonella typhimurium, varies greatly among genes, even when only "silent" sites are considered. The degree of divergence is clearly related to the level of gene expression, reflecting constraints on synonymous codon choice. However, where this constraint is weak, among genes not expressed at high levels, divergence is also related to the chromosomal location of the gene; it appears that genes furthest away from oriC, the origin of replication, have a mutation rate approximately two times that of genes near oriC.

Bias↗

Complementarity of Bacillus subtilis 16S rRNA with sites of antibiotic-dependent ribosome stalling in cat and erm leaders.

Inducible cat and erm genes are regulated by translational attenuation. In this regulatory model, gene activation results from chloramphenicol- or erythromycin-dependent stalling of a ribosome at a precise site in the leader region of cat or erm transcripts. The stalled ribosome is believed to destabilize a downstream region of RNA secondary structure that sequesters the ribosome-binding site for the cat or erm coding sequence. Here we show that the ribosome stall sites in cat and erm leader mRNAs, designated crb and erb, respectively, are largely complementary to an internal sequence in 16S rRNA of Bacillus subtilis. A tetracycline resistance gene that is likely regulated by translational attenuation also contains a sequence in its leader mRNA, trb, which is complementary to a sequence in 16S rRNA that overlaps with the crb and erb complements. An in vivo assay is described which is designed to test whether 16S rRNA of a translating ribosome can interact with the crb sequence in mRNA in an inducer-dependent reaction. The assay compares the growth rate of cells expressing crb-86 with the growth rate of cells lacking crb-86 in the presence of subinhibitory levels of inducers of cat-86, chloramphenicol, fluorothiamphenicol, amicetin, or erythromycin. Under these conditions, crb-86 retarded growth. Deletion of the crb-86 sequence, insertion of ochre mutations into crb-86, or synonymous codon changes in crb-86 that decreased its complementarity with 16S rRNA all eliminated from detection inducer-dependent growth retardation. Lincomycin, a ribosomally targeted antibiotic that is not an inducer of cat-86, failed to selectively retard the growth of cells expressing crb-86. We suggest that cat-86 inducers enable the crb-86 sequence in mRNA to base pair with 16S rRNA of translating ribosome. When the base pairing is extensive, as with crb-86, ribosomes become transiently trapped on crb and are temporarily withdrawn from protein synthesis to the extent that growth rate declines. Site-specific positioning of an antibiotic-stalled ribosome is a hallmark of the translational attenuation model. The proposed rRNA-mRNA interaction may precisely position the ribosome on the stall site and perhaps contributes to stabilizing the ribosome leader mRNA complex.

Bacillus subtilis↗

Identification of two sequences in the cytoplasmic tail of the human immunodeficiency virus type 1 envelope glycoprotein that inhibit cell surface expression.

During synthesis and export of protein, the majority of the human immunodeficiency virus type 1 (HIV-1) Env glycoprotein gp160 is retained in the endoplasmic reticulum (ER) and subsequently ubiquitinated and degraded by proteasomes. Only a small fraction of gp160 appears to be correctly folded and processed and is transported to the cell surface, which makes it difficult to identify negative sequence elements regulating steady-state surface expression of Env at the post-ER level. Moreover, poorly localized mRNA retention sequences inhibiting the nucleocytoplasmic transport of viral transcripts interfere with the identification of these sequence elements. Using two heterologous systems with CD4 or immunoglobulin extracellular/transmembrane domains in combination with the gp160 cytoplasmic domain, we were able to identify two membrane-distal, neighboring motifs, is1 (amino acids 750 to 763) and is2 (amino acids 764 to 785), which inhibited surface expression and induced Golgi localization of the chimeric proteins. To prove that these two elements act similarly in the homologous context of the Env glycoprotein, we generated a synthetic gp160 gene with synonymous codons, the transcripts of which are not retained within the nucleus. In accordance with the results in heterologous systems, an internal deletion of both elements considerably increased surface expression of gp160.

Amino Acid Sequence↗

Complementation of a deletion in the rubella virus p150 nonstructural protein by the viral capsid protein.

Rubella virus (RUB) replicons with an in-frame deletion of 507 nucleotides between two NotI sites in the P150 nonstructural protein (DeltaNotI) do not replicate (as detected by expression of a reporter gene encoded by the replicon) but can be amplified by wild-type helper virus (Tzeng et al., Virology 289:63-73, 2001). Surprisingly, virus with DeltaNotI was viable, and it was hypothesized that this was due to complementation of the NotI deletion by one of the virion structural protein genes. Introduction of the capsid (C) protein gene into DeltaNotI-containing replicons as an in-frame fusion with a reporter gene or cotransfection with both DeltaNotI replicons and RUB replicon or plasmid constructs containing the C gene resulted in replication of the DeltaNotI replicon, confirming the hypothesis that the C gene was the structural protein gene responsible for complementation and demonstrating that complementation could occur either in cis or in trans. Approximately the 5' one-third of the C gene was necessary for complementation. Mutations that prevented translation of the C protein while minimally disturbing the C gene sequence abrogated complementation, while synonymous codon mutations that changed the C gene sequence without affecting the amino acid sequence at the 5' end of the C gene had no effect on complementation, indicating that the C protein, not the C gene RNA, was the moiety responsible for complementation. Complementation occurred at a basic step in the virus replication cycle, because DeltaNotI replicons failed to accumulate detectable virus-specific RNA.

Amino Acid Sequence↗

Evolutionary dynamics of the chloroplast genome in Abutilon (Malvoideae, Malvaceae).

The genus Abutilon Mill. (Malvaceae) comprises approximately 178 species distributed across tropical and subtropical regions, many of which hold significant ornamental, economic, and medicinal value; yet its taxonomic classification remains challenging. In this study, six species were sequenced from herbarium specimens, and the chloroplast (cp.) genomes of ten additional species were assembled de novo from publicly available raw data. Three previously reported cp. genomes were also incorporated to characterise cp. genome structure, identify polymorphic loci, and perform phylogenetic analyses. The cp. genomes ranged from 159,458 to 160,454 bp and exhibited the typical quadripartite structure, with each genome containing 112 unique genes (78 protein-coding, 30 tRNA, and 4 rRNA) that showed conserved content and organisation. These genomes exhibited high similarity in GC content, inverted repeat boundaries, relative synonymous codon usage, amino acid frequencies, and substitution patterns. However, notable variation was observed in the total number of simple sequence repeats, ranging from 70 to 97 per genome. Selection analyses indicated predominant purifying selection, with evidence of episodic positive selection detected in rpoC2, rbcL, and ycf1. Two codons in rbcL were clade-specific and provided phylogenetic signal distinguishing Australian and Old World pantropical species. Nucleotide diversity analysis identified six highly polymorphic intergenic spacers (trnH-psbA, rps19-rpl2, psbT-pbf1, psaC-ndhD, trnR-atpA, and ndhJ-ndhK) that may be suitable for taxonomic studies. The phylogeny from maximum likelihood (ML) and Bayesian inference (BI) resolved two major clades: one comprising an exclusively Australian lineage occurring predominantly in arid and semi-arid environments, and the other a pantropical lineage spanning multiple continents. Abutilon grandifolium was recovered as sister to the remaining sampled Abutilon taxa in both ML and BI analyses, although no biogeographic origin inference can be drawn from this placement pending broader taxon sampling and integration of nuclear genomic data. These findings provide insights into the evolutionary dynamics of the cp. genome in Abutilon and offer a foundational genomic framework for refining Abutilon taxonomy.

Genome, Chloroplast↗

Multilayered nucleotide organization reveals purifying selection and host-driven adaptation in CPV and FPV.

Since feline panleukopenia virus (FPV) is considered the most likely ancestor of canine parvovirus (CPV), comprehensive comparisons of nucleotide organization in corresponding viral genes between CPV and FPV may provide novel insights into the evolutionary dynamics underlying the divergence of these two viruses. Here, we characterize the evolutionary patterns of CPV and FPV genes across multiple levels of nucleotide organization. Both viruses exhibited highly conserved nucleotide usage at nonsynonymous sites, with Ka/Ks patterns consistent with strong purifying selection, whereas synonymous sites showed greater variability. CpG dinucleotides were markedly underrepresented across all four viral genes, suggesting host-associated selective pressure and/or intrinsic nucleotide compositional constraints. Extensive nonrandom biases in synonymous codon usage, codon neighboring nucleotide context, and codon pair usage further revealed fine-scale genomic optimization shaped by natural selection and nucleotide compositional constraints. Structural protein genes (VP1 and VP2) displayed stronger codon usage bias and higher tRNA adaptation than nonstructural genes. Moreover, CPV genes showed greater translational adaptation to feline hosts than to canine hosts. These findings highlight how closely related parvoviruses exploit flexible nucleotide organization to facilitate host adaptation while maintaining essential protein functions.

Animals↗

RNA secondary structure and compensatory evolution.

The classic concept of epistatic fitness interactions between genes has been extended to study interactions within gene regions, especially between nucleotides that are important in maintaining pre-mRNA/mRNA secondary structures. It is shown that the majority of linkage disequilibria found within the Drosophila Adh gene are likely to be caused by epistatic selection operating on RNA secondary structures. A recently proposed method of RNA secondary structure prediction based on DNA sequence comparisons is reviewed and applied to several types of RNAs, including tRNA, rRNA, and mRNA. The patterns of covariation in these RNAs are analyzed based on Kimura's compensatory evolution model. The results suggest that this model describes the substitution process in the pairing regions (helices) of RNA secondary structures well when the helices are evolutionarily conserved and thermodynamically stable, but fails in some other cases. Epistatic selection maintaining pre-mRNA/mRNA secondary structures is compared to weak selective forces that determine features such as base composition and synonymous codon usage. The relationships among these forces and their relative strengths are addressed. Finally, our mutagenesis experiments using the Drosophila Adh locus are reviewed. These experiments analyze long-range compensatory interactions between the 5' and 3' ends of Adh mRNA, the different constraints on secondary structures in introns and exons, and the possible role of secondary structures in RNA splicing.

Alcohol Dehydrogenase↗

An analysis of allelic variation in the ABCA4 gene.

PURPOSE: To assess the allelic variation of the ATP-binding transporter protein (ABCA4). METHODS: A combination of single-strand conformation polymorphism (SSCP) and automated DNA sequencing was used to systematically screen this gene for sequence variations in 374 unrelated probands with a clinical diagnosis of Stargardt disease, 182 patients with age-related macular degeneration (AMD), and 96 normal subjects. RESULTS: There was no significant difference in the proportion of any single variant or class of variant between the control and AMD groups. In contrast, truncating variants, amino acid substitutions, synonymous codon changes, and intronic variants were significantly enriched in patients with Stargardt disease when compared with their presence in subjects without Stargardt disease (Kruskal-Wallis P < 0.0001 for each variant group). Overall, there were 2480 instances of 213 different variants in the ABCA4 gene, including 589 instances of 97 amino acid substitutions, and 45 instances of 33 truncating variants. CONCLUSIONS: Of the 97 amino acid substitutions, 11 occurred at a frequency that made them unlikely to be high-penetrance recessive disease-causing variants (HPRDCV). After accounting for variants in cis, one or more changes that were compatible with HPRDCV were found on 35% of all Stargardt-associated alleles overall. The nucleotide diversity of the ABCA4 coding region, a collective measure of the number and prevalence of polymorphic sites in a region of DNA, was found to be 1.28, a value that is 9 to 400 times greater than that of two other macular disease genes that were examined in a similar fashion (VMD2 and EFEMP1).

ATP-Binding Cassette Transporters↗

Nucleotide sequence of the coding portion of human alpha globin messenger RNA.

The nucleotide sequence of the coding portion of human alpha globin mRNA has been determined by sequence analysis using human alpha globin cDNA cloned in bacterial plasmids. The sequence was obtained by a combination of direct sequence analysis of the cloned cDNA and analysis of cDNA obtained by primer extension, using short restriction endonuclease fragments of cloned alpha cDNA that were hybridized to human globin mRNA and elongated on the mRNA template by viral reverse transcriptase. The human alpha globin mRNA has an unexpectedly high G + C base composition (64.7%), similar to that observed for rabbit globin alpha mRNA, and displays a striking bias in the use of synonym codons for various amino acids. The bias in codon usage of human alpha globin mRNA is similar, with some exceptions, to that previously observed for rabbit alpha globin mRNA as well as for human and rabbit beta globin mRNAs. A detailed restriction endonuclease map of the human alpha globin cDNA is presented.

Amino Acid Sequence↗

[Studies on the molecular evolution of apolipoprotein multigene family].

The apolipoprotein genes represent a large family of genes encoding various binding proteins for plasma lipid transport. Because of their long divergence history, it is not known whether and how these genes have evolved through gene duplication from a common ancestor. To test this possibility and reconstruct a reliable phylogenetic tree, a simple method to evaluate the branch length and its divergence time in unrooted parsimony tree under the condition of non-even evolutionary rate was developed. The tree built from the 26 apolipoprotein sequences by above method clearly shows: (1) The common ancestor of ApoA-I ApoA-II, ApoA-IV, ApoE may appear 460 million years ago in an ordovician vertebrate which may be related with the major apolipoprotein LAL1 and LAL2 in Lamprey from the evidence of sequence alignment; (2) The central role of different selection pressure upon the ancestor gene of apolipoprotein made them evolved into different subgroups; (3) The high evolution rate in rodent ApoE molecules may be related with the existence of a large amount of hidden substitutions and the disruption of synonym codon usage clock in their genome; (4) The evolutionary rate of various branches in parsimony tree is significantly different in which the average UEP of ApoA-I, ApoA-IV is 2.0 MY, ApoA-II 1.7MY, ApoE 2.4MY; (5) The receptor domain in ApoE seems to be more conservative than other fragments. These data suggest a long, complex evolutionary history for apolipoprotein genes in which the gene duplication events of different origins took place.

Amino Acid Sequence↗

[Molecular evolution of MHC DQA genes. II. Phylogenetic analysis based on nucleotide substitution and SCU bias].

Phylogenetics of 23 alleles at MHC DQA loci in 7 mammalian species was studied based on their nucleotide (NT) substitution and synonymous codon usage (SCU) bias. (1) It was demonstrated that the NT substitution rates are 1.0 x 10(-9) NT/site/yr for exon2 and 1.3 x 10(-9) NT/site/yr for exon2-4 in a large time scale, which is similar to other nuclear genes, while for mouse and rat the rates are nearly twice as high as above mentioned. (2) The DQA locus diversity and their interallelic diversity developed long after the radiation of mammalian 80Mya (million years ago). The bovine counterpart, of, and with the same recent ancestor of ovine DQA2, remains to be discovered. HLA-DQA2 locus split from HLA-DQA1 ancestor at the time between 12 approximately 20 Mya while allele diversity of HLA-DQA1 emerged and developed from 24 Mya to less than 1 Mya. (3) The phylogenetic trees based on SCU divergence reflect the phylogenetics of MHC DQA genes quite well generally in a new respect and reveal that HLA-DQA2 has a distinctive SCU bias different from all other MHC DQA locianalyzed. It indicates that SCU statistics plays an important and unique role in phylogenetic analysis of orthologous genes. The method to estimate the SCU divergence and SCU similarity was improved in this research.

Animals↗