PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Internal structure of the silk fibroin gene of Bombyx mori. I The fibroin gene consists of a homogeneous alternating array of repetitious crystalline and amorphous coding sequences.

The DNA sequence orgainzation of the protein encoding region of the gene for silk fibroin has been analyzed. The accompanying paper (Manningm R. F., and Gage, L. P. (1980) J. Biol. Chem. 255, 9451-9457) shows that the total length of the gene, and its protein, as well as the pattern of restriction sites in the gene is highly polymorphic among inbred stocks of Bombyx mori, In this paper, those features of fibroin gene structure which are invariant among these alleles are presented. Fibroin is composed primarily of relatively short "crystalline" and "amorphous" peptides of known sequence whose arrangement in the protein is unknown. Knowledge of the codons most commonly used in fibroin mRNA allowed utilization of particular restriction inzymes as a means for determing the nature and organization of crystalline and amorphous coding sequences in the fibroin gene. Three restriction endonucleases were identified that cleve sequences coding for amorphous region peptides. Their cleavage pattern revelaed that the repetitive coding sequence of the gene core (approximately 15 kilobases) is divided into at least 10 large crystalline coding domains interrupted by smaller amorphous coding domains. Many restriction endoncleases do not cleave the fibroin core at all, three of them with four gase recognition sequences. Specific deductions as to codon usage and repetitive sequence homogeneity in the gene follow from these results. One novel finding is the rigorous exclusion of the glycine codon GGA prior to serine codons even though this glycine codon is used frequently prior to alanine codons. The sequence homogeneity and the regularly alternating arrangement of crystalline and amorphous coding sequences of the gene are discussed in terms of the function of fibroin protein and the evolution of highly repetitive DNA.

Alleles↗

[Why does the DNA code contain 4 letters?].

The answer to this question is not yet known. There are two ways to express information, i.e., to reduce the number of letters in alphabet (n), which simplifies the decoding machine, but leads to longer informational sequences, or to increase n, which shortens sequences, but complicates the informational machine. The compromise between these two possibilities would be to obtain the minimum of one of summary informational component's parameters. The summary component is the sum of corresponding decoding machine's and the program's parameters. In this work it was demonstrated that DNA four-letter code is optimal, for it allows the minimal volume of summary cell informational contents. But it is so only for the most simple DNA. Our calculations may indirectly show that such DNA (and not more complicated) was the object of "projecting" at one of the biological evolution's early stages.

Amino Acid Sequence↗

Mitochondrial genetic codes evolve to match amino acid requirements of proteins.

Mitochondria often use genetic codes different from the standard genetic code. Now that many mitochondrial genomes have been sequenced, these variant codes provide the first opportunity to examine empirically the processes that produce new genetic codes. The key question is: Are codon reassignments the sole result of mutation and genetic drift? Or are they the result of natural selection? Here we present an analysis of 24 phylogenetically independent codon reassignments in mitochondria. Although the mutation-drift hypothesis can explain reassignments from stop to an amino acid, we found that it cannot explain reassignments from one amino acid to another. In particular--and contrary to the predictions of the mutation-drift hypothesis--the codon involved in such a reassignment was not rare in the ancestral genome. Instead, such reassignments appear to take place while the codon is in use at an appreciable frequency. Moreover, the comparison of inferred amino acid usage in the ancestral genome with the neutral expectation shows that the amino acid gaining the codon was selectively favored over the amino acid losing the codon. These results are consistent with a simple model of weak selection on the amino acid composition of proteins in which codon reassignments are selected because they compensate for multiple slightly deleterious mutations throughout the mitochondrial genome. We propose that the selection pressure is for reduced protein synthesis cost: most reassignments give amino acids that are less expensive to synthesize. Taken together, our results strongly suggest that mitochondrial genetic codes evolve to match the amino acid requirements of proteins.

Amino Acids↗

Primate retroviruses: envelope glycoproteins of endogenous type C and type D viruses possess common interspecies antigenic determinants.

The major 70,000- to 80,000-molecular-weight envelope glycoproteins of the squirrel monkey retrovirus, Mason-Pfizer monkey virus, and M7 baboon virus and the related endogenous feline virus, RD114, were isolated and immunologically characterized. Immunoprecipitation and competition immunoassay analysis revealed these viral envelope glycoproteins to possess several distinct classes of immunological determinants. These include species-specific determinants, group-specific antigenic determinants unique to endogenous primate type C viruses, and group-specific determinants for type D viruses such as Mason-Pfizer monkey virus and squirrel monkey retrovirus. In addition, a class of broadly reactive antigenic determinants shared by envelope glycoproteins of both type C viruses of the baboon/RD114 group and type D viruses of the Mason-Pfizer monkey virus/squirrel monkey virus group are described. Other mammalian oncornaviruses tested, including isolates of nonprimate origin and representative type B viruses, lacked these determinants. The demonstration of antigenic determinants specific to envelope glycoproteins of type C and type D primate viruses indicates either that these viruses are evolutionarily related or that genetic recombination occurred between their progenitors. Alternatively, endogenous type D oncornaviruses may be replication defective, and acquisition of endogenous type C viral genetic sequences coding for envelope glycoprotein determinants may be necessary for their isolation as infectious virus.

Animals↗

Chromosomal fragility, structural rearrangements and mobile element activity may reflect dynamic epigenetic mechanisms of importance in neurobehavioural genetics.

Advances in human genome analyses have not yet allowed identification of specific genetic mechanisms underlying the expression of human neurobehavioural disorders. There is an increasing awareness that several genes may contribute to behavioural phenotypes and these genes appear to interact in as yet undetermined ways. It has been suggested that the problem needs elucidation from an epigenetic, gene expression perspective. Cytogenetic instability manifesting as chromosomal fragile sites, translocations, duplications, deletions and inversions, when co-occurring with neurobehavioural disorders, may offer a doorway to the investigation of such chromatin level, regulatory region, epigenetic processes. Due to earlier indications of non-specificity of chromosomal aberrations, poor phenotype:genotype correlations and a shift to analysing candidate coding regions on high resolution map level, the only utility of chromosomal breakpoints came to be seen as harbouring possible candidate genes of interest when segregating together with particular neurobehavioural disorders. More recent findings of the expression of highly specific subsets of fragile sites in association with Tourette and Rett syndromes need to be extended to other neurobehavioural disorders to ascertain whether observed patterns can be considered representative of 'chromatin endophenotypes' correlating with discrete sets of neurobehavioural symptoms. Environmental/epigenetic factors could affect the chromatin characteristics of the genome arising through DNA strand breakage, mobile element activity and retroinsertion, establishing new architectural features of regulatory control networks very rapidly in comparison to coding region evolution rates. Microarray-based techniques for the genome-wide mapping of in vivo protein-DNA interactions offer increasingly comprehensive views of genetic and epigenetic regulatory networks. It may be informative to include functionally significant chromatin structural variation analyses when considering candidate genes for neurobehavioural disorders.

Cell Cycle↗

The evolution of transcriptional regulation in eukaryotes.

Gene expression is central to the genotype-phenotype relationship in all organisms, and it is an important component of the genetic basis for evolutionary change in diverse aspects of phenotype. However, the evolution of transcriptional regulation remains understudied and poorly understood. Here we review the evolutionary dynamics of promoter, or cis-regulatory, sequences and the evolutionary mechanisms that shape them. Existing evidence indicates that populations harbor extensive genetic variation in promoter sequences, that a substantial fraction of this variation has consequences for both biochemical and organismal phenotype, and that some of this functional variation is sorted by selection. As with protein-coding sequences, rates and patterns of promoter sequence evolution differ considerably among loci and among clades for reasons that are not well understood. Studying the evolution of transcriptional regulation poses empirical and conceptual challenges beyond those typically encountered in analyses of coding sequence evolution: promoter organization is much less regular than that of coding sequences, and sequences required for the transcription of each locus reside at multiple other loci in the genome. Because of the strong context-dependence of transcriptional regulation, sequence inspection alone provides limited information about promoter function. Understanding the functional consequences of sequence differences among promoters generally requires biochemical and in vivo functional assays. Despite these challenges, important insights have already been gained into the evolution of transcriptional regulation, and the pace of discovery is accelerating.

Animals↗

A maximum likelihood method for detecting functional divergence at individual codon sites, with application to gene family evolution.

The tailoring of existing genetic systems to new uses is called genetic co-option. Mechanisms of genetic co-option have been difficult to study because of difficulties in identifying functionally important changes. One way to study genetic co-option in protein-coding genes is to identify those amino acid sites that have experienced changes in selective pressure following a genetic co-option event. In this paper we present a maximum likelihood method useful for measuring divergent selective pressures and identifying the amino acid sites affected by divergent selection. The method is based on a codon model of evolution and uses the nonsynonymous-to-synonymous rate ratio (omega) as a measure of selection on the protein, with omega = 1, < 1, and > 1 indicating neutral evolution, purifying selection, and positive selection, respectively. The model allows variation in omega among sites, with a fraction of sites evolving under divergent selective pressures. Divergent selection is indicated by different omega's between clades, such as between paralogous clades of a gene family. We applied the codon model to duplication followed by functional divergence of (i) the epsilon and gamma globin genes and (ii) the eosinophil cationic protein (ECP) and eosinophil-derived neurotoxin (EDN) genes. In both cases likelihood ratio tests suggested the presence of sites evolving under divergent selective pressures. Results of the epsilon and gamma globin analysis suggested that divergent selective pressures might be a consequence of a weakened relationship between fetal hemoglobin and 2,3-diphosphoglycerate. We suggest that empirical Bayesian identification of sites evolving under divergent selective pressures, combined with structural and functional information, can provide a valuable framework for identifying and studying mechanisms of genetic co-option. Limitations of the new method are discussed.

Amino Acids↗

Molecular evolution of the Est-6 gene in Drosophila melanogaster: contrasting patterns of DNA variability in adjacent functional regions.

We have investigated nucleotide polymorphism at the esterase 6 gene (Est-6) gene, including the complete coding region (1686 bp), as well as the 5'-flanking (1183 bp) and 3'-flanking (193 bp) regions of the gene, in 30 strains of Drosophila melanogaster and in one strain of Drosophila simulans. The level of silent variation is similar in the coding and in the 3'-flanking region, but smaller in the 5'-flanking region. Strong linkage disequilibrium occurs within each region; and also, although less pronounced, between the 5'-flanking region and the rest of the gene, including the 3'-flanking region. We suggest that the pattern of nucleotide polymorphism of Est-6 may be shaped by: (1) directional and balancing selection acting on the promoter and the coding region; and (2) interactions between the two regions that involve variable degrees of hitchhiking. The patterns of linkage disequilibrium, as well as the statistics Z(nS) (Genetics 146 (1997)1197) and B and Q (Genet. Res. 74 (1999) 65), may be interpreted as there being multiple targets of selection within the gene. The previously reported Est-6 allozyme latitudinal clines may be accounted for by the interaction between selective processes in the promoter and coding regions.

Animals↗

Analysis of a new human parechovirus allows the definition of parechovirus types and the identification of RNA structural domains.

Human parechoviruses (HPeV), members of the Parechovirus genus of Picornaviridae, are frequent pathogens but have been comparatively poorly studied, and little is known of their diversity, evolution, and molecular biology. To increase the amount of information available, we have analyzed 7 HPeV strains isolated in California between 1973 and 1992. We found that, on the basis of VP1 sequences, these fall into two genetic groups, one of which has not been previously observed, bringing the number of known groups to five. While these correlate partly with the three known serotypes, two members of the HPeV2 serotype belong to different genetic groups. In view of the growing importance of molecular techniques in diagnosis, we suggest that genotype is an important criterion for identifying viruses, and we propose that the genetic groups we have defined should be termed human parechovirus types 1 to 5. Complete nucleotide sequence analysis of two of the Californian isolates, representing two types, confirmed the identification of a new genetic group and suggested a role for recombination in parechovirus evolution. It also allowed the identification of a putative HPeV1 cis-acting replication element, which is located in the VP0 coding region, as well as the refinement of previously predicted 5' and 3' untranslated region structures. Thus, the results have significantly improved our understanding of these common pathogens.

Amino Acid Sequence↗

Recent advances in producing and selecting functional proteins by using cell-free translation.

Prokaryotic and eukaryotic in vitro translation systems have recently become the focus of increasing interest for tackling fundamental problems in biochemistry. Cell-free systems can now be used to study the in vitro assembly of membrane proteins and viral particles, rapidly produce and analyze protein mutants, and enlarge the genetic code by incorporating unnatural amino acids. Using in vitro translation systems, display techniques of great potential have been developed for protein selection and evolution. Furthermore, progress has been made to efficiently produce proteins in batch or continuous cell-free translation systems and to elucidate the molecular causes of low yield and find possible solutions for this problem.

Animals↗

Azolla--a model organism for plant genomic studies.

The aquatic ferns of the genus Azolla are nitrogen-fixing plants that have great potentials in agricultural production and environmental conservation. Azolla in many aspects is qualified to serve as a model organism for genomic studies because of its importance in agriculture, its unique position in plant evolution, its symbiotic relationship with the N2-fixing cyanobacterium, Anabaena azollae, and its moderate-sized genome. The goals of this genome project are not only to understand the biology of the Azolla genome to promote its applications in biological research and agriculture practice but also to gain critical insights about evolution of plant genomes. Together with the strategic and technical improvement as well as cost reduction of DNA sequencing, the deciphering of their genetic code is imminent.

Cyanobacteria↗

Chromosome-level genome assembly of the longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae).

The longhorn beetle Arhopalus rusticus (Coleoptera: Cerambycidae) is a widely distributed wood-boring pest of conifers. Here, we assembled a chromosome-level genome of A. rusticus using Illumina, Oxford Nanopore, and Hi-C sequencing technologies. The assembled genome is 1180.40&#x2009;Mb, with a scaffold N50 of 125.01&#x2009;Mb, and BUSCO completeness of 93.6%. All contigs were assembled into ten pseudo-chromosomes. The genome contains 69.87% repeat sequences. We identify 18, 377 protein-coding genes in the genome, of which 11,368 were functionally annotated. This genome provides a valuable resource for understanding the ecology, genetics, and evolution of A. rusticus, as well as for controlling wood-boring pests.

Animals↗

Taking variation of evolutionary rates between sites into account in inferring phylogenies.

As methods of molecular phylogeny have become more explicit and more biologically realistic following the pioneering work of Thomas Jukes, they have had to relax their initial assumption that rates of evolution were equal at all sites. Distance matrix and likelihood methods of inferring phylogenies make this assumption; parsimony, when valid, is less limited by it. Nucleotide sequences, including RNA sequences, can show substantial rate variation; protein sequences show rates that vary much more widely. Assuming a prior distribution of rates such as a gamma distribution or lognormal distribution has deservedly been popular, but for likelihood methods it leads to computational difficulties. These can be resolved using hidden Markov model (HMM) methods which approximate the distribution by one with a modest number of discrete rates. Generalized Laguerre quadrature can be used to improve the selection of rates and their probabilities so as to more nearly approach the desired gamma distribution. A model based on population genetics is presented predicting how the rates of evolution might vary from locus to locus. Challenges for the future include allowing rates at a given site to vary along the tree, as in the "covarion" model, and allowing them to have correlations that reflect three-dimensional structure, rather than position in the coding sequence. Markov chain Monte Carlo likelihood methods may be the only practical way to carry out computations for these models.

Algorithms↗

Molecular evolution and phylogenetic utility of the petD group II intron: a case study in basal angiosperms.

Sequences of spacers and group I introns in plant chloroplast genomes have recently been shown to be very effective in phylogenetic reconstruction at higher taxonomic levels and not only for inferring relationships among species. Group II introns, being more frequent in those genomes than group I introns, may be further promising markers. Because group II introns are structurally constrained, we assumed that sequences of a group II intron should be alignable across seed plants. We designed universal amplification primers for the petD intron and sequenced this intron in a representative selection of 47 angiosperms and three gymnosperms. Our sampling of taxa is the most representative of major seed plant lineages to date for group II introns. Through differential analysis of structural partitions, we studied patterns of molecular evolution and their contribution to phylogenetic signal. Nonpairing stretches (loops, bulges, and interhelical nucleotides) were considerably more variable in both substitutions and indels than in helical elements. Differences among the domains are basically a function of their structural composition. After the exclusion of four mutational hotspots accounting for less than 18% of sequence length, which are located in loops of domains I and IV, all sequences could be aligned unambiguously across seed plants. Microstructural changes predominantly occurred in loop regions and are mostly simple sequence repeats. An indel matrix comprising 241 characters revealed microstructural changes to be of lower homoplasy than are substitutions. In showing Amborella first branching and providing support for a magnoliid clade through a synapomorphic indel, the petD data set proved effective in testing between alternative hypotheses on the basal nodes of the angiosperm tree. Within angiosperms, group II introns offer phylogenetic signal that is intermediate in information content between that of spacers and group I introns on the one hand and coding sequences on the other.

Base Sequence↗

MyESL: A Software for Evolutionary Sparse Learning in Molecular Phylogenetics and Genomics.

Evolutionary sparse learning uses supervised machine learning to build evolutionary models where genomic sites loci are parameters. It uses the Least Absolute Shrinkage and Selection Operator with bi-level sparsity to connect a specific phylogenetic hypothesis with sequence variation across genomic loci. The MyESL software addresses the need for open-source tools to perform evolutionary sparse learning analyses, offering features to preprocess input phylogenomic alignments, post-process output models to generate molecular evolutionary metrics, and make Least Absolute Shrinkage and Selection Operator regression adaptable and efficient for phylogenetic trees and alignments. The core of MyESL, which constructs models with logistic regressions using bi-level sparsity, is written in C++. Its input data preprocessing and result post-processing tools are developed in Python. Compared to other tools, MyESL is more computationally efficient and provides evolution-friendly inputs and outputs. These features have already enabled the use of MyESL in two phylogenomic applications, one to identify outlier sequences and fragile clades in inferred phylogenies and another to build genetic models of convergent traits. In addition to the use in a Python environment, MyESL is available as a standalone executable compatible across multiple platforms, which can be directly integrated into scripts and third-party software. The source code, executable, and documentation for MyESL are openly accessible at https://github.com/kumarlabgit/MyESL.

Phylogeny↗

A comparison of methods for self-adaptation in evolutionary algorithms.

Evolutionary algorithms, including evolutionary programming and evolution strategies, have often been applied to real-valued function optimization problems. These algorithms generally operate directly on the real values to be optimized, in contrast with genetic algorithms which usually operate on a separately coded transformation of the objective variables. Evolutionary algorithms often rely on a second-level optimization of strategy parameters, tunable variables that in part determine how each parent will generate offspring. Two alternative methods for performing this second-level optimization have been proposed and are compared across a series of function optimization tasks. The results appear to favor the approach offered originally in evolution strategies, although the applicability of the findings may be limited to the case where each parameter of a parent solution is perturbed independently of all others.

Algorithms↗

Neo-Lamarckian medicine.

Darwinian medicine is the treatment of disease based on evolution. The underlying assumption of Darwinian medicine is that traits are coded by genes, which are often assumed to be sequences of DNA nucleotides. The quantitative genetic ramification of this perspective is that traits, including disease susceptibility, are either caused by genes or by the environment, with genotype-by-environment interactions usually considered statistical artefacts. I emphasize also examining those epigenetic signals that can be altered by environmental perturbations and then transmitted to subsequent generations. Although seldom studied, environmentally-alterable meiotically-heritable epigenetic signals exist and provide a mechanism underlying genotype-by-environment interactions. Environment of a parent can affect its descendants by heritably altering epigenetic signals. Neo-Lamarckian medicine is the application of these evolutionary epigenetic notions to diseases and could have enormous public health and environmental policy implications. If industrial contaminants adversely affect organisms by meiotically-heritably altering their epigenetic signals, then cleaning up these contaminants will not remedy the problem. Once contaminants have adversely altered an individual's epigenetic signals, this harm will be transmitted to future generations even if they are not exposed to the contaminant. Exposure to environmental shocks such as free radicals or other carcinogens can alter cytosine methylation patterns on regulatory genes. This can cause cancer by up-regulating genes for cell division or by down-regulating tumour suppressor genes. Environmentally-alterable meiotically-heritable epigenetic signals could also underlie other diseases, such as diabetes, Prader-Willi syndrome, and many complex diseases. If environmentally-altered meiotically-heritable epigenetic effects are widespread - which is an important open empirical question - they have the potential to alter paradigmatic views of evolutionary medicine and the putative dichotomy of nature versus nurture. Neo-Lamarckian medicine would thereby shift emphasis from cure to prevention of diseases.

Animals↗