PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

The evolution of cooperation and altruism--a general framework and a classification of models.

One of the enduring puzzles in biology and the social sciences is the origin and persistence of intraspecific cooperation and altruism in humans and other species. Hundreds of theoretical models have been proposed and there is much confusion about the relationship between these models. To clarify the situation, we developed a synthetic conceptual framework that delineates the conditions necessary for the evolution of altruism and cooperation. We show that at least one of the four following conditions needs to be fulfilled: direct benefits to the focal individual performing a cooperative act; direct or indirect information allowing a better than random guess about whether a given individual will behave cooperatively in repeated reciprocal interactions; preferential interactions between related individuals; and genetic correlation between genes coding for altruism and phenotypic traits that can be identified. When one or more of these conditions are met, altruism or cooperation can evolve if the cost-to-benefit ratio of altruistic and cooperative acts is greater than a threshold value. The cost-to-benefit ratio can be altered by coercion, punishment and policing which therefore act as mechanisms facilitating the evolution of altruism and cooperation. All the models proposed so far are explicitly or implicitly built on these general principles, allowing us to classify them into four general categories.

Altruism↗

The fidelity of the translation of the genetic code.

Aminoacyl-tRNA synthetases play a central role in maintaining accuracy during the translation of the genetic code. To achieve this challenging task they have to discriminate against amino acids that are very closely related not only in structure but also in chemical nature. A 'double-sieve' editing model was proposed in the late seventies to explain how two closely related amino acids may be discriminated. However, a clear understanding of this mechanism required structural information on synthetases that are faced with such a problem of amino acid discrimination. The first structural basis for the editing model came recently from the crystal structure of isoleucyl-tRNA synthetase, a class I synthetase, which has to discriminate against valine. The structure showed the presence of two catalytic sites in the same enzyme, one for activation, a coarse sieve which binds both isoleucine and valine, and another for editing, a fine sieve which binds only valine and rejects isoleucine. Another structure of the enzyme in complex with tRNA showed that the tRNA is responsible for the translocation of the misactivated amino-acid substrate from the catalytic site to the editing site. These studies were mainly focused on class I synthetases and the situation was not clear about how class II enzymes discriminate against similar amino acids. The recent structural and enzymatic studies on threonyl-tRNA synthetase, a class II enzyme, reveal how this challenging task is achieved by using a unique zinc ion in the active site as well as by employing a separate domain for specific editing activity. These studies led us to propose a model which emphasizes the mirror symmetrical approach of the two classes of enzymes and highlights that tRNA is the key player in the evolution of these class of enzymes.

Amino Acids↗

From immunogenetics to immunomics: functional prospecting of genes and transcripts.

Human and mouse genome and transcriptome projects have expanded the field of 'immunogenetics' beyond the traditional study of the genetics and evolution of MHC, TCR and Ig loci into the new interdisciplinary area of 'immunomics'. Immunomics is the study of the molecular functions associated with all immune-related coding and non-coding mRNA transcripts. To unravel the function, regulation and diversity of the immunome requires that we identify and correctly categorize all immune-related transcripts. The importance of intercalated genes, antisense transcripts and non-coding RNAs and their potential role in regulation of immune development and function are only just starting to be appreciated. To better understand immune function and regulation, transcriptome projects (e.g. Functional Annotation of the Mouse, FANTOM), that focus on sequencing full-length transcripts from multiple tissue sources, ideally should include specific immune cells (e.g. T cell, B cells, macrophages, dendritic cells) at various states of development, in activated and unactivated states and in different disease contexts. Progress in deciphering immune regulatory networks will require the cooperative efforts of immunologists, immunogeneticists, molecular biologists and bioinformaticians. Although primary sequence analysis remains useful for annotation of new transcripts it is less useful for identifying novel functions of known transcripts in a new context (protein interaction network or pathway). The most efficient approach to mine useful information from the vast a priori knowledge contained in biological databases and the scientific literature, is to use a combination of computational and expert-driven knowledge discovery strategies. This paper will illustrate the challenges posed in attempts to functionally infer transcriptional regulation and interaction of immune-related genes from text and sequence-based data sources.

Alternative Splicing↗

Relative rates of nucleotide substitution in the chloroplast genome.

Coding sequences from maize, rice, tobacco, and liverwort chloroplasts are aligned and subjected to relative rate tests. Results of rate tests suggest that coding sequences from maize and rice are evolving with homogeneous rates of nucleotide substitution while coding sequences from the grass lineages (i.e., maize and rice) are evolving at a faster rate than coding sequences from the tobacco chloroplast. Rate tests also suggest that particular loci evolve at significantly faster rates in grass chloroplast genomes than the tobacco chloroplast genome. These loci encode proteins important to RNA polymerase, the H(+)-ATPase complex, and the ribosomal proteins. Much of the variation at these loci can be attributed to differences in nonsynonymous substitution rates. Taken together, these studies suggest that the chloroplast DNA molecular clock varies both between evolutionary lineages and between protein coding loci.

Biological Evolution↗

siRNA-mediated transcriptional gene silencing: the potential mechanism and a possible role in the histone code.

Epigenetics is the study of meiotically and mitotically heritable changes in gene expression which are not coded for in the DNA. Three distinct mechanisms appear intricately related in initiating and sustaining epigenetic modifications: RNA-associated silencing, DNA methylation and histone modification. Recently, in human cells small-interfering RNAs (siRNAs) have been shown to mediate transcriptional gene silencing (TGS). The observation that siRNAs can function to suppress gene expression at the level of transcription has created a major paradigm shift in mammalian RNA interference. The putative mechanism(s) of siRNA-mediated TGS in both yeast and human cells will be discussed. Undoubtedly, the ramifications from this paradigm shift in which RNA has demonstrated a potent and specific capability to regulate the expression of the gene are immeasurable both therapeutically (i. e. directed control of gene expression) and biologically in understanding the evolution of the cell.

Animals↗

Avoidance of inter-repeat recombination by sequence divergence and a mechanism of neutral evolution.

Eucaryotic genomes are loaded with diverse repeated sequences and are therefore threatened by rearrangements via inter-repeat crossovers and by gene-inactivating conversions between genes and their inactive pseudogenes. Such repeated DNA sequences are usually diverged and polymorphic. Sequence divergence by well-spread point mutations is a potent inhibitor of homologous recombination due to the loss of recombination initiation sites and to the editing of recombinational intermediates by the mismatch repair system. Evidence is reviewed suggesting that a germ line process can identify duplicated sequences by homologous pairing, modify them by methylation and mutate by C----T transitions. Since this process requires a minimum contiguous homology that is larger than the average exon size, it is proposed that fragmentation by intron inserts protects the coding sequences from inactivation by homologous interactions with their pseudogene sequences.

Animals↗

[Possible vertebrate use of a mechanism of chromatin reduction to fuse the V and C genes of immunoglobulins].

Diminution of a chromatin fragment separating V- and C-immunoglobulin genes is suggested as possible mechanism of fusion of V- and C-genes into one functionally active cisrtron. During the early stage of immunodifferentiation of lymphocytes, the chromatin fibrills unfold, and V- and C-genes bring together by DNA-DNA and DNA-protein interaction. The endonucleases excise the chromatin fragment interlocated between V- and C-genes, then the repair enzymes restore the intactness of chromating structure by fusion of V- and C-genes. A number of human pathologic syndromes, the elimination of nucleus during erythrocyte differentiation in mammals indicate the possibility of manifestation on diminution mechanism in vertebrates. Fusion of V- and C-genes by chromatin diminution fits the number of immunologic phenomena, and it may display for different paths of V-gene coding and may be experimentally verified.

Animals↗

Phylogenetic analysis of the evolution of lactose digestion in adults.

In most of the world's population the ability to digest lactose declines sharply after infancy. High lactose digestion capacity in adults is common only in populations of European and circum-Mediterranean origin and is thought to be an evolutionary adaptation to millennia of drinking milk from domestic livestock. Milk can also be consumed in a processed form, such as cheese or soured milk, which has a reduced lactose content. Two other selective pressures for drinking fresh milk with a high lactose content have been proposed: promotion of calcium uptake in high-latitude populations prone to vitamin-D deficiency and maintainance of water and electrolytes in the body in highly and environments. These three hypotheses are all supported by the geographic distribution of high lactose digestion capacity in adults. However, the relationships between environmental variables and adult lactose digestion capacity are highly confounded by the shared ancestry of many populations whose lactose digestion capacity has been tested. The three hypotheses for the evolution of high adult lactose digestion capacity are tested here using a comparative method of analysis that takes the problem of phylogenetic confounding into account. This analysis supports the hypothesis that high adult lactose digestion capacity is an adaptation to dairying but does not support the hypotheses that lactose digestion capacity is additionally selected for either at high latitudes or in highly arid environments. Furthermore, methods using maximum likelihood are used to show that the evolution of milking preceded the evolution of high lactose digestion.

Adult↗

Concerted evolution in the repeats of an immunomodulating cell surface protein, SOWgp, of the human pathogenic fungi Coccidioides immitis and C. posadasii.

Genome dynamics that allow pathogens to escape host immune responses are fundamental to our understanding of host-pathogen interactions. Here we present the first population-based study of the process of concerted evolution in the repetitive domain of a protein-coding gene. This gene, SOWgp, encodes the immunodominant protein in the parasitic phase of the human pathogenic fungi Coccidioides immitis and C. posadasii. We sequenced the entire gene from strains representing the geographic ranges of the two Coccidioides species. By using phylogenetic and genetic distance analyses we discovered that the repetitive part of SOWgp evolves by concerted evolution, predominantly by the mechanism of unequal crossing over. We implemented a mathematical model originally developed for multigene families to estimate the rate of homogenization and recombination of the repetitive array, and the results indicate that the pattern of concerted evolution is a result of homogenization of repeat units proceeding at a rate close to the nucleotide point mutation rate. The release of the SOWgp molecules by the pathogen during proliferation may mislead the host: we speculate that the pathogen benefits from concerted evolution of repeated domains in SOWgp by an enhanced ability to misdirect the host's immune system.

Antigens, Fungal↗

Comparative genetics and evolution of annexin A13 as the founder gene of vertebrate annexins.

Annexin A13 (ANXA13) is believed to be the original founder gene of the 12-member vertebrate annexin A family, and it has acquired an intestine-specific expression associated with a highly differentiated intracellular transport function. Molecular characterization of this subfamily in a range of vertebrate species was undertaken to assess coding region conservation, gene organization, chromosomal linkage, and phylogenetic relationships relevant to its progenitor role in the structure-function evolution of the annexin gene superfamily. Protein diagnostic features peculiar to this subfamily include an alternate isoform containing a KGD motif, an elevated basic amino acid content with polyhistidine expansion in the 5'-translated region, and the conservation of 15% core tetrad residues specific to annexin A13 members. The 12 coding exons comprising the 58-kb human ANXA13 gene were deduced from BAC clone sequencing, whereas internal repetitive elements and neighboring genes in chromosome 8q24.12 were identified by contig analysis of the draft sequence from the human genome project. A unique exon splicing pattern in the annexin A13 gene was corroborated by coanalysis of mouse, rat, zebrafish, and pufferfish genomic DNA and determined to be the most distinct of all vertebrate annexins. The putative promoter region was identified by phylogenetic footprinting of potential binding sites for intestine-specific transcription factors. Mouse annexin A13 cDNA was used to map the gene to an orthologous linkage group in mouse chromosome 15 (between Sdc2 and Myc by backcross analysis), and the zebrafish cDNA permitted its localization to linkage group 24. Comparative analysis of annexin A13 from nine species traced this gene's speciation history and assessed coding region variation, whereas phylogenetic analysis showed it to be the deepest-branching vertebrate annexin, and computational analysis estimated the gene age and divergence rate. The unique, conserved aspects of annexin A13 primary structure, gene organization, and genetic maps identify it as the probable common ancestor of all vertebrate annexins, beginning with the sequential duplication to annexins A7 and A11 approximately 700 MYA, before the emergence of chordates.

Alternative Splicing↗

Deviations from compositional randomness in eukaryotic and prokaryotic proteins: the hypothesis of selective-stochastic stability and a principle of charge conservation.

Eight proteins of diverse lengths, functions, and origin, are examined for compositional non-randomness amino acid by amino acid. The proteins investigated are human fibrinopeptide A, guinea pig Insulin, rattlesnake cytochrome c, MS2 phage coat protein, rabbit triosephosphate isomerase, bovine pancreatic deoxyribonuclease A, bovine glutamate dehydrogenase, and Bacillus thermoproteolyticus thermolysin. As a result of this study the experimentally testable hypothesis is put forth that for a large class of proteins the ratio of that fraction of the molecule which exhibits compositional non-randomness to that fraction which does not is on the average, stable about a mean value (estimated as 0.32 plus or minus 0.17) and (nearly) independent of protein length. Stochastic and selective evolutionary forces are viewed as interacting rather than independent phenomena. With respect to amino acid composition, this coupling ameliorates the current controversy over Darwinian vs. non-Darwinian evolution, selectionist vs. neutralist, in favor of neither: Within the context of the quantitative data, the evolution of real proteins is seen as a compromise between the two viewpoints, both important. The compositional fluctuations of the electrically charged amino acids glutamic and aspartic acid, lysine and arginine, are examined in depth for over eighty protein families, both prokaryotic and eukaryotic. For both taxa, each of the acidic amino acids is present in amounts roughly twice that predicted from the genetic code. The presence of an excess of glutamic acid is independent of the presence of an excess of aspartic acid and vice versa.

Amino Acids↗

Sequence and structural organization of the human gene encoding ciliary neurotrophic factor.

Ciliary neurotrophic factor (CNTF) is a potent polypeptide hormone whose actions appear to be restricted to the nervous system where it promotes survival, neurotransmitter synthesis and neurite outgrowth in certain neuronal populations. We have cloned the gene encoding human CNTF (hCNTF) and have characterized its structure and organization. The hCNTF gene appears to be a unique-copy gene with a simple genetic organization, since only a single intron interrupts the coding domain. The hCNTF gene is located on chromosome 11, as determined using human-hamster somatic cell hybrids. The CNTF protein is highly conserved in evolution. The amino acid (aa) sequences of rat and rabbit CNTF translated from cDNAs display approx. 85% homology with the deduced aa sequence encoding hCNTF.

Amino Acid Sequence↗

Determinants of rate variation in mammalian DNA sequence evolution.

Attempts to analyze variation in the rates of molecular evolution among mammalian lineages have been hampered by paucity of data and by nonindependent comparisons. Using phylogenetically independent comparisons, we test three explanations for rate variation which predict correlations between rate variation and generation time, metabolic rate, and body size. Mitochondrial and nuclear genes, protein coding, rRNA, and nontranslated sequences from 61 mammal species representing 14 orders are used to compare the relative rates of sequence evolution. Correlation analyses performed on differences in genetic distance since common origin of each pair against differences in body mass, generation time, and metabolic rate reveal that substitution rate at fourfold degenerate sites in two out of three protein sequences is negatively correlated with generation time. In addition, there is a relationship between the rate of molecular evolution and body size for two nuclear-encoded sequences. No evidence is found for an effect of metabolic rate on rate of sequence evolution. Possible causes of variation in substitution rate between species are discussed.

Age Factors↗

Predicting genome-wide functional constraints with GPN-Star.

Genomic language models have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences1. However, standard genomic language models adapted from natural language processing often require large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks2-4. Here we introduce a genomic pretrained network with species tree and alignment representations (GPN-Star), which is a biologically grounded genomic language model featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammal and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales show task-dependent advantages of modelling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms previous methods in prioritizing pathogenic and fine-mapped genome-wide association study variants, yields strong enrichments of complex trait heritability and improves power in rare variant association testing5. Extending beyond humans, we train GPN-Star for five model organisms-Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana-demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful and flexible tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article↗

Broad-spectrum resistance to Bacillus thuringiensis toxins in Heliothis virescens.

Evolution of pest resistance to insecticidal proteins produced by Bacillus thuringiensis (Bt) would decrease our ability to control agricultural pests with genetically engineered crops designed to express genes coding for these proteins. Previous genetic and biochemical analyses of insect strains with resistance to Bt toxins indicate that (i) resistance is restricted to single groups of related Bt toxins, (ii) decreased toxin sensitivity is associated with changes in Bt-toxin binding to sites in brush-border membrane vesicles of the larval midgut, and (iii) resistance is inherited as a partially or fully recessive trait. If these three characteristics were common to all resistant insects, specific crop-variety deployment strategies could significantly diminish problems associated with resistance in field populations of pests. We present data on Bt-toxin resistance in Heliothis virescens, a major agricultural pest targeted for control with Bt-toxin-producing crops. A laboratory strain of H. virescens developed resistance in response to selection with the Bt toxin CryIA(c). In contrast to other cases of Bt-toxin resistance, this H. virescens strain exhibits cross-resistance to Bt toxins that differ significantly in structure and activity. Furthermore, the resistance in this strain is not accompanied by significant changes in toxin binding, and resistance is inherited as an additive trait when larvae are treated with high doses of CryIA(c) toxin. These findings have important implications for Bt-toxin-based pest control.

Journal Article↗

Predicting functional constraints across evolutionary timescales with phylogeny-informed genomic language models.

Genomic language models (gLMs) have emerged as a powerful approach for learning genome-wide functional constraints directly from DNA sequences. However, standard gLMs adapted from natural language processing often require extremely large model sizes and computational resources, yet still fall short of classical evolutionary models in predictive tasks. Here, we introduce GPN-Star (Genomic Pretrained Network with Species Tree and Alignment Representation), a biologically grounded gLM featuring a phylogeny-aware architecture that leverages whole-genome alignments and species trees to model evolutionary relationships explicitly. Trained on alignments spanning vertebrate, mammalian, and primate evolutionary timescales, GPN-Star achieves state-of-the-art performance across a wide range of variant effect prediction tasks in both coding and non-coding regions of the human genome. Analyses across timescales reveal task-dependent advantages of modeling more recent versus deeper evolution. To demonstrate its potential to advance human genetics, we show that GPN-Star substantially outperforms prior methods in prioritizing pathogenic and fine-mapped GWAS variants; yields unprecedented enrichments of complex trait heritability; and improves power in rare variant association testing. Extending beyond humans, we train GPN-Star for five model organisms - Mus musculus, Gallus gallus, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana - demonstrating the robustness and generalizability of the framework. Taken together, these results position GPN-Star as a scalable, powerful, and flexible new tool for genome interpretation, well suited to leverage the growing abundance of comparative genomics data.

Journal Article↗