PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Evolution of the major histocompatibility complex.

The major histocompatibility complex is a group of closely linked loci that code for molecules used by T-lymphocytes as context for the recognition of antigens. The loci fall into two classes: I, coding for molecules used as context by cytotoxic T-lymphocytes and II, used as context by helper and other regulatory T-cells. The Mhc is present in all mammals and perhaps all vertebrates. Some of the Mhc loci are highly polymorphic, while others are not. This article will summarize what is known about the genetic organization of the Mhc in different species and will discuss the selection pressures acting on the individual loci and the tempo and the mechanisms of their evolution.

Animals↗

Adaptive evolution of non-coding DNA in Drosophila.

A large fraction of eukaryotic genomes consists of DNA that is not translated into protein sequence, and little is known about its functional significance. Here I show that several classes of non-coding DNA in Drosophila are evolving considerably slower than synonymous sites, and yet show an excess of between-species divergence relative to polymorphism when compared with synonymous sites. The former is a hallmark of selective constraint, but the latter is a signature of adaptive evolution, resembling general patterns of protein evolution in Drosophila. I estimate that about 40-70% of nucleotides in intergenic regions, untranslated portions of mature mRNAs (UTRs) and most intronic DNA are evolutionarily constrained relative to synonymous sites. However, I also use an extension to the McDonald-Kreitman test to show that a substantial fraction of the nucleotide divergence in these regions was driven to fixation by positive selection (about 20% for most intronic and intergenic DNA, and 60% for UTRs). On the basis of these observations, I suggest that a large fraction of the non-translated genome is functionally important and subject to both purifying selection and adaptive evolution. These results imply that, although positive selection is clearly an important facet of protein evolution, adaptive changes to non-coding DNA might have been considerably more common in the evolution of D. melanogaster.

Adaptation, Physiological↗

Economy, speed and size matter: evolutionary forces driving nuclear genome miniaturization and expansion.

BACKGROUND: Nuclear genome size varies 300 000-fold, whereas transcriptome size varies merely 17-fold. In the largest genomes nearly all DNA is non-genic secondary DNA, mostly intergenic but also within introns. There is now compelling evidence that secondary DNA is functional, i.e. positively selected by organismal selection, not the purely neutral or 'selfish' outcome of mutation pressure. The skeletal DNA theory argued that nuclear volumes are genetically determined primarily by nuclear DNA amounts, modulated somewhat by genes affecting the degree of DNA packing or unfolding; the huge spread of nuclear genome sizes is the necessary consequence of the origin of the nuclear envelope and the nucleation of its assembly by DNA, plus the adaptively significant 300 000-fold range of cell volumes and selection for balanced growth by optimizing karyoplasmic volume ratios (essentially invariant with cell volume in growing/multiplying cells). This simple explanation of the C-value paradox is refined here in the light of new insights into the nature of heterochromatin and the nuclear lamina, the genetic control of cell volume, and large-scale eukaryote phylogeny, placing special emphasis on protist test cases of the basic principles of nuclear genome size evolution. GENOME MINIATURIZATION: and Expansion Intracellular parasites (e.g. Plasmodium, microsporidia) dwarfed their genomes by gene loss and eliminating virtually all secondary DNA. The primary driving forces for genome reduction are metabolic and spatial economy and cell multiplication speed. Most extreme nuclear shrinkage yielded genomes as tiny as 0.38 Mb (making the nuclear genome size range effectively 1.8 million-fold!) in some minute enslaved nuclei (nucleomorphs) of cryptomonads and chlorarachneans, chimaeric cells that also retain a separate normal large nucleus. The latter shows typical correlation between genome size and cell volume, but nucleomorphs do not despite co-existing in the same cell for >500 My. Thus mutation pressure does not inexorably increase genome size; selection can eliminate essentially all non-coding DNA if need be. Nucleomorphs and microsporidia even reduced gene size. Expansion of secondary DNA in the main nucleus, and in large-celled eukaryotes generally, must be positively selected for function. Ciliate nuclear dimorphism provides a key test that refutes the selfish DNA and strongly supports the skeletal DNA/karyoplasmic ratio interpretation of genome size evolution. GENETIC CONTROL OF CELL VOLUME IS MULTIGENIC: The quantitatively proportional correlation between genome size and cell size cannot be explained by purely mutational theories, as eukaryote cell volumes are causally determined by cell cycle control genes, not by DNA amounts.

Animals↗

Analytical expression of the purine/pyrimidine autocorrelation function after and before random mutations.

The mutation process is a classical evolutionary genetic process. The type of mutations studied here is the random substitutions of a purine base R (adenine or guanine) by a pyrimidine base Y (cytosine or thymine) and reciprocally (transversions). The analytical expressions derived allow us to analyze in genes the occurrence probabilities of motifs and d-motifs (two motifs separated by any d bases) on the R/Y alphabet under transversions. These motif probabilities can be obtained after transversions (in the evolutionary sense; from the past to the present) and, unexpectedly, also before transversions (after back transversions, in the inverse evolutionary sense, from the present to the past). This theoretical part in Section 2 is a first generalization of a particular formula recently derived. The application in Section 3 is based on the analytical expression giving the autocorrelation function (the d-motif probabilities) before transversions. It allows us to study primitive genes from actual genes. This approach solves a biological problem. The protein coding genes of chloroplasts and mitochondria have a preferential occurrence of the 6-motif YRY(N)6YRY (maximum of the autocorrelation function for d = 6, N = R or Y) with a periodicity modulo 3. The YRY(N)6YRY preferential occurrence without the periodicity modulo 3 is also observed in the RNA coding genes (ribosomal, transfer, and small nuclear RNA genes) and in the noncoding genes (introns and 5' regions of eukaryotic nuclei). However, there are two exceptions to this YRY(N)6YRY rule: the protein coding genes of eukaryotic nuclei, and prokaryotes, where YRY(N)6YRY has the second highest value after YRY(N)0YRY (YRYYRY) with a periodicity modulo 3. When we go backward in time with the analytical expression, the protein coding genes of both eukaryotic nuclei and prokaryotes retrieve the YRY(N)6YRY preferential occurrence with a periodicity modulo 3 after 0.2 back transversions per base. In other words, the actual protein coding genes of chloroplasts and mitochondria are similar to the primitive protein coding genes of eukaryotic nuclei and prokaryotes. On the other hand, this application represents the first result concerning the mutation process in the model of DNA sequence evolution we recently proposed. According to this model, the actual genes on the R/Y alphabet derive from two successive evolutionary genetic processes: an independent mixing of a few nonrandom types of oligonucleotides leading to genes called primitive followed by a mutation process in these primitive genes.(ABSTRACT TRUNCATED AT 400 WORDS)

Base Sequence↗

Genetic evidence of a strong functional constraint of neurotrypsin during primate evolution.

Neurotrypsin is one of the extra-cellular serine proteases that are predominantly expressed in the brain and involved in neuronal development and function. Mutations in humans are associated with autosomal recessive non-syndromic mental retardation (MR). We studied the molecular evolution of neurotrypsin by sequencing the coding region of neurotrypsin in 11 representative non-human primate species covering great apes, lesser apes, Old World monkeys and New World monkeys. Our results demonstrated a strong functional constraint of neurotrypsin that was caused by strong purifying selection during primate evolution, an implication of an essential functional role of neurotrypsin in primate cognition. Further analysis indicated that the purifying selection was in fact acting on the SRCR domains of neurotrypsin, which mediate the binding activity of neurotrypsin to cell surface or extra-cellular proteins. In addition, by comparing primates with three other mammalian orders, we demonstrated that the absence of the first copy of the SRCR domain (exon 2 and 3) in mouse and rat was due to the deletion of this segment in the murine lineage.

Amino Acid Sequence↗

Evolution of a simian immunodeficiency virus pathogen.

Analysis of disease induction by simian immunodeficiency viruses (SIV) in macaques was initially hampered by a lack of molecularly defined pathogenic strains. The first molecularly cloned SIV strains inoculated into macaques, SIVmacBK28 and SIVmacBK44 (hereafter designated BK28 and BK44, respectively), were cases in point, since they failed to induce disease within 1 year postinoculation in any inoculated animal. Here we report the natural history of infection with BK28 and BK44 in inoculated rhesus macaques and efforts to increase the pathogenicity of BK28 through genetic manipulation and in vivo passage. BK44 infection resulted in no disease in four animals infected for more than 7 years, whereas BK28 induced disease in less than half of animals monitored for up to 7 years. Elongation of the BK28 transmembrane protein (TM) coding sequence truncated by prior passage in human cells marginally increased pathogenicity, with two of four animals dying in the third year and one dying in the seventh year of infection. Modification of the BK28 long terminal repeat to include four consensus nuclear factor SP1 and two consensus NF-kappaB binding sites enhanced early virus replication without augmenting pathogenicity. In contrast, in vivo passage of BK28 from the first animal to die from immunodeficiency disease (1.5 years after infection) resulted in a consistently pathogenic strain and a 50% survival time of about 1.3 years, thus corresponding to one of the most pathogenic SIV strains identified to date. To determine whether the diverse viral quasispecies that evolved during in vivo passage was required for pathogenicity or whether a more virulent virus variant had evolved, we generated a molecular clone composed of the 3' half of the viral genome derived from the in vivo-passaged virus (H824) fused with the 5' half of the BK28 genome. Kinetics of disease induction with this cloned virus (BK28/H824) were similar to those with the in vivo-passaged virus, with four of five animals surviving less than 1.7 years. Thus, evolution of variants with enhanced pathogenicity can account for the increased pathogenicity of this SIV strain. The genetic changes responsible for this virulent transformation included at most 59 point mutations and 3 length-change mutations. The critical mutations were likely to have been multiple and dispersed, including elongation of the TM and Nef coding sequences; changes in RNA splice donor and acceptor sites, TATA box sites, and Sp1 sites; multiple changes in the V2 region of SU, including a consensus neutralization epitope; and five new N-linked glycosylation sites in SU.

Amino Acid Sequence↗

A novel RNA splicing-mediated gene silencing mechanism potential for genome evolution.

Over 90% of the human genome consists of non-protein-coding regions. Introns constitute most of the non-coding regions located in precursor messenger RNAs (pre-mRNAs). During pre-mRNA maturation, the introns are excised out of mRNA and thought to be completely digested prior to translation. If the introns were merely metabolic "leavings," why would the genome hold such a large amount of extraneous genetic materials? Here we show a novel posttranscriptional gene silencing system identified within mammalian introns. By packaging human spliceosome-recognition sites along with an exonic insert into an artificial intron, we observed that the splicing and processing of such an exon-containing intron in either sense or antisense conformation produced equivalent gene silencing effects, while a palindromic hairpin insert containing both sense and antisense strands resulted in synergistic effects. These findings may explain how cells respond to the presence of transgenic introns that are homologous to pre-existing exons during genomic evolution.

Animals↗

Origins of translation: the hypothesis of permanently attached adaptors.

A mechanism for prebiotic translation is proposed in which primeval transfer-RNA (adaptors) are assumed to be permanently associated with messenger nucleic acid molecules. Residual 'fossil' evidences are found to be present within the base sequences of contemporary tRNAs, suggesting the existence of inter-primal-tRNA interactions necessary for the mechanism. The structure of proposed primal-tRNA is such that it can not only choose its own amino acid in the absence of aminoacyl synthetase, but can also associate nonspecifically with adjacent primal-tRNA molecules attached to the neighbouring codons. Such associations can give rise, through cooperative binding between message and adaptors to the 'static template surfaces' which can direct translation of nucleotide sequences into those of amino acids. The origins of ribosomes and contemporary genetic code are suggested by this hypothesis. Proposed structures and processes are thermodynamically compatible. The approximate date of occurrence of the proposed system is calculated, which is consistent with the period of occurrence of the earliest organism with ribosomes.

Amino Acids↗

The complete structure of the rat VIP gene.

Vasoactive intestinal polypeptide (VIP) is a regulatory neuropeptide/neurotransmitter of 28 amino acids involved in a wide variety of physiological functions. Using synthetic oligodeoxynucleotide probes related to the rat VIP-cDNA, we have isolated and characterized the gene encoding the rat pre-pro VIP/PHI-27 and compared it to the human VIP gene. The rat VIP gene spanned 7400 base pairs, and contained 7 exons interrupted by 6 introns. 100% identity was found between the gene exons and the cDNA sequence. Differences in sizes of introns 2, 4 and 5 (shorter in the rat gene) are the reason for the shorter rat gene compared with the human gene of 8837 base pairs. Comparison of the genes in the two species showed a high homology in the exon sequences, 80-90% in exons 2, 4, 5, 6 and 30-50% in exons 1 and 7. In addition, the exon-intron junctions shared high identity between the genes. The rat untranslated exon 1 had little homology (30%) with human exon 1 and was 13 base pairs shorter. Interestingly, the 160 base pairs at the 5'-flanking region upstream of the cap-site share more than 75% identity between the two genes, including the exact position of TATA-boxes in positions -28, -145, -155, a cAMP-responsive element in position -80 and a CAAT sequence in position -127. The conservation of the 5'-flanking region of the VIP gene in parallel with the conservation of its coding exons emphasize the importance of these sequences during evolution.

Amino Acid Sequence↗

Sharing Ia antigens between species. III. Ia specificities shared between mice and human beings.

Certain mouse alloantisera have been found to detect immunologic cross-reactions between human and murine Ia antigens. Almost every anti-Iaa, -Iak and -Iad serum tested exhibited such cross-reactions. Sera prepared against the products of limited segments of the mouse I region and tested on human B cells revealed that anti-I-E/Ck cross-reactions were more readily detectable than anti-I-A, B, Jk cross-reactions. Most of the mouse alloantisera were cytotoxic to bells from almost every individual tested, although a few sera exhibited more restricted patterns of lysis, permitting limited segregation analyses. The cytotoxicity of these mouse alloantisera in family studies was consistent with HLA linkage of the genes responsible for the cross-reacting Ia determinants. Immunochemical analysis on radiolabelled detergent lysates of human lymphocytes indicated that the sera reacted with molecules of 34,000 and 28,000 daltons. Thus, by cellular distribution, HLA association, and immunochemical criteria, these mouse alloantisera detect human Ia antigens. These cross-reactive sera should be of practical value for the detection of human Ia antigens and may also have theoretical implications for the evolution of genes coding for Ia antigens.

Animals↗

Natural selection and molecular evolution in primate PAX9 gene, a major determinant of tooth development.

Large differences in relation to dental size, number, and morphology among and within modern human populations and between modern humans and other primate species have been observed. Molecular studies have demonstrated that tooth development is under strict genetic control, but, the genetic basis of primate tooth variation remains unknown. The PAX9 gene, which codes for a paired domain-containing transcription factor that plays an essential role in the development of mammal dentition, has been associated with selective tooth agenesis in humans and mice, which mainly involves the posterior teeth. To determine whether this gene is polymorphic in humans, we sequenced approximately 2.1 kb of the entire four-exon region (exons 1, 2, 3 and 4; 1,026 bp) and exon-intron (1.1 kb) boundaries of 86 individuals sampled from Asian, European, and Native American populations. We provided evidence that human PAX9 polymorphisms are limited to exon 3 only and furnished details about the distribution of a mutation there in 350 Polish subjects. To investigate the pattern of selective pressure on exon 3, we sequenced ortholog regions of this exon in four species of New World monkeys and one gorilla. In addition, orthologous sequences of PAX9 available in public databases were also analyzed. Although several differences were identified between humans and other species, our findings support the view that strong purifying selection is acting on PAX9. New World and Old World primate lineages may, however, have different degrees of restriction for changes in this DNA region.

Amino Acid Sequence↗

An evolutionary model for the origin of non-randomness, long-range order and fractality in the genome.

We present a model for genome evolution, comprising biologically plausible events such as transpositions inside the genome and insertions of exogenous sequences. This model attempts to formulate a minimal proposition accounting for key statistical properties of genomes, avoiding, as far as possible, unsupportable hypotheses for the remote evolutionary past. The statistical properties that are observed in genomic sequences and are reproduced by the proposed model are: (i) deviations from randomness at different length scales, measured by suitable algorithms, (ii) a special form of size distribution (power law distribution) characterising different levels of genome organisation in the non-coding, and (iii) extensive resemblance in the alternation of coding and non-coding regions at several length scales (self-similarity) in long genomic sequences of higher eukaryotes.

DNA↗

Rose: generating sequence families.

MOTIVATION: We present a new probabilistic model of the evolution of RNA-, DNA-, or protein-like sequences and a software tool, Rose, that implements this model. Guided by an evolutionary tree, a family of related sequences is created from a common ancestor sequence by insertion, deletion and substitution of characters. During this artificial evolutionary process, the 'true' history is logged and the 'correct' multiple sequence alignment is created simultaneously. The model also allows for varying rates of mutation within the sequences, making it possible to establish so-called sequence motifs. RESULTS: The data created by Rose are suitable for the evaluation of methods in multiple sequence alignment computation and the prediction of phylogenetic relationships. It can also be useful when teaching courses in or developing models of sequence evolution and in the study of evolutionary processes. AVAILABILITY: Rose is available on the Bielefeld Bioinformatics WebServer under the following URL: http://bibiserv.TechFak.Uni-Bielefeld.DE/rose/ The source code is available upon request. CONTACT: folker@TechFak.Uni-Bielefeld.DE

Algorithms↗

GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence model.

MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on &#x223c;149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.

Phylogeny↗

A genome-wide study of dual coding regions in human alternatively spliced genes.

Alternative splicing is a major mechanism for gene product regulation in many multicellular organisms. By using different exon combinations, some coding regions can encode amino acids in multiple reading frames in different transcripts. Here we performed a systematic search through a set of high-quality human transcripts and show that approximately 7% of alternatively spliced genes contain dual (multiple) coding regions. By using a conservative criterion, we found that in these regions most secondary reading frames evolved recently in mammals, and a significant proportion of them may be specific to primates. Based on the presence of in-frame stop codons in orthologous sequences in other animals, we further classified ancestral and derived reading frames in these regions. Our results indicated that ancestral reading frames are usually under stronger selection than are derived reading frames. Ancestral reading frames mainly influence the coding properties of these dual coding regions. Compared with coding regions of the whole genome, ancestral reading frames largely maintain similar nucleotide composition at each codon position and amino acid usage, while derived reading frames are significantly different. Our results also indicated that prior to acquisition of a new reading frame, the suppression of in-frame stop codons in the ancestral state is mainly achieved by one-step transition substitutions at the first or second codon position. Finally, the selective forces imposed on these dual coding regions will also be discussed.

Alternative Splicing↗

The apolipoprotein multigene family: structure, expression, evolution, and molecular genetics.

The plasma apolipoproteins can be classified into two subgroups: the soluble apolipoproteins including apolipoprotein (apo) A-I, A-II, A-IV, C-I, C-II, C-III, and E, and the apoBs including apoB-100 and apoB-48. The soluble apolipoproteins have very similar genomic structures, each having a total of three introns at the same locations; apoA-IV is an exception in that it has lost its first intron. Using the exon/intron junctions as reference points, we can obtain an alignment of the coding regions of all the soluble apolipoprotein genes. The mature peptide regions of the genes are almost completely made up of tandem repeats of 11 codons. The part of mature peptide region encoded by exon 3 contains a common block of 33 codons, whereas the part encoded by exon 4 contains a much more variable number of internal repeats of 11 codons. On the basis of the degree of homology of the various sequences, and the pattern of the internal repeats in these genes, an evolutionary tree has been proposed for the soluble apolipoprotein genes. ApoB-100 differs considerably from the soluble apolipoproteins. It is the largest apolipoprotein containing 4536 amino acid residues. Two types of internal repeats are identified in apoB-100: amphipathic alpha-helical repeats and proline-containing repeats with high beta-sheet content. The apoB gene contains 29 exons and 28 introns. Its evolutionary relationship to the soluble apolipoprotein genes is unclear. The 3' end of the apoB gene contains a region of variable number of tandem 12-16-base pair repeats.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Experimental studies on the origin of the genetic code and the process of protein synthesis: a review update.

This article is an update of our earlier review (Lacey and Mullins, 1983) in this journal on the origin of the genetic code and the process of protein synthesis. It is our intent to discuss only experimental evidence published since then although there is the necessity to mention the old enough to place the new in context. We do not include theoretical nor hypothetical treatments of the code or protein synthesis. Relevant data regarding the evolution of tRNAs and the recognition of tRNAs by aminoacyl-tRNA-synthetases are discussed. Our present belief is that the code arose based on a core of early assignments which were made on a physico-chemical and anticodonic basis and this was expanded with new assignments later. These late assignments do not necessarily show an amino acid-anticodon relatedness. In spite of the fact that most data suggest a code origin based on amino acid-anticodon relationships, some new data suggesting preferential binding of Arg to its codons are discussed. While information regarding coding is not increasing very rapidly, information regarding the basic chemistry of the process of protein synthesis has increased significantly, principally relating to aminoacylation of mono- and polyribonucleotides. Included in those studies are several which show stereoselective reactions of L-amino acids with nucleotides having D-sugars. Hydrophobic interactions definitely play a role in the preferences which have been observed.

5'-Nucleotidase↗

Evolution of genomic diversity and sex at extreme environments: fungal life under hypersaline Dead Sea stress.

We have found that genomic diversity is generally positively correlated with abiotic and biotic stress levels (1-3). However, beyond a high-threshold level of stress, the diversity declines to a few adapted genotypes. The Dead Sea is the harshest planetary hypersaline environment (340 g.liter-1 total dissolved salts, approximately 10 times sea water). Hence, the Dead Sea is an excellent natural laboratory for testing the "rise and fall" pattern of genetic diversity with stress proposed in this article. Here, we examined genomic diversity of the ascomycete fungus Aspergillus versicolor from saline, nonsaline, and hypersaline Dead Sea environments. We screened the coding and noncoding genomes of A. versicolor isolates by using >600 AFLP (amplified fragment length polymorphism) markers (equal to loci). Genomic diversity was positively correlated with stress, culminating in the Dead Sea surface but dropped drastically in 50- to 280-m-deep seawater. The genomic diversity pattern paralleled the pattern of sexual reproduction of fungal species across the same southward gradient of increasing stress in Israel. This parallel may suggest that diversity and sex are intertwined intimately according to the rise and fall pattern and adaptively selected by natural selection in fungal genome evolution. Future large-scale verification in micromycetes will define further the trajectories of diversity and sex in the rise and fall pattern.

Biological Evolution↗