PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Short sequence repeats in microbial pathogenesis and evolution.

Repetitive DNA is ubiquitous in microbial genomes. Different classes of short sequence repeats (SSRs) have been identified and demonstrated to be generally heterogeneous in a locus-dependent manner, reflected in variation in the number of repeat units present at a given genomic site or by sequence heterogeneity among individual units. Both types of variability can be used to assess intra-species genetic diversity. Repeat variability often affects the coding potential of the region in which the repetitive element is located. This implies that determination of the primary structure of variable numbers of tandem repeats can be used for epidemiological identification purposes, and also for the analysis of gene function. Precise assessment of SSR structure can also generate insight into the regulation of gene expression. Together, DNA repeat analysis in microbial species provides information on both functional and evolutionary aspects of genetic diversity among microbial isolates.

Candida albicans↗

Recombination patterns in aphthoviruses mirror those found in other picornaviruses.

Foot-and-mouth disease virus (FMDV) is thought to evolve largely through genetic drift driven by the inherently error-prone nature of its RNA polymerase. There is, however, increasing evidence that recombination is an important mechanism in the evolution of these and other related picornoviruses. Here, we use an extensive set of recombination detection methods to identify 86 unique potential recombination events among 125 publicly available FMDV complete genome sequences. The large number of events detected between members of different serotypes suggests that horizontal flow of sequences among the serotypes is relatively common and does not incur severe fitness costs. Interestingly, the distribution of recombination breakpoints was found to be largely nonrandom. Whereas there are clear breakpoint cold spots within the structural genes, two statistically significant hot spots precisely separate these from the nonstructural genes. Very similar breakpoint distributions were found for other picornovirus species in the genera Enterovirus and Teschovirus. Our results suggest that genome regions encoding the structural proteins of both FMDV and other picornaviruses are functionally interchangeable modules, supporting recent proposals that the structural and nonstructural coding regions of the picornaviruses are evolving largely independently of one another.

Aphthovirus↗

Biochemical and immunological properties of gag genecoded structural proteins of endogenous tyep C RNA tumor viruses of diverse mammalian species.

The major nonglycosylated structure proteins of mammalian type C RNA tumor viruses are synthesized in the form of a high molecular weight precursor coded for by a viral gene disignated "gag". Previous genetic analysis of a prototype virus isolated of mouse origin has led to a determination of the internal arragnement of the regions within the gag gene coding for individual structural proteins. In the present study, the biochemical properties of structural proteins of type C virus isolates of additional mammalian species were analyzed. The results obtained indicate that the biochemical properties of immunologically cross-reactive proteins have been highly conserved throughout the evolution of this group of viruses. Moreover, these findings provide a means of mapping the gag genes of a broad range of mammalian type C viruses. In view of the results obtained, a new nomenclature system for type C viral gag gene-coded translational products is proposed.

Animals↗

Gene algebra from a genetic code algebraic structure.

By considering two important factors involved in the codon-anticodon interactions, the hydrogen bond number and the chemical type of bases, a codon array of the genetic code table as an increasing code scale of interaction energies of amino acids in proteins was obtained. Next, in order to consecutively obtain all codons from the codon AAC, a sum operation has been introduced in the set of codons. The group obtained over the set of codons is isomorphic to the group (Z(64), +) of the integer module 64. On the Z(64)-algebra of the set of 64(N) codon sequences of length N, gene mutations are described by means of endomorphisms f:(Z(64))(N)-->(Z(64))(N). Endomorphisms and automorphisms helped us describe the gene mutation pathways. For instance, 77.7% mutations in 749 HIV protease gene sequences correspond to unique diagonal endomorphisms of the wild type strain HXB2. In particular, most of the reported mutations that confer drug resistance to the HIV protease gene correspond to diagonal automorphisms of the wild type. What is more, in the human beta-globin gene a similar situation appears where most of the single codon mutations correspond to automorphisms. Hence, in the analyses of molecular evolution process on the DNA sequence set of length N, the Z(64)-algebra will help us explain the quantitative relationships between genes.

Anticodon↗

Fungal evolution: the case of the vanishing mitochondrion.

Mitochondria, the energy-producing organelles of the eukaryotic cell, are derived from an ancient endosymbiotic alpha-Proteobacterium. These organelles contain their own genetic system, a remnant of the endosymbiont's genome, which encodes only a fraction of the mitochondrial proteome. The majority of mitochondrial proteins are translated from nuclear genes and are imported into mitochondria. Recent studies of phylogenetically diverse representatives of Fungi reveal that their mitochondrial DNAs are among the most highly derived, encoding only a limited set of genes. Much of the reduction in the coding content of the mitochondrial genome probably occurred early in fungal evolution. Nevertheless, genome reduction is an ongoing process. Fungi in the chytridiomycete order Neocallimastigales and in the pathogenic Microsporidia have taken mitochondrial reduction to the extreme and have permanently lost a mitochondrial genome. These organisms have organelles derived from mitochondria that retain traces of their mitochondrial ancestry.

DNA, Mitochondrial↗

Primary structure of influenza virus genome regions coding for polypeptides from the major antigenic sites of H3 hemagglutinin.

Nucleotide sequences for some regions of the hemagglutinin (HA) gene of influenza virus A/Leningrad/385/80 (H3N2) were analyzed. A double-stranded complementary DNA was synthesized on the influenza genome RNA in the presence of synthetic oligodeoxyribonucleotides A GCAAAAGCAGG and A GTAGAAACAAG and inserted into the Pst I-site of the pBR322 plasmid through G-C-tailing. Nucleotide sequences were determined by a solid phase modification of the Maxam and Gilbert procedure. A comparison of our data with those for influenza virus A/Bangkok/1/79, which is the nearest sequenced prototype of the strain investigated, revealed only two base changes in the variable region of the HA gene. One of them was accompanied by a change in the coded amino acid (Asp53----Tyr) located in the antigenic site E of the HA glycoprotein. There also were some point mutations in the constant region of the gene. The results obtained are discussed in terms of the evolution of "Hong Kong" influenza viruses during their circulation.

Antigens, Viral↗

Molecular cloning of the Drosophila Fanconi anaemia gene FANCD2 cDNA.

Fanconi anaemia (FA) is a rare disease characterized by chromosome instability and cancer susceptibility. With the exception of FANCD2, none of the Fanconi anaemia genes are conserved in evolution, limiting the study of the Fanconi anaemia pathway in genetically tractable models. Here we report the cloning and sequencing of a Drosophila full length cDNA homologous to human FANCD2 (dmFANCD2) as a first step in using Drosophila in Fanconi anaemia research. dmFANCD2 is composed of 14 exons coding for a protein of 1478 aminoacids. Southern blot and in situ hybridization analysis indicated that dmFANCD2 is present at single copy in the Drosophila genome and maps at the chromosomal band 92-F3. Sequence and structural biocomputational analysis indicated that, although the aminoacidic sequence, and specially the N-terminus region, is not highly conserved between humans and flies (23% identity and 43% similarity), both proteins are of the same size, globular and compact, with several transmembrane helixes and related to nuclear membrane proteins. Interestingly, the human ATM phosphorylation site at S222 and the complex-dependent monoubiquitination site at K561 are highly conserved in Drosophila at positions S267 and K595, respectively. The same is true for other putative ATM sites and their aminoacidic environment and for two out of three aminoacid mutations associated with human pathology. These results suggest that the key FANCD2 features have been conserved during over 500 million years of divergent evolution, highlighting their biological importance.

Amino Acid Sequence↗

Larger genetic differences within africans than between Africans and Eurasians.

The worldwide pattern of single nucleotide polymorphism (SNP) variation is of great interest to human geneticists, population geneticists, and evolutionists, but remains incompletely understood. We studied the pattern in noncoding regions, because they are less affected by natural selection than are coding regions. Thus, it can reflect better the history of human evolution and can serve as a baseline for understanding the maintenance of SNPs in human populations. We sequenced 50 noncoding DNA segments each approximately 500 bp long in 10 Africans, 10 Europeans, and 10 Asians. An analysis of the data suggests that the sampling scheme is adequate for our purpose. The average nucleotide diversity (pi) for the 50 segments is only 0.061% +/- 0.010% among Asians and 0.064% +/- 0.011% among Europeans but almost twice as high (0.115% +/- 0.016%) among Africans. The African diversity estimate is even higher than that between Africans and Eurasians (0.096% +/- 0.012%). From available data for noncoding autosomal regions (total length = 47,038 bp) and X-linked regions (47,421 bp), we estimated the pi-values for autosomal regions to be 0.105, 0.070, 0.069, and 0.097% for Africans, Asians, Europeans, and between Africans and Eurasians, and the corresponding values for X-linked regions to be 0.088, 0.042, 0.053, and 0.082%. Thus, Africans differ from one another slightly more than from Eurasians, and the genetic diversity in Eurasians is largely a subset of that in Africans, supporting the out of Africa model of human evolution. Clearly, one must specify the geographic origins of the individuals sampled when studying pi or SNP density.

Africa↗

A population genetics study of single nucleotide polymorphisms in the interleukin 4 receptor alpha (IL4RA) gene.

Interleukin 4 (IL4) plays a critical role in T helper 2 (Th2) immune responses. Here we report a population genetics study of variation in the gene encoding the alpha-chain of the IL4 receptor (IL4RA) in three ethnic groups: African Americans, European Americans and East Asians. A 2941-bp region spanning exon 12 of IL4RA gene was sequenced in 12 individuals from each group. A total of 24 single nucleotide polymorphisms (SNPs) were identified in the combined sample. The genetic variation of the coding region of exon 12 is two to three times higher than in other reported genes. A significant departure from the expectation of evolutionary neutrality was observed, suggesting that natural selection may have influenced the evolution of this gene. We propose a model in which past selection by pathogens contributed to the increasing prevalence of atopic disorders in Western societies.

Black or African American↗

Transposon-mediated expansion and diversification of a family of ULP-like genes.

Transposons comprise a major component of eukaryotic genomes, yet it remains controversial whether they are merely genetic parasites or instead significant contributors to organismal function and evolution. In plants, thousands of DNA transposons were recently shown to contain duplicated cellular gene fragments, a process termed transduplication. Although transduplication is a potentially rich source of novel coding sequences, virtually all appear to be pseudogenes in rice. Here we report the results of a genome-wide survey of transduplication in Mutator-like elements (MULEs) in Arabidopsis thaliana, which shows that the phenomenon is generally similar to rice transduplication, with one important exception: KAONASHI (KI). A family of more than 97 potentially functional genes and apparent pseudogenes, evidently derived at least 15 MYA from a cellular small ubiquitin-like modifier-specific protease gene, KI is predominantly located in potentially autonomous non-terminal inverted repeat MULEs and has evolved under purifying selection to maintain a conserved peptidase domain. Similar to the associated transposase gene but unlike cellular genes, KI is targeted by small RNAs and silenced in most tissues but has elevated expression in pollen. In an Arabidopsis double mutant deficient in histone and DNA methylation with elevated KI expression compared to wild type, at least one KI-MULE is mobile. The existence of KI demonstrates that transduplicated genes can retain protein-coding capacity and evolve novel functions. However, in this case, our evidence suggests that the function of KI may be selfish rather than cellular.

Amino Acid Sequence↗

Structure and organization of Marchantia polymorpha chloroplast genome. I. Cloning and gene identification.

We have determined the complete nucleotide sequence of chloroplast DNA from a liverwort, Marchantia polymorpha, using a clone bank of chloroplast DNA fragments. The circular genome consists of 121,024 base-pairs and includes two large inverted repeats (IRA and IRB, each 10,058 base-pairs), a large single-copy region (LSC, 81,095 base-pairs), and a small single-copy region (SSC, 19,813 base-pairs). The nucleotide sequence was analysed with a computer to deduce the entire gene organization, assuming the universal genetic code and the presence of introns in the coding sequences. We detected 136 possible genes. 103 gene products of which are related to known stable RNA or protein molecules. Stable RNA genes for four species of ribosomal RNA and 32 species of tRNA were located, although one of the tRNA genes may be defective. Twenty genes encoding polypeptides involved in photosynthesis and electron transport were identified by comparison with known chloroplast genes. Twenty-five open reading frames (ORFs) show structural similarities to Escherichia coli RNA polymerase subunits, 19 ribosomal proteins and two related proteins. Seven ORFs are comparable with human mitochondrial NADH dehydrogenase genes. A computer-aided homology search predicted possible chloroplast homologues of bacterial proteins; two ORFs for bacterial 4Fe-4S-type ferredoxin, two for distinct subunits of a protein-dependent transport system, one ORF for a component of nitrogenase, and one for an antenna protein of a light-harvesting complex. The other 33 ORFs, consisting of 29 to 2136 codons, remain to be identified, but some of them seem to be conserved in evolution. Detailed information on gene identification is presented in the accompanying papers. We postulated that there were 22 introns in 20 genes (8 tRNA genes and 12 ORFs), which may be classified into the groups I and II found in fungal mitochondrial genes. The structural gene for ribosomal protein S12 is trans-split on the opposite DNA strand. The universal genetic code was confirmed by the substitution pattern of simultaneous codons, and by possible codon recognition of the chloroplast-encoded tRNA molecules, assuming no importation of tRNA molecules from the cytoplasm. The nucleotide residue A or T is preferred at the third position of the codons (G+C, 11.9%) and in intergenic spacers (G+C, 19.5%), resulting in an overall G+C content that is low (28.8%) throughout the liverwort chloroplast genome. Possible gene expression signals such as promoters and terminators for transcription, predicted locations of gene products, and DNA replicative origins are discussed.

Base Sequence↗

Structure and properties of the region of homology between plasmids pMB1 and ColE1.

Physical maps of the two independently isolated Escherichia coli plasmids, pMB1 and ColE1, were prepared with 13 restriction endonucleases and compared. A 5.1 kilobase continuous region covering 55% of pMB1 and 75% of colE1 was found to have similar, but non-identical, restriction maps. The differences in the maps of this region probably arose by localized mutational events rather than by major sequence rearrangements. The F-factor was found to mobilize pMB1 efficiently for conjugal transfer. A region on pMB1 required for its F-mediated transfer was mapped. Results of our study combined with results of other investigators suggest that pMB1 and ColE1 share functional properties such as colicin production, colicin immunity, mode of replication, and mobilization by the F-factor, and that the sequences required to code these functions are contained within the 5.1 kilobase homologous region.

Bacteriocin Plasmids↗

Tsbrowse: an interactive browser for ancestral recombination graphs.

SUMMARY: Ancestral recombination graphs (ARGs) represent the interwoven paths of genetic ancestry of a set of recombining sequences. The ability to capture the evolutionary history of samples makes ARGs valuable in a wide range of applications in population and statistical genetics. ARG-based approaches are increasingly becoming a part of genetic data analysis pipelines due to breakthroughs enabling ARG inference at biobank-scale. However, there is a lack of visualization tools, which are crucial for validating inferences and generating hypotheses. We present tsbrowse, an open-source, web-based Python application for the interactive visualization of the fundamental building blocks of ARGs, i.e. nodes, edges and mutations. We demonstrate the application of tsbrowse to various data sources and scenarios, and highlight its key features of browsability along the genome, user interactivity, and scalability to very large sample sizes. AVAILABILITY AND IMPLEMENTATION: Tsbrowse is installed as a Python package from PyPI (https://pypi.org/project/tsbrowse/), while a development version is maintained at https://github.com/tskit-dev/tsbrowse. Documentation is available at https://tskit.dev/tsbrowse/docs/. Source code is archived on Zenodo with DOI, https://doi.org/10.5281/zenodo.15683039.

Software↗

Clearcut: a fast implementation of relaxed neighbor joining.

SUMMARY: Clearcut is an open source implementation for the relaxed neighbor joining (RNJ) algorithm. While traditional neighbor joining (NJ) remains a popular method for distance-based phylogenetic tree reconstruction, it suffers from a O(N(3)) time complexity, where N represents the number of taxa in the input. Due to this steep asymptotic time complexity, NJ cannot reasonably handle very large datasets. In contrast, RNJ realizes a typical-case time complexity on the order of N(2)logN without any significant qualitative difference in output. RNJ is particularly useful when inferring a very large tree or a large number of trees. In addition, RNJ retains the desirable property that it will always reconstruct the true tree given a matrix of additive pairwise distances. Clearcut implements RNJ as a C program, which takes either a set of aligned sequences or a pre-computed distance matrix as input and produces a phylogenetic tree. Alternatively, Clearcut can reconstruct phylogenies using an extremely fast standard NJ implementation. AVAILABILITY: Clearcut source code is available for download at: http://bioinformatics.hungry.com/clearcut

Algorithms↗

Does heterochromatin variation potentiate speciation? A study in Nesokia.

An increasing incidence of sex-chromosome variation in constitutive heterochromatin, including individuals with mosaic genotypes, has been observed in a single natural population of Nesokia indica, the Indian mole rat. Variations in the heterochromatic areas of the X chromosome are largely due to deletions at R-band-positive regions corresponding to folate-sensitive fragile sites. All individuals with either a pre- or post-zygotic loss or gain of sex-chromosome heterochromatin have so far proved to be infertile. Whether such F1 sterility is due to abnormal gonadal development, gametic incompetence, or other factors is not clear. More important is the indication that the constitutive heterochromatin of this species may contain coding DNA sequences with putative regulatory functions.

Animals↗

[The role of conserved sequences in the regulatory elements of the Antp-like homeobox-containing genes of vertebrates].

By the present time the homeobox genes have been found in the representatives of the main invertebrate and vertebrate taxa. It has been demonstrated that these genes play the key role in the space and time genome expression orchestration in ontogenesis. The autoregulatory and cross-regulatory functional interactions integrate the homeobox genes into the gene networks. We found a correlation in variability of the coding and regulatory regions for vertebrate homeobox genes. The phylogenetic relations of structure and regulatory elements involved into the cross- and autoregulatory connections have been investigated in detail. The comprehensive phylogenetic analysis of the promoter region for these genes compared to results of such analysis of their homeoboxes has revealed two opposed tendencies in evolution of the regulatory elements of genes. The first trend is conservation of many regulatory elements in evolution of vertebrate homeobox genes and the second one is high variability of other non-coding gene regions.

Amino Acid Sequence↗

A multiple-site-specific heteroduplex tracking assay as a tool for the study of viral population dynamics.

Rapidly evolving entities, such as viruses, can undergo complex genetic changes in the face of strong selective pressure. We have developed a modified heteroduplex tracking assay (HTA) capable of detecting the presence of single, specific mutations or sets of linked mutations. The initial application of this approach, termed multiple-site-specific (MSS) HTA, was directed toward the detection of mutations in the HIV-1 pro gene at positions 46, 48, 54, 82, 84, and 90, which are associated with resistance to multiple protease inhibitors. We demonstrate that MSS HTA is sensitive and largely specific to all targeted mutations. The assay allows the accurate and reproducible quantitation of viral subpopulations comprising 3% or more of the total population. Furthermore, we used MSS HTA in longitudinal studies of pro gene evolution in vitro and in vivo. In the examples shown here, populations turned over rapidly and more than one population was present frequently. To demonstrate the versatility of MSS HTA, we also constructed a probe sensitive to changes at positions 181 and 184 of the RT coding domain. Changes at these positions are involved in resistance to nevirapine and 2',3'-dideoxy-3'-thiacytidine (3TC), respectively. This assay easily detected the evolution of resistance to 3TC. MSS HTA provides a rapid and sensitive approach for detecting the presence of and quantifying complex mixtures of distinct genotypes, including genetically linked mutations, and, as one example, represents a useful tool for following the evolution of drug resistance during failure of HIV-1 antiviral therapy.

Antiviral Agents↗

Pattern of nucleotide substitution and the extent of purifying selection in retroviruses.

The patterns of point mutation and nucleotide substitution are inferred from nucleotide differences in three coding and two noncoding regions of retroviral genomes. Evidence is presented in favor of the view that the majority of mutations accumulate at the reverse transcription stage. Purifying selection is apparently very weak at the amino acid level, and almost nonexistent between synonymous codons. The pattern of purifying selection obeys the rules previously established in vertebrates [Gojobori T, Li W-H, Graur D (1982) J Mol Evol 18:360-369]; i.e., the magnitude of purifying selection at the amino acid level is negatively correlated with Grantham's [Grantham R (1974) Science 185: 862-864] chemical distances between the amino acids interchanged. We refute Modiano et al.'s [Modiano G, Battistuzzi G, Motulsky AG (1981) Proc Natl Acad Sci USA 78:1110-1114] hypothesis, according to which the pattern of mutation is preadapted to buffer against deleterious mutations. On the contrary, the pattern of mutation reduces the level of conservativeness from that imposed on the amino acid substitution pattern by the structure of the genetic code. The extraordinarily high rate of nucleotide substitution in retroviruses in comparison with that in other organisms is apparently caused by an extremely high rate of mutation coupled with a lack of stringent purifying selection at both the codon and the amino acid levels.

Animals↗