PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genetic code evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

Phylogenetics of eggshell morphogenesis in Antheraea (lepidoptera: saturniidae): unique origin and repeated reduction of the aeropyle crown.

Integrated phylogenetic and developmental analyses should enhance our understanding of morphological evolution and thereby improve systematists' ability to utilize morphological characters, but case studies are few. The eggshell (chorion) of Lepidoptera (Insecta) has proven especially tractable experimentally for such analyses because its morphogenesis proceeds by extracellular assembly of proteins. This study focuses on a morphological novelty, the aeropyle crown, that arises at the end of choriogenesis in the wild silkmoth genus Antheraea. Aeropyle crowns are cylindrical projections, ending in prominent prongs, that surround the openings of breathing tubes (aeropyle channels) traversing the chorion. They occur over the entire egg surface in some species, are localized to a circumferential band in many others, and in some are missing entirely, thus exhibiting variation typical of discrete characters analyzed in morphological phylogenetics. Seeking an integrated developmental-phylogenetic view, we first survey aeropyle crown variation broadly across Antheraea and related genera. We then map these observations onto a robust phylogeny, based on three nuclear genes, to test the adequacy of character codings for aeropyle crown variation and to estimate the frequency and direction of change in those characters. Thirdly, we draw on previous studies of choriogenesis, supplemented by new data on gene expression, to hypothesize developmental-genetic bases for the inferred chorion character transformations. Aeropyle crowns are inferred to arise just once, in the ancestor of Antheraea, but to undergo four or more subsequent reductions without regain, a pattern consistent with Dollo's Law. Spatial distribution shows an analogous trend, though less clear-cut, toward reduction of coverage by aeropyle crowns. These trends suggest either that there is little or no natural selection on the details of the aeropyle crown structure or that evolution toward functional optima is ongoing, although no direct evidence exists for either. Genetic, biochemical, and microscopy studies point to at least two developmental changes underlying the origin of the aeropyle crown, namely, reinitiation of deposition of chorionic lamellae after the end of normal choriogenesis (i.e., heterochrony), and sharply increased production of underlying "filler" proteins that push the nascent final lamellae upward to form the crown (i.e., heteroposy). Identification of a unique putative cis-regulatory element shared by unrelated genes involved in aeropyle crown formation suggests a possible simple mechanism for repeated evolutionary reduction and spatial restriction of aeropyle crowns.

Animals↗

Application of comparative phylogenomics to study the evolution of Yersinia enterocolitica and to identify genetic differences relating to pathogenicity.

Yersinia enterocolitica, an important cause of human gastroenteritis generally caused by the consumption of livestock, has traditionally been categorized into three groups with respect to pathogenicity, i.e., nonpathogenic (biotype 1A), low pathogenicity (biotypes 2 to 5), and highly pathogenic (biotype 1B). However, genetic differences that explain variation in pathogenesis and whether different biotypes are associated with specific nonhuman hosts are largely unknown. In this study, we applied comparative phylogenomics (whole-genome comparisons of microbes with DNA microarrays combined with Bayesian phylogenies) to investigate a diverse collection of 94 strains of Y. enterocolitica consisting of 35 human, 35 pig, 15 sheep, and 9 cattle isolates from nonpathogenic, low-pathogenicity, and highly pathogenic biotypes. Analysis confirmed three distinct statistically supported clusters composed of a nonpathogenic clade, a low-pathogenicity clade, and a highly pathogenic clade. Genetic differences revealed 125 predicted coding sequences (CDSs) present in all highly pathogenic strains but absent from the other clades. These included several previously uncharacterized CDSs that may encode novel virulence determinants including a hemolysin, a metalloprotease, and a type III secretion effector protein. Additionally, 27 CDSs were identified which were present in all 47 low-pathogenicity strains and Y. enterocolitica 8081 but absent from all nonpathogenic 1A isolates. Analysis of the core gene set for Y. enterocolitica revealed that 20.8% of the genes were shared by all of the strains, confirming this species as highly heterogeneous, adding to the case for the existence of three subspecies of Y. enterocolitica. Further analysis revealed that Y. enterocolitica does not cluster according to source (host).

Animals↗

Sex chromosome evolution: molecular aspects of Y-chromosome degeneration in Drosophila.

Ancient Y-chromosomes of various organisms contain few active genes and an abundance of repetitive DNA. The neo-Y chromosome of Drosophila miranda is in transition from an ordinary autosome to a genetically inert Y-chromosome, while its homolog, the neo-X chromosome, is evolving partial dosage compensation. Here, I compare four large genomic regions located on the neo-sex chromosomes that contain a total of 12 homologous genes. In addition, I investigate the partial coding sequence for 56 more homologous gene pairs from the neo-sex chromosomes. Little modification has occurred on the neo-X chromosome, and genes are highly constrained at the protein level. In contrast, a diverse array of molecular changes is contributing to the observed degeneration of the neo-Y chromosome. In particular, the four large regions surveyed on the neo-Y chromosome harbor several transposable element insertions, large deletions, and a large structural rearrangement. About one-third of all neo-Y-linked genes are nonfunctional, containing either premature stop codons and/or frameshift mutations. Intact genes on the neo-Y are accumulating amino acid and unpreferred codon changes. In addition, both 5'- and 3'-flanking regions of genes and intron sequences are less constrained on the neo-Y relative to the neo-X. Despite heterogeneity in levels of dosage compensation along the neo-X chromosome of D. miranda, the neo-Y chromosome shows surprisingly uniform signs of degeneration.

Animals↗

Glycoprotein evolution of vesicular stomatitis virus New Jersey.

A T1 ribonuclease fingerprinting study of a large number of virus isolates had previously demonstrated that considerable genetic variability existed among natural isolates of the vesicular stomatitis virus (VSV) New Jersey (NJ) serotype [S.T. Nichol (1988) J. Virol. 62, 572-579]. Based on these results, 34 virus isolates were chosen as representing the extent of genetic diversity within the VSV NJ serotype. We report the entire glycoprotein (G) gene nucleotide sequence and the deduced amino acid sequence for each of these viruses. Up to 19.8% G gene sequence differences could be seen among NJ serotype isolates. Analysis of the distribution of nucleotide substitutions relative to nucleotide codon position revealed that third position changes were distributed randomly throughout the gene. Third base changes constituted 84% of the observed nucleotide substitutions and affected 89% of the third base positions located in the G gene. Only three short oligonucleotide stretches of complete sequence conservation were observed. The remaining nucleotide changes located in the first and second positions were not distributed randomly, indicating that most of the amino acids coded by the G gene cannot be altered without reducing the fitness of the VSV NJ serotype viruses. Despite these constraints, up to 8.5% amino acid differences were observed between virus isolates. These differences were located throughout the G protein including regions adjacent to defined major antibody neutralization epitopes. Apparent clusters of amino acid substitutions were present in the hydrophobic signal sequence, transmembrane domain, and within the cytoplasmic domain of the G protein. A maximum parsimony analysis of the G gene nucleotide sequences allowed construction of a phylogram indicating the evolutionary relationship of these viruses. The VSV NJ serotype appears to contain at least three distinct lineages or subtypes. All recent virus isolates from the United States and Mexico are within subtype I and appear to have evolved from an ancestor more closely related to the Hazelhurst historic strain than other older strains. The implications of these findings for the evolution, epizootiology, and classification of these viruses are discussed.

Amino Acid Sequence↗

Structure, expression and duplication of genes which encode phosphoglyceromutase of Drosophila melanogaster.

We report here the isolation and characterization of genes from Drosophila that encode the glycolytic enzyme phosphoglyceromutase (PGLYM). Two genomic regions have been isolated that have potential to encode PGLYM. Their cytogenetic localizations have been determined by in situ hybridization to salivary gland chromosomes. One gene, Pglym78, is found at 78A/B and the other, Pglym87, at 87B4,5 of the Drosophila polytene map. Pglym78 transcription follows a developmental pattern similar to other glycolytic genes in Drosophila, i.e., substantial maternal transcript deposited during oogenesis; a decline in abundance in the first half of embryogenesis; a subsequent increase in the second half of embryogenesis which continues throughout larval life; a decline in pupae and a second increase to a plateau in adults. This transcript has been mapped by cDNA and genomic sequence comparison, RNase protection, and primer extension. Using similar analyses transcripts of Pglym87 could not be detected. Pglym78 has two introns which interrupt the coding region, while the Pglym87 gene lacks introns. This and other features support a model of retrotransposition mediated gene duplication for the origin of Pglym87. The apparent absence of a complete, intact coding frame and transcript suggest that Pglym87 is a pseudogene. However, retention of reading frame and codon bias suggests that Pglym87 may retain coding function, or may have been inactivated recently, substantially after the time of duplication, or that the molecular evolution of Pglym87 is unusual. Similarities of the unusual molecular evolution of Pglym87 and other proposed pseudogenes are discussed.

Amino Acid Sequence↗

Phenotypic reversion of an IS1-mediated deletion mutation: a combined role for point mutations and deletions in transposon evolution.

We have physically characterised a deletion mutant of the R plasmid R100 which has lost all of the antibiotic resistances, including chloramphenicol resistance (Cmr), coded by its IS1-flanked r-determinant. The deletion was mediated by one of the flanking IS1 elements and terminates within the carboxyl terminus of the Cmr gene. DNA sequence analysis showed that the mutated gene would produce a protein 20 amino acids longer than the wild-type due to fusion with an open reading frame in the IS element. Surprisingly for a deletion mutation, rare, spontaneous Cmr revertants could be recovered. Two of the four revertants studied had frame shifts due to the insertion of a single AT base pair at the same position; the revertants would code for a protein five amino acids shorter than the wild-type. The other two revertants had acquired duplications of the 34-bp inverted terminal repeat sequences of the IS1 element and would direct the synthesis of a protein six amino acids longer than the wild-type. The reverted Cmr markers were still capable of transposition. These observations suggest a role for point mutations and small DNA rearrangements in the formation of new gene organisations produced by mobile genetic elements.

Base Sequence↗

Organization and structure of Volvox beta-tubulin genes.

Genomic clones encoding two Volvox beta-tubulin genes have been isolated and shown to represent the only two beta-tubulin genes in the genome. Restriction fragment length polymorphism analysis was used to demonstrate that the two genes are genetically linked. One of these genes was sequenced and the mRNA start site(s) determined by primer extension. A comparison of its sequence to those of the two beta-tubulin genes of Chlamydomonas revealed: (1) a high degree of conservation of the coding region, with the predicted amino acid sequence differing only in the C-terminal residue; (2) extensive sequence conservation in the 5' untranslated leader region and a 16 bp (putative regulatory) sequence in the promoter region; (3) the same number and location of introns, with a short region of homology in intron 1, but little significant homology in introns 2 and 3.

Amino Acid Sequence↗

Statistical analysis of L-tuple frequencies in eubacteria and organelles.

This work is an attempt to study the structural features and evolutionary patterns of nucleotide sequences by analyzing their 1- through 4-plet frequencies and statistical relations between them. We present mathematical apparatus for this analysis. In particular, we introduce criteria to estimate the degree of homogeneity of L-plet composition in a given set of sequences and the dependence of the L-plet frequencies on the composition of lower orders. We apply these criteria to the study of eubacteria, mitochondria and chloroplasts. We demonstrate that L-plet frequencies are quite useful for revealing evolutionary relationship between DNA sequences and that the non-random distribution is more typical for doublets than to triplets. Non-randomness of triplet composition is more characteristic to coding than to non-coding regions, while no significant differences in dinucleotide composition can be observed. The obtained results can be used for revealing possible mechanisms of the codon usage phenomena.

Analysis of Variance↗

Protein-coding genes as molecular markers for ecologically distinct populations: the case of two Bacillus species.

Bacillus globisporus and Bacillus psychrophilus are one among many pairs of ecologically distinct taxa that are distinguished by very few nucleotide differences in 16S rRNA gene sequence. This study has investigated whether the lack of divergence in 16S rRNA between such species stems from the unusually slow rate of evolution of this molecule, or whether other factors might be preventing neutral sequence divergence at 16S rRNA as well as every other gene. B. globisporus and B. psychrophilus were each surveyed for restriction-site variation in two protein-coding genes. These species were easily distinguished as separate DNA sequence clusters for each gene. The limited ability of 16S rRNA to distinguish these species is therefore a consequence of the extremely slow rate of 16S rRNA evolution. The present results, and previous results involving two Mycobacterium species, demonstrate that there exist closely related species which have diverged long enough to have formed clearly separate sequence clusters for protein-coding genes, but not for 16S rRNA. These results support an earlier argument that sequence clustering in protein-coding genes could be a primary criterion for discovering and identifying ecologically distinct groups, and classifying them as separate species.

Alanine Dehydrogenase↗

Ubiquitous selective constraints in the Drosophila genome revealed by a genome-wide interspecies comparison.

Non-coding DNA comprises approximately 80% of the euchromatic portion of the Drosophila melanogaster genome. Non-coding sequences are known to contain functionally important elements controlling gene expression, but the proportion of sites that are selectively constrained is still largely unknown. We have compared the complete D. melanogaster and Drosophila simulans genome sequences to estimate mean selective constraint (the fraction of mutations that are eliminated by selection) in coding and non-coding DNA by standardizing to substitution rates in putatively unconstrained sequences. We show that constraint is positively correlated with intronic and intergenic sequence length and is generally remarkably strong in non-coding DNA, implying that more than half of all point mutations in the Drosophila genome are deleterious. This fraction is also likely to be an underestimate if many substitutions in non-coding DNA are adaptively driven to fixation. We also show that substitutions in long introns and intergenic sequences are clustered, such that there is an excess of substitutions <8 bp apart and a deficit farther apart. These results suggest that there are blocks of constrained nucleotides, presumably involved in gene expression control, that are concentrated in long non-coding sequences. Furthermore, we infer that there is more than three times as much functional non-coding DNA as protein-coding DNA in the Drosophila genome. Most deleterious mutations therefore occur in non-coding DNA, and these may make an important contribution to a wide variety of evolutionary processes.

Animals↗

vrrB, a hypervariable open reading frame in Bacillus anthracis.

Bacillus anthracis appears to be the most molecularly homogeneous bacterial species known. Extensive surveys of worldwide isolates have revealed vanishingly small amounts of genomic variation. The biological importance of the resting-stage spore may lead to very low evolutionary rates and, perhaps, to the lack of potentially adaptive genetic variation. In contrast to the overall homogeneity, some gene coding regions contain hypervariability that is translated into protein variation. During marker analysis of diverse strains, we have discovered a novel ca. 750-nucleotide open reading frame (ORF) that contains in-frame, variable-number tandem-repeat sequences. Four distinct variable regions exist within vrrB, giving rise to 11 distinct alleles in eight different length categories among B. anthracis strains. This ORF putatively codes for a 241- to 265-amino-acid protein, rich in glutamine (13.2%), glycine (23.4%), and histidine (23.0%). The variable-region amino acids of the vrrB ORF are strongly hydrophilic. Coupled with putative transmembrane domains flanking the variable regions, this suggests a membrane-anchored cytosolic or extracellular location for the putative protein. Sequence analysis of the complete ORFs from three Bacillus cereus strains shows maintenance of the ORF across species boundaries, including strong conservation of the amino acid sequence and the capacity to vary among strains. The presence of 11 different alleles of the vrrB locus is in stark contrast to the near homogeneity of B. anthracis. Evolution of hypervariable genes can negate the lack of genetic variability in species such as B. anthracis and provide select rapid evolution in other more variable species.

Amino Acid Sequence↗

Genome organization and reorganization in evolution: formatting for computation and function.

This volume deals with the role of epigenetics in life and evolution. The most dynamic forms of functional genome formatting involve DNA interacting with cellular complexes that do not alter sequence information. Such important epigenetic phenomena are the main subjects of other articles in this volume. This article focuses on the long-lived form of genome formatting that lies within the DNA sequence itself. I argue for a computational view of genome function as the long-term information storage organelle of each cell. Structural formatting consists of organizing various signals and coding sequences into computationally ready systems facilitating genome expression and genome transmission. The basic features of genome organization can be understood by examining the E. coli lac operon as a paradigmatic genomic system. Multiple systems are connected through distributed signals and repetitive DNA to form higher-order genome system architectures. Molecular discoveries about mechanisms of DNA restructuring show that cells possess the natural genetic engineering functions necessary for evolutionary change by rearranging genomic components and reorganizing system architectures. The concepts of cellular computation and decision-making, genome system architecture, and natural genetic engineering combine to provide a new way of framing evolutionary theories and understanding genome sequence information.

Animals↗

Hemoglobins, XLVIIII. The primary structure of a monomeric hemoglobin from the hagfish, Myxine glutinosa L.: evolutionary aspects and comparative studies of the function with special reference to the heme linkage.

Hagfish hemoglobin has three main components, one of which is Hb III. It is monomeric and consists of 148 amino acid residues (M = 17 350). Its complete primary structure, previously published, is discussed here. The proximal amino acid (F8) of the heme linkage is histidine as always in the hemoglobins, but the regularly expected distal histidine E7 is substituted by glutamine. This substitution, leading to a new kind of heme linkage, has hitherto only been demonstrated in opossum hemoglobin. It is suggested that E7, Gln, is directed out of the heme pocket, and that the adjacent Ell, Ile, is directed toward the inside of the pocket, giving the distal heme contact instead of histidine. Myxine Hb III has an additional tail of 9 amino acid residues at its N-terminal end, as has the hemoglobin of Lampetra fluviatilis. The genetic codes of Myxine and Lampetra hemoglobins show 117 differences, in spite of many morphological resemblances between hagfish and lamprey. Their primary hemoglobin structures show differences substantial enough to be compatible with the divergence of the two families some 400-500 million years ago.

Amino Acid Sequence↗

Catalog of 162 single nucleotide polymorphisms (SNPs) in a 4.7-kb region of the HLA-DP loci in southern Chinese ethnic groups.

HLA class-II proteins are cell-surface molecules that present antigens to T cells, and their expressional regulation is crucial to the immune reaction. Sequence variation at the regulatory region can directly affect the gene expression level. We cloned and sequenced a 4.7-kb region containing the regulatory region, exon1, and partial intron1 of both HLA-DPA1 and DPB1 genes in 25 variable sequences from southern Chinese ethnic groups and got a high-density map of 162 single nucleotide polymorphisms (SNPs): seven in 5'-flanking regions, four in 5'-untranslated regions, and four in the coding regions. By comparing these data with SNPs in dbSNP database in the NCBI, 145 SNPs (89.5%) were novel. In addition, eight genetic variations of insertion-deletion polymorphisms (INDELs) were discovered within the 4.7-kb region. These high-resolution maps can be used as resources of markers for association studies of complex diseases, assessment of individuals' predisposition to diseases, and tailoring of therapies, as well as research markers for population genetics and evolution.

Base Sequence↗

Concordant evolution of coding and noncoding regions of DNA made possible by the universal rule of TA/CG deficiency-TG/CT excess.

The universal rule of TA/CG deficiency-TG/CT excess previously proposed as the construction principle of coding sequences applies to noncoding regions of the gene as well. Analysis of a 1989-base-long gene sequence for mouse immunoglobulin gamma 2a heavy-chain constant region as well as the 19,002-base-long gene sequence for human serum albumin revealed deficiency and overabundance of very similar sets of base trimers and tetramers in the coding and noncoding regions of the same gene, in spite of the fact that noncoding regions were considerably richer in A + T. Inasmuch as this universal rule does not discriminate one strand of DNA double helix from another, two complementary DNA strands of the entire gene maintained nearly perfect symmetry. That is to say, the degrees of excesses, deficiencies of the 64-base trimers remained nearly identical between two complementary strands, and this symmetry was only slightly disturbed in the coding region. It would thus appear that the universal rule as an intrinsic force has been exerting far greater influence than natural selection in the evolution of genes.

Animals↗

Orientation of loci in the major histocompatibility complex of the rat and its comparison to man and the mouse.

Among 290 F2 progeny of an r10 X ACP cross were two recombinants which allowed the loci for glyoxalase-1 and neuraminidase-1 to be mapped relative to the RT1.A and dw-3 loci in the major histocompatibility complex (MHC) of the rat. In 673 progeny of the same cross there was a recombinant between ft and dw-3, and in 403 progeny of the backcross BY1 X (BY1 X BDIX)F1 there was another recombinant between ft and dw-3. These data, combined with those from previous studies, provide the information for constructing a detailed map of the rat major histocompatibility complex: the gene order and size in the rat are very similar to those in the mouse and different from those in man and in the other species that have been studied. Comparison of the structures of the MHC in the various species leads to a hypothesis about the evolution of the MHC which involves sequential duplications of the genes coding for class I and class II loci and an inversion in the prototypic muridae which placed the class II loci between the class I loci.

Animals↗

Multiple distinct splicing enhancers in the protein-coding sequences of a constitutively spliced pre-mRNA.

We have identified multiple distinct splicing enhancer elements within protein-coding sequences of the constitutively spliced human beta-globin pre-mRNA. Each of these highly conserved sequences is sufficient to activate the splicing of a heterologous enhancer-dependent pre-mRNA. One of these enhancers is activated by and binds to the SR protein SC35, whereas at least two others are activated by the SR protein SF2/ASF. A single base mutation within another enhancer element inactivates the enhancer but does not change the encoded amino acid. Thus, overlapping protein coding and RNA recognition elements may be coselected during evolution. These studies provide the first direct evidence that SR protein-specific splicing enhancers are located within the coding regions of constitutively spliced pre-mRNAs. We propose that these enhancers function as multisite splicing enhancers to specify 3' splice-site selection.

Cross-Linking Reagents↗

Structure of a gene coding for human HMG2 protein.

A human genomic library was screened with the pig thymus cDNA coding for chromosomal protein HMG2. A 4341-base pair fragment containing the entire gene encoding this protein was isolated and characterized. The HMG2 gene is 2665 base pairs long from the start site to the end of transcription and comprises 5 exons. The size of mRNA postulated from the exons is 1125 base pairs long, consistent with that obtained by Northern hybridization analysis. The canonical 5'-regulatory motifs, CCAAT, are present, while the TATA element is absent from the gene. Southern analysis suggested that HMG2 protein is encoded by a single or only a few genes of high homology. The primary structure of the human HMG2 protein consists of 208 amino acid residues, deduced from the coding region of the gene, is different from that of pig HMG2 in only two amino acids; one is exchanged and the other is missing. The amino acid sequences of two DNA binding domains, "HMG-box," also recently found in several transcription factors, are completely homologous in human and pig HMG2. The present study, which is the first one on the isolation and characterization of complete gene coding for HMG2 protein, may be useful for evolutional and genomic analysis of the proteins containing the HMG-box sequences for DNA binding.

Amino Acid Sequence↗