PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Nucleotide sequence and molecular evolution of the gene coding for glyceraldehyde-3-phosphate dehydrogenase in the thermoacidophilic archaebacterium Sulfolobus solfataricus.

A Sulfolobus solfataricus genomic library cloned in the EMBL3 phage was screened using as probes synthetic oligonucleotides designed from the known amino acid sequence of a peptide obtained from the purified glyceraldehyde-3-phosphate dehydrogenase (aGAPD) protein. The screening led to the isolation of six recombinant phages (lambda G1-lambda G6) and one of them (lambda G4) contained the entire GAPD gene. The deduced amino acid sequence accounts for a protein made of 341 amino acids and the initial methionine is encoded by a GTG triplet. Alignment of the S. solfataricus aGAPD sequence versus GAPD from archaea, eukarya, and bacteria showed that aGAPD is very similar to other archaebacterial but not to eukaryotic or eubacterial GAPD. For known archaebacterial GAPD sequences, the rate of nucleotide substitutions per site per year showed that these sequences are homologous not only at the amino acid but also at the nucleotide level. The evolutionary rates are nearly similar to those reported for other eukaryotic genes.

Amino Acid Sequence↗

Conservation of engrailed-like homeobox sequences during vertebrate evolution.

The Drosophila melanogaster developmental gene engrailed (en) is a member of a distinct subfamily of homeobox genes with a wide phylogenetic distribution. Here we report the use of reduced stringency polymerase chain reaction (PCR) to amplify and clone 8 genes related to en from 5 vertebrate species, including representatives of the most ancient vertebrate lineages. Nucleotide and deduced amino acid sequence comparisons between mouse, toad, zebrafish, lamprey and hagfish genes reveal extensive evolutionary conservation, and suggests that 2 en-like genes have been retained in most vertebrate lineages.

Amino Acid Sequence↗

Hypervariable noncoding sequences in Saccharomyces cerevisiae.

Compared to protein-coding sequences, the evolution of noncoding sequences and the selective constraints placed on these sequences is not well characterized. To compare the evolution of coding and noncoding sequences, we have conducted a survey for DNA polymorphism at five randomly chosen loci among a diverse collection of 81 strains of Saccharomyces cerevisiae. Average rates of both polymorphism and divergence are 40% lower at noncoding sites and 90% lower at nonsynonymous sites in comparison to synonymous sites. Although noncoding and coding sequences show substantial variability in ratios of polymorphism to divergence, two of the loci, MLS1 and PDR10, show a higher rate of polymorphism at noncoding compared to synonymous sites. The high rate of polymorphism is not accompanied by a high rate of divergence and is limited to a few small regions. These hypervariable regions include sites with three segregating bases at a single site and adjacent polymorphic sites. We show that this clustering of polymorphic sites is significantly greater than one would expect on the basis of the spacing between polymorphic fourfold degenerate sites. Although hypervariable noncoding sequences could result from selection on regulatory mutations, they could also result from transient mutational hotspots.

Base Pairing↗

Implications of thermodynamics of protein folding for evolution of primary sequences.

Natural proteins exhibit essentially two-state thermodynamics, with one stable fold that dominates thermodynamically over a vast number of possible folds, a number that increases exponentially with the size of the protein. Here we address the question of whether this feature of proteins is a rare property selected by evolution or whether it is in fact true of a significant proportion of all possible protein sequences. Using statistical procedures developed to study spin glasses, we show that, given certain assumptions, the probability that a randomly synthesized protein chain will have a dominant fold (which is the global minimum of free energy) is a function of temperature, and that below a critical temperature the probability rapidly increases as the temperature decreases. Our results suggest that a significant proportion of all possible protein sequences could have a thermodynamically dominant fold.

Amino Acid Sequence↗

A cluster of cuticle protein genes of Drosophila melanogaster at 65A: sequence, structure and evolution.

A 36-kb genomic DNA segment of the Drosophila melanogaster genome containing 12 clustered cuticle genes has been mapped and partially sequenced. The cluster maps at 65A 5-6 on the left arm of the third chromosome, in agreement with the previously determined location of a putative cluster encompassing the genes for the third instar larval cuticle proteins LCP5, LCP6 and LCP8. This cluster is the largest cuticle gene cluster discovered to date and shows a number of surprising features that explain in part the genetic complexity of the LCP5, LCP6 and LCP8 loci. The genes encoding LCP5 and LCP8 are multiple copy genes and the presence of extensive similarity in their coding regions gives the first evidence for gene conversion in cuticle genes. In addition, five genes in the cluster are intronless. Four of these five have arisen by retroposition. The other genes in the cluster have a single intron located at an unusual location for insect cuticle genes.

Animals↗

Rh gene evolution in primates: study of intron sequences.

By amplification and sequencing of RH gene intron 4 of various primates we demonstrate that an Alu-Sx-like element has been inserted in the RH gene of the common ancestor of humans, apes, Old World monkeys, and New World monkeys. The study of mouse and lemur intron 4 sequences allowed us to precisely define the insertion point of the Alu-Sx element in intron 4 of the RH gene ancestor common to Anthropoidea. Like humans, chimpanzees and gorillas possess two types of RH intron 4, characterized by the presence (human RHCE and ape RHCE-like genes) or absence (human RHD and ape RHD-like genes) of the Alu-Sx element. This led us to conclude that in the RH common ancestor of humans, chimpanzees, and gorillas, a duplication of the common ancestor gene gave rise to two genes, one differing from the other by a 654-bp deletion encompassing an Alu-Sx element. Moreover, most of chimpanzees and some gorillas posses two types of RHD-like intron 4. The introns 4 of type 1 have a length similar to that of human RHD intron 4, whereas introns 4 of type 2 display an insertion of 12 bp. The latest insertion was not found in the human genome (72 individuals tested). The study of RH intron 3 length polymorphism confirmed that, like humans, chimpanzees and gorillas possess two types of intron 3, with the RHD-type intron 3 being 289 bases shorter than the RHCE intron 3. By amplification and sequencing of regions encompassing introns 3 and 4, we demonstrated that chimpanzee and gorilla RH-like genes displayed associations of introns 3 and 4 distinct to those found in man. Altogether, the results demonstrate that, as in humans, chimpanzee and gorilla RH genes experienced intergenic exchanges.

Animals↗

Evolution of simple sequence in proteins.

The proteins of Saccharomyces cerevisiae contain a high proportion of low-complexity, simple sequences. These are protein segments composed almost exclusively or largely of a single repetitive amino acid polymer and are the most commonly shared feature between proteins. We have examined a survey of other species to determine how widespread this phenomenon might be. This was done by comparing how frequently segments from one protein are present in other proteins. Any recently evolutionarily related proteins were excluded. It was found that the most commonly shared features of eukaryotic proteins were repetitive but that prokaryotes did not contain such shared, extensively redundant repeats. The proportion of eukaryotic proteins that contain a significantly repetitive fraction changes dramatically from species to species. In addition the individual amino acids present in these repeats change between species. This suggests that the primary sequence of the repeats may not be important for their function. Further tests of the yeast repeats confirmed that these repeats evolve more quickly than the remainder of the protein sequence within which they are embedded. These results show that these rapid evolving, simple sequence repeats are in fact the most commonly shared pattern between all of the genomic proteins of eukaryotes.

Amino Acid Sequence↗

The production of de novo folded proteins by a stepwise chain elongation: a model for prebiotic chemical evolution of macromolecular sequences.

We describe an experimental procedure to mimic the formation of long (over 40 residues) co-oligopetide sequences in many identical copies which may have occurred in the prebiotic molecular evolution. The basic hypothesis is that chain formation is based on the stepwise fragment condensation of randomly generated short oligopeptides, whereby the elongation takes place under the contingent environmental constraints (solubility, pH, salinity), which eliminate most of the products, and thus determine the selection towards one particular small set of chains. The present work aims at verifying the validity of this scheme. In order to do so, we utilize a classic synthetic procedure based on the Merrifield solid-phase synthesis of peptides for the synthesis of randomly produced peptides as well as for their stepwise fragment condensation. Thus, starting from a library of peptides with n=10, the first condensation step produces a library of 16 peptides with 20 residues each (n=20), of which only four remain water-soluble and, therefore, capable to undergo the next fragment condensation step. This gives rise to 16 peptides with n=30, out of which twelve precipitate out under the chosen pH and buffer conditions and are eliminated. Finally, a 44-residue-long water-soluble de novo protein is obtained. This has no homologies or similarities with extant proteins, and, based on circular dichroism (CD), it assumes a stable three-dimensional folding. In agreement with CD data, molecular-modelling simulations suggest an helical fold for the protein with poor, if any, structural homology with known proteins. The implication of this procedure as a general mechanism for the etiology of de novo macromolecular sequences and globular proteins in the origin of life is briefly discussed.

Amino Acid Sequence↗

Allelic variation in HLA-B and HLA-C sequences and the evolution of the HLA-B alleles.

Several new HLA-B (B8, B51, Bw62)- and HLA-C (Cw6, Cw7)-specific genes were isolated either as genomic cosmid or cDNA clones to study the diversity of HLA antigens. The allele specificities were identified by sequence analysis in comparison with published HLA-B and -C sequences, by transfection experiments, and Southern and northern blot analysis using oligonucleotide probes. Comparison of the classical HLA-A, -B, and -C sequences reveals that allele-specific substitutions seem to be rare events. HLA-B51 codes only for one allele-specific residue: arginine at position 81 located on the alpha 1 helix, pointing toward the antigen binding site. HLA-B8 contains an acidic substitution in amino acid position 9 on the first central beta sheet which might affect antigen binding capacity, perhaps in combination with the rare replacement at position 67 (F) on the alpha 1 helix. HLA-B8 shows greatest homology to HLA-Bw42, -Bw41, -B7, and -Bw60 antigens, all of which lack the conserved restriction sites Pst I at position 180 and Sac I at position 131. Both sites associated with amino acid replacements seem to be genetic markers of an evolutionary split of the HLA-B alleles, which is also observed in the leader sequences. HLA-Cw7 shows 98% sequence identity to the JY328 gene. In general, the HLA-C alleles display lower levels of variability in the highly polymorphic regions of the alpha 1 and alpha 2 domains, and have more distinct patterns of locus-specific residues in the transmembrane and cytoplasmic domains. Thus we propose a more recent origin for the HLA-C locus.

Amino Acid Sequence↗

First case of mother-to-infant HIV type 1 group O transmission and evolution of C2V3 sequences in the infected child. French HIV Pediatric Cohort Study Group.

We report the first case of mother-to-infant transmission and follow-up for an HIV-1 group O virus from Cameroon. Isolates were obtained from the mother at delivery and from the child at birth and when 16 and 30 months old. We analyzed the viral evolution within mother and child by examining 51 sequences spanning C2V3 regions of the viral envelope gene. The mother carried two genotypes, v1 and v2. The genotype v1 was dominant in the child at birth, and persisted as a minor genotype at age 30 months. The genotype v2 was absent in the child sequences. The variability of the nucleotide sequences of the isolates from the child increased with age from 0.8 to 6%, and a novel genotype (v3) appeared at age 30 months. The nonsynonymous-to-synonymous mutation ratio increased with the age of the child, from 0.75 at birth to 1.86 at 30 months, indicating a high rate of fixation of amino acid changes in the child. The overall pattern was similar to that reported by Ganeshan et al. (J Virol 1997;71:663-677) for group M viruses infecting child with a slow-developing form of the disease.

Acquired Immunodeficiency Syndrome↗

Partial nucleotide sequencing and molecular evolution of epidemic causing Dengue 2 strains.

To study the genetic variability and to detect evolutionary changes and movement of dengue 2 (DEN-2) strains, nucleotide sequencing of the envelope protein gene and the nonstructural protein 1 gene junction was performed for 9 isolates from the 1996 Delhi epidemic and 1 isolate from the 1967 Delhi epidemic. The epidemic strains had a divergence of 10%-11% from the 1967 strains, but were quite similar to DEN-2 isolates from Seychelles, Somalia, and Torres Strait. In addition, the sequence data were compared to the prototype DEN-2 strain, New Guinea C, and other published DEN-2 sequences from different parts of the world. The phylogenetic analysis by the Molecular Evolutionary Genetics Analysis program suggests that the 1996 Delhi isolates of DEN-2 were genotype IV. The 1967 isolate was similar to a 1957 isolate of DEN-2, P9-122, from India, and was classified as genotype V. This study indicates that earlier DEN-2 strains of genotype V have been replaced by genotype IV.

Base Sequence↗

Evolution of DNA sequence homologies between the sex chromosomes in primate species.

Cloned DNA sequences from 18 X-Y homologous loci have been used to examine the evolution of regions of homology between the human X and Y chromosomes. The pattern of X-Y linkage in different primate species has enabled the charting of the chronology of their appearance and removal from the sex chromosomes during evolution. Examination of the pattern of differences in restriction enzyme sites at different loci has been used to estimate the degree of divergence in three different regions of homology. These studies have indicated that (1) blocks of homology have arisen at different points in evolution, (2) different regions of homology are heterogeneous in composition in that they contain X-Y homologous sequences of different age, and (3) the combination of X and Y locations together with the point of evolutionary origin has defined five new patterns of homology.

Animals↗

Molecular evolution of chloroplast DNA sequences.

Comparative data on the evolution of chloroplast genes are reviewed. The chloroplast genome has maintained a similar structural organization over most plant taxa so far examined. Comparisons of nucleotide sequence divergence among chloroplast genes reveals marked similarity across the plant kingdom and beyond to the cyanobacteria (blue-green algae). Estimates of rates of nucleotide substitution indicate a synonymous rate of 1.1 x 10(-9) substitutions per site per year. Noncoding regions also appear to be constrained in their evolution, although addition/deletion events are common. There have also been evolutionary changes in the distribution of introns in chloroplast encoded genes. Relative to mammalian mitochondrial DNA, the chloroplast genome evolves at a conservative rate.

Biological Evolution↗

Baboon and cotton-top tamarin B2m cDNA sequences and the evolution of primate beta 2-microglobulin.

Nonhuman primates represent phylogenetic intermediates for studying the divergence of human and murine beta 2Ms. We report the nucleotide sequences of B2m cDNA clones from a baboon cell line, 26CB-1 (Papio hamadryas; primates: Cercopithecoidea), and a cotton-top tamarin cell line, 1605L (Saguinus oedipus; primates: Ceboidea). The baboon and tamarin B2m sequences indicate a very slow rate of B2m evolution in primates relative to that in murid rodents. Phenotypic evolution of beta 2M has also been very conservative in primates, with only 9-14 substitutions separating baboon or tamarin beta 2Ms from those of humans or orangutans. Analyses of silent and amino-acid-altering nucleotide substitutions provide evidence that negative selection has acted to limit variability in beta strands of primate beta 2Ms, while positive selection has promoted diversity in non-beta-strand regions of murine beta 2Ms. No evidence for the action of selection upon beta 2M residues that contact the class I heavy chain was found in primates or mice. The finding that different selective forces have operated upon primate and murine beta 2Ms suggests that beta 2M may have evolved to serve distinct functions in primates and mice.

Amino Acid Sequence↗

The contribution of slippage-like processes to genome evolution.

Simple sequences present in long (> 30 kb) sequences representative of the single-copy genome of five species (Homo sapiens, Caenorhabditis elegans, Saccharomyces cerevisiae, E. coli, and Mycobacterium leprae) have been analyzed. A close relationship was observed between genome size and the overall level of sequence repetition. This suggested that the incorporation of simple sequences had accompanied increases of genome size during evolution. Densities of simple sequence motifs were higher in noncoding regions than in coding regions in eukaryotes but not in eubacteria. All five genomes showed very biased frequency distributions of simple sequence motifs in all species, particularly in eukaryotes where AAA and TTT predominated. Interspecific comparisons showed that noncoding sequences in eukaryotes showed highly significantly similar frequency distributions of simple sequence motifs but this was not true of coding sequences. ANOVA of the frequency distributions of simple sequence motifs indicated strong contributions from motif base composition and repeat unit length, but much of the variation remained unexplained by these parameters. The sequence composition of simple sequences therefore appears to reflect both underlying sequence biases in slippage-like processes and the action of selection. Frequency distributions of simple sequence motifs in coding sequences correlated weakly or not at all with those in noncoding sequences. Selection on coding sequences to eliminate undesirable sequences may therefore have been strong, particularly in the human lineage.

Animals↗