PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Trends in protein evolution inferred from sequence and structure analysis.

Complementary developments in comparative genomics, protein structure determination and in-depth comparison of protein sequences and structures have provided a better understanding of the prevailing trends in the emergence and diversification of protein domains. The investigation of deep relationships among different classes of proteins involved in key cellular functions, such as nucleic acid polymerases and other nucleotide-dependent enzymes, indicates that a substantial set of diverse protein domains evolved within the primordial, ribozyme-dominated RNA world.

Evolution, Molecular↗

Molecular structure and evolution of DNA sequences located at the alpha satellite boundary of chromosome 20.

We have isolated and characterised one PAC clone (dJ233C1) containing a linkage between alphoid and non-alphoid DNA. The non-alphoid DNA was found to map at the pericentromeric region of chromosome 20, both on p and q sides, and to contain homologies with one contig (ctg176, Sanger Centre), also located in the same chromosome region. At variance with the chromosome specificity shown by the majority of non-alphoid DNA, a subset of alphoid repeats derived from the PAC yielded FISH hybridisation signals located at the centromeric region of several human chromosomes, belonging to three different suprachromosomal families. The evolutionary conservation of this boundary region was investigated by comparative FISH experiments on chromosomes from great apes. The non-alphoid DNA was found to have undergone events of expansion and transposition to different pericentromeric regions of great apes chromosomes. Alphoid sequences revealed a very wide distribution of FISH signals in the great apes. The pattern was substantially discordant with the data available in the literature, which is essentially derived from the central alphoid subset. These results add further support to the emerging opinion that the pericentromeric regions are high plastics, and that the alpha satellite junctions do not share the evolutionary history with the main subsets.

Animals↗

The evolution of repetitive DNA sequences in sea urchins.

Molecular hybridization of nuclear DNAs has been employed to study the evolution of the repetitive DNA sequences in four species of sea urchin. The data show that relative to S. purpuratus there has been approximately 0.1% sequence divergence per million years in the repetitive DNA sequences of S. droebachiensis, S. franciscanus, and L. pictus. These results confirm that repetitive DNA sequences are strongly conserved during evolution. However, comparison of the extent of base pair mismatch in the repetitive DNA heteroduplexes formed at Cot 20 with those formed at Cot 200 during the hybridization of S. purpuratus and L. pictus DNAs reveals that highly repetitive sequences of sea urchins may diverge more rapidly than do the more moderately repetitive sequences.

Animals↗

Structural constraints and emergence of sequence patterns in protein evolution.

The aim of this work was to study the relationship between structure conservation and sequence divergence in protein evolution. To this end, we developed a model of structurally constrained protein evolution (SCPE) in which trial sequences, generated by random mutations at gene level, are selected against departure from a reference three-dimensional structure. Since at the mutational level SCPE is completely unbiased, any emergent sequence pattern will be due exclusively to structural constraints. In this first report, it is shown that SCPE correctly predicts the characteristic hexapeptide motif of the left-handed parallel beta helix (LbetaH) domain of UDP-N-acetylglucosamine acyltransferases (LpxA).

Acyltransferases↗

Evolution of simple sequence repeats.

Simple Sequence Repeats (SSRs) are common and frequently polymorphic in eukaryote DNA. Many are subject to high rates of length mutation in which a gain or loss of one repeat unit is most often observed. Can the observed abundances and their length distributions be explained as the result of an unbiased random walk, starting from some initial repeat length? In order to address this question, we have considered two models for an unbiased random walk on the integers, n (n0 < or = n). The first is a continuous time process (Birth and Death Model or BDM) in which the probability of a transition to n + 1 or n - 1 is lambda k, with k = n - n0 + 1 per unit time. The second is a discrete time model (Random Walk Model or RWM), in which a transition is made at each time step, either to n - 1 or to n + 1. In each case the walks start at length n0, with new walks being generated at a steady rate, S, the source rate, determined by a base substitution rate of mutation from neighboring sequences. Each walk terminates whenever n reaches n0 - 1 or at some time, T, which reflects the contamination of pure repeat sequences by other mutations that remove them from consideration, either because they fail to satisfy the criteria for repeat selection from some database or because they can no longer undergo efficient length mutations. For infinite T, the results are particularly simple for N(k), the expected number of repeats of length n = k + n0 - 1, being, for BDM, N(k) = S/k lambda, and for RWM, N(k) = 2S. In each case, there is a cut-off value of k for finite T, namely k = T lambda ln2 for BDM and k = 0.57 square root of T for RWM; for larger values of k, N(k) becomes rapidly smaller than the infinite time limit. We argue that these results may be compared with SSR length distributions averaged over many loci, but not for a particular locus, for which founder effects are important. For the data of Beckmann & Weber [(1992), Genomics 12, 627] on GT.AC repeats in the human, each model gives a reasonable fit to the data, with the source at two repeat units (n0 = 2). Both the absolute number of loci and their length distribution are well represented.

DNA↗

Evaluation of models for the evolution of protein sequences and functions under structural constraint.

In the field of evolutionary structural genomics, methods are needed to evaluate why genomes evolved to contain the fold distributions that are observed. In order to study the effects of population dynamics in the evolved genomes we need fast and accurate evolutionary models which can analyze the effects of selection, drift and fixation of a protein sequence in a population that are grounded by physical parameters governing the folding and binding properties of the sequence. In this study, various knowledge-based, force field, and statistical methods for protein folding have been evaluated with four different folds: SH2 domains, SH3 domains, Globin-like, and Flavodoxin-like, to evaluate the speed and accuracy of the energy functions. Similarly, knowledge-based and force field methods have been used to predict ligand binding specificity in SH2 domain. To demonstrate the applicability of these methods, the dynamics of evolution of new binding capabilities by an SH2 domain is demonstrated.

Computational Biology↗

Catabolic ornithine transcarbamylase of Halobacterium halobium (salinarium): purification, characterization, sequence determination, and evolution.

Halobacterium halobium (salinarium) is able to grow fermentatively via the arginine deiminase pathway, which is mediated by three enzymes and one membrane-bound arginine-ornithine antiporter. One of the enzymes, catabolic ornithine transcarbamylase (cOTCase), was purified from fermentatively grown cultures by gel filtration and ammonium sulfate-mediated hydrophobic chromatography. It consists of a single type of subunit with an apparent molecular mass of 41 kDa. As is common for proteins of halophilic Archaea, the cOTCase is unstable below 1 M salt. In contrast to the cOTCase from Pseudomonas aeruginosa, the halophilic enzyme exhibits Michaelis-Menten kinetics with both carbamylphosphate and ornithine as substrates with Km values of 0.4 and 8 mM, respectively. The N-terminal sequences of the protein and four peptides were determined, comprising about 30% of the polypeptide. The sequence information was used to clone and sequence the corresponding gene, argB. It codes for a polypeptide of 295 amino acids with a calculated molecular mass of 32 kDa and an amino acid composition which is typical of halophilic proteins. The native molecular mass was determined to be 200 kDa, and therefore the cOTCase is a hexamer of identical subunits. The deduced protein sequence was compared to the cOTCase of P. aeruginosa and 14 anabolic OTCases, and a phylogenetic tree was constructed. The halobacterial cOTCase is more distantly related to the cOTCase than to the anabolic OTCase of P. aeruginosa. It is found in a group with the anabolic OTCases of Bacillus subtilis, P. aeruginosa, and Mycobacterium bovis.

Allosteric Regulation↗

Strong male-driven evolution of DNA sequences in humans and apes.

Studies of human genetic diseases have suggested a higher mutation rate in males than in females and the male-to-female ratio (alpha) of mutation rate has been estimated from DNA sequence and microsatellite data to be about 4-6 in higher primates. Two recent studies, however, claim that alpha is only about 2 in humans. This is even smaller than the estimates (alpha > 4) for carnivores and birds; humans should have a higher alpha than carnivores and birds because of a longer generation time and a larger sex difference in the number of germ cell cycles. To resolve this issue, we sequenced a noncoding fragment on Y of about 10.4 kilobases (kb) and a homologous region on chromosome 3 in humans, greater apes, and lesser apes. Here we show that our estimate of alpha from the internal branches of the phylogeny is 5.25 (95% confidence interval (CI) 2.44 to infinity), similar to the previous estimates, but significantly higher than the two recent ones. In contrast, for the external (short, species-specific) branches, alpha is only 2.23 (95% CI: 1.47-3.84). We suggest that closely related species are not suitable for estimating alpha, because of ancient polymorphism and other factors. Moreover, we provide an explanation for the small estimate of alpha in a previous study. Our study reinstates a high alpha in hominoids and supports the view that DNA replication errors are the primary source of germline mutation.

Animals↗

Lineage-specific variations of congruent evolution among DNA sequences from three genomes, and relaxed selective constraints on rbcL in Cryptomonas (Cryptophyceae).

BACKGROUND: Plastid-bearing cryptophytes like Cryptomonas contain four genomes in a cell, the nucleus, the nucleomorph, the plastid genome and the mitochondrial genome. Comparative phylogenetic analyses encompassing DNA sequences from three different genomes were performed on nineteen photosynthetic and four colorless Cryptomonas strains. Twenty-three rbcL genes and fourteen nuclear SSU rDNA sequences were newly sequenced to examine the impact of photosynthesis loss on codon usage in the rbcL genes, and to compare the rbcL gene phylogeny in terms of tree topology and evolutionary rates with phylogenies inferred from nuclear ribosomal DNA (concatenated SSU rDNA, ITS2 and partial LSU rDNA), and nucleomorph SSU rDNA. RESULTS: Largely congruent branching patterns and accelerated evolutionary rates were found in nucleomorph SSU rDNA and rbcL genes in a clade that consisted of photosynthetic and colorless species suggesting a coevolution of the two genomes. The extremely accelerated rates in the rbcL phylogeny correlated with a shift from selection to mutation drift in codon usage of two-fold degenerate NNY codons comprising the amino acids asparagine, aspartate, histidine, phenylalanine, and tyrosine. Cysteine was the sole exception. The shift in codon usage seemed to follow a gradient from early diverging photosynthetic to late diverging photosynthetic or heterotrophic taxa along the branches. In the early branching taxa, codon preferences were changed in one to two amino acids, whereas in the late diverging taxa, including the colorless strains, between four and five amino acids showed changes in codon usage. CONCLUSION: Nucleomorph and plastid gene phylogenies indicate that loss of photosynthesis in the colorless Cryptomonas strains examined in this study possibly was the result of accelerated evolutionary rates that started already in photosynthetic ancestors. Shifts in codon usage are usually considered to be caused by changes in functional constraints and in gene expression levels. Thus, the increasing influence of mutation drift on codon usage along the clade may indicate gradually relaxed constraints and reduced expression levels on the rbcL gene, finally correlating with a loss of photosynthesis in the colorless Cryptomonas paramaecium strains.

Asparagine↗

Phylogeny of vertebrate nuclear receptors--analysis of variance components in protein sequences.

Nuclear receptors (NR) constitute a large family of proteins and play a crucial role in regulating mineral metabolism and physiological homeostasis of various organ systems. The aim of this study was to elucidate whether the variance among NRs of estrogen, androgen and vitamin-D in various vertebrate species including humans is attributed to differences between the taxonomic groups within a specific receptor (i.e. between orthologous) or between the different proteins within the taxon (i.e. between paralogous genes). Published data on 57 protein sequences of the above NRs were used for phylogenetic analysis. The results showed that in DNA- and ligand-binding regions, 94% and 70% of variance is due to differences between the three proteins. However, in non-binding regions, 47% of the variance results from differences between the three paralogous proteins. Human sequences consistently clustered with their mammal orthologous within the three groups of NR sequences, clearly indicating that evolution of human sequences is not distinct from mammal sequence evolution.

Analysis of Variance↗

Amino acid sequence of heavy chain from Xenopus laevis IgM deduced from cDNA sequence: implications for evolution of immunoglobulin domains.

Present understanding of the evolution of immunoglobulins is derived almost entirely from studies of a few mammalian species. To obtain information about immunoglobulin genes in Xenopus laevis, a cDNA library was prepared in the expression vector lambda gt11 from mitogen-stimulated splenocytes of this species. Of approximately equal to 50,000 clones screened, 18 were found to express IgM epitopes. One of these, lambda XIg14, hybridized with RNA of RNA of approximately equal to 2 kilobases from splenocytes. The insert of this clone appears to encode a variable region and part of a mu constant region; that of another clone, lambda XIg8, appears to encode a variable region and a complete mu constant region. Both inserts contain sequence corresponding to the three gene segments (VH, DH, and JH) that encode heavy-chain variable regions. The heavy-chain constant region (CH) encoded by lambda XIg8 has the characteristic features of C mu, including a four-domain structure and a carboxyl-terminal tail. The amino acid sequences of two mu-chain peptides agree with the cDNA sequence. The identity in amino acid sequence between the corresponding Xenopus and mouse C mu domains ranges from 31 to 47%. The C mu domains vary in the extent to which their sequences resemble the sequences of other immunoglobulins, consistent with previous suggestions that the immunoglobulin domains have an independent evolutionary history.

Amino Acid Sequence↗

The stereospecificity of sequential nicotinamide-adenine dinucleotide-dependent oxidoreductases in relation to the evolution of metabolic sequences.

The generalization that 'when a metabolic sequence involves consecutive nicotinamide-adenine dinucleotide-dependent reactions, the dehydrogenases have the same stereospecificity' was tested and confirmed for three metabolic sequences. (1) NAD+-xylitol (D-xylulose) dehydrogenase and NADP+-xylitol (L-xylulose) dehydrogenase are both B-specific. (2) D-Mannitol 1-phosphate dehydrogenase and D-sorbitol 6-phosphate dehydrogenase are both B-specific. (3) meso Tartrate dehydrogenase and oxaloglycollate reductive decarboxylase are both A-specific. Other dehydrogenases associated with the metabolism of meso-tartrate in Pseudomonas putida, such as hydroxypyruvate reductase and tartronate semialdehyde reductase, were also shown to be A-specific. Malate dehydrogenase from Pseudomonas putida was A-specific, and the proposition is discussed that the common A-stereospecificity among the dehydrogenases involved in meso-tartrate metabolism reflects their origin from malate dehydrogenase.

Alcohol Oxidoreductases↗

A new polymorphism in the human factor VIII gene: implications for linkage analysis in haemophilia A and for the evolution of int22h sequences.

A new polymorphism in the human factor VIII gene has been localized and characterized. It is a biallelic, single nucleotide polymorphism located in intron 22 of the gene, within the 9.5 kb int22h-1 segment. The allelic forms are G (frequency 0.65) and A (frequency 0.35), giving a predicted rate of heterozygosity of 0.46. The polymorphism occurs within a CG dinucleotide and affects an MspI site (CCGG). Int22h-1 is duplicated twice extragenically at Xq28; both extragenic copies (int22h-2 and -3) are also polymorphic with respect to MspI. Investigation of 156 MspI [-] alleles, comprising 30 intragenic and 126 extragenic sites, indicated that all were due to A alleles and none had arisen by C to T transition within the CG dinucleotide. The intragenic MspI site (designated MspI A) is located 737 bases downstream of a previously described XbaI restriction fragment length polymorphism. Despite their close proximity, the polymorphisms are not in complete linkage disequilibrium; haplotype analysis in 85 factor VIII genes from a Caucasian population predicts an informativity of approximately 60% in linkage studies using both, compared with an informativity of approximately 47% in studies using either on its own.

Chromosome Mapping↗

Changes in viral loads of lamivudine-resistant mutants and evolution of HBV sequences during adefovir dipivoxil therapy.

The addition of adefovir dipivoxil (ADV) to ongoing lamivudine therapy is effective against lamivudine-resistant virus in patients with hepatitis B virus (HBV) infection. We studied 39 patients who received ADV added to lamivudine for breakthrough hepatitis. We determined early viral changes (12 weeks) in YMDD mutants (rtM204I [YIDD sequence], rtM204V [YVDD]) and rtL180M in all 39 patients as well as amino acid changes in the polymerase reverse transcriptase (rt) region and precore/core promoter mutations in 15 patients who received long-term treatment (more than 1 year). Changes in rtM204I and rtL180M viral loads were greater than that of the rtM204V, albeit statistically insignificant. Moreover, the greatest change in viral load was seen for rtM204I without hepatitis B e antigen (HBeAg). The precore mutant was replaced with wild-type virus in three of eight patients after 1 year of added ADV therapy. Compared to baseline with lamivudine therapy only, new amino acid mutations were seen in the rt region at baseline with ADV in seven patients. At 1 year after ADV coadministration, the YMDD motif was replaced with wild-type (rt204M) in two patients, in whom mutations were fewer and of a different type. We conclude that the rtM204I may be more sensitive to ADV in vivo. ADV tended to select wild-type virus from precore mutants. Moreover, viruses that were wild-type in the rt region reappeared after 1 year of ADV coadministration in some patients.

Adenine↗