PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Evolution of the recombination signal sequences in the Ig heavy-chain variable region locus of mammals.

The Ig and T cell receptor (TCR) loci have an exceptionally dynamic evolutionary history, but the mechanisms responsible remain a subject of speculation. Ig and TCR genes are unique in vertebrates in that they are assembled from V, D, and J segments by site-specific recombination in developing lymphocytes. Here we examine the extent to which the V(D)J recombination in germline cells may have been responsible for remodeling Ig and TCR loci in mammals by asking whether gene segments have evolved as a unit, or whether, instead, recombination signal sequences (RSSs) and coding sequences have different phylogenies. Four distinct types of RSS have been defined in the human Ig heavy-chain variable region (Vh) locus, namely H1, H2, H3, and H5, and no other RSS type has been detected in other mammalian species. There is a well-supported discrepancy between the evolutionary history of the RSSs as compared with the Vh coding sequences: the RSS type H2 of one Vh gene segment has clearly become replaced by a RSS type H3 during mammalian evolution, between 115 and 65 million years ago. Two general models might explain the RSS swap: the first involves an unequal crossing over, and the second implicates germline activation of V(D)J recombination. The Vh-H2/RSS-H3 recombination product has likely been selected during the evolution of mammals because it provides better V(D)J recombination efficiency.

Base Sequence↗

MCALIGN2: faster, accurate global pairwise alignment of non-coding DNA sequences based on explicit models of indel evolution.

BACKGROUND: Non-coding DNA sequences comprise a very large proportion of the total genomic content of mammals, most other vertebrates, many invertebrates, and most plants. Unraveling the functional significance of non-coding DNA depends on how well we are able to align non-coding DNA sequences. However, the alignment of non-coding DNA sequences is more difficult than aligning protein-coding sequences. RESULTS: Here we present an improved pair-hidden-Markov-Model (pair HMM) based method for performing global pairwise alignment of non-coding DNA sequences. The method uses an explicit model of indel length frequency distribution which can be specified, and allows any time reversible model of nucleotide substitution. The method uses a deterministic global optimiser to find the alignment with the highest posterior probability. We test MCALIGN2 in simulations, and compare it to a previous Monte Carlo based method (MCALIGN), to the pair HMM method of Knudsen and Miyamoto, and to a heuristic method (AVID) that performed very well in a previous simulation study. We show that the pair HMM methods have excellent performance for all combinations of parameter values we have considered. MCALIGN2 is up to ten times faster than MCALIGN. MCALIGN2 is more accurate in resolving indels given an accurate explicit model than heuristic methods, but is computationally slower. CONCLUSION: MCALIGN2 produces better quality alignments by explicitly using biological knowledge about the indel length distribution and time reversible models of nucleotide substitution. As a result, it can outperform other available sequence alignment methods for the cases we have considered to align non-coding DNA sequences.

Algorithms↗

Evolution of human polyomavirus JC.

More than 20 near full-length genome sequences have been reported for human polyomavirus JC (JCV). These have previously been classified into seven genotypes, and additional subtypes, which exhibit geographical associations. One of these genotypes, Type 4, has been suggested to be a recombinant of Types 1 and 3. We have investigated the pattern of diversity, and evolutionary relationships, among these sequences. In direct contradiction of a recent report, we found that different phylogenetic methods gave consistent results for the phylogenetic relationships among strains. The single known strain representing Type 5 was shown to be a mosaic of sequences from Types 2 and 6, although whether this recombination occurred in vivo or in vitro is not clear. In contrast, there was no substantial evidence that Type 4 strains are recombinant; rather they seem to be simply divergent examples of Type 1. On the assumption that the major genotypes of JCV diverged with human populations, the rate of synonymous nucleotide substitution was estimated to be around 4x10(-7) per site per year, about 10 times higher than a previous estimate for primate polyomaviruses.

Base Sequence↗

The mammalian alphaD-globin gene lineage and a new model for the molecular evolution of alpha-globin gene clusters at the stem of the mammalian radiation.

We have explored the evolution of the alpha-globin gene family by comparative sequence and phylogenetic analyses of mammalian alpha-globin genes. Our analyses reveal the existence of a new alpha-globin gene lineage in mammals that is related to the alpha(D)-globin genes of birds, squamates and turtles. The gene is located in the middle of the alpha-globin gene cluster of a marsupial, Sminthopsis macroura and of humans. It exists in a wide variety of additional mammals, including pigs, cows, cats, and dogs, but is a pseudogene in American marsupials. Evolutionary analyses suggest that the gene has generally evolved under purifying selection, indicative of a functional gene. The presence of mRNA products in humans, pigs, and cows also suggest that the gene is expressed and likely to be functional. The analyses support the hypothesis that the alpha(D)-globin gene lineage has an ancient evolutionary origin that predates the divergence of amniotes. The structural similarity of alpha-globin gene clusters of marsupials and humans suggest that an eight gene cluster (5'-zeta2-zeta1-alpha(D)-alpha3-alpha2-alpha1-theta-omega-3'), including seven alpha-like genes and one beta-like globin gene (omega-globin) existed in the common ancestor of all marsupial and eutherian mammals. This basic structure has remained relatively stable in marsupials and in the lineage leading to humans, although omega-globin has been lost from the alpha-globin gene cluster of humans.

Animals↗

Gene loss, protein sequence divergence, gene dispensability, expression level, and interactivity are correlated in eukaryotic evolution.

Lineage-specific gene loss, to a large extent, accounts for the differences in gene repertoires between genomes, particularly among eukaryotes. We derived a parsimonious scenario of gene losses for eukaryotic orthologous groups (KOGs) from seven complete eukaryotic genomes. The scenario involves substantial gene loss in fungi, nematodes, and insects. Based on this evolutionary scenario and estimates of the divergence times between major eukaryotic phyla, we introduce a numerical measure, the propensity for gene loss (PGL). We explore the connection among the propensity of a gene to be lost in evolution (PGL value), protein sequence divergence, the effect of gene knockout on fitness, the number of protein-protein interactions, and expression level for the genes in KOGs. Significant correlations between PGL and each of these variables were detected. Genes that have a lower propensity to be lost in eukaryotic evolution accumulate fewer substitutions in their protein sequences and tend to be essential for the organism viability, tend to be highly expressed, and have many interaction partners. The dependence between PGL and gene dispensability and interactivity is much stronger than that for sequence evolution rate. Thus, propensity of a gene to be lost during evolution seems to be a direct reflection of its biological importance.

Amino Acid Substitution↗

Molecular evolution of the Escherichia coli chromosome. IV. Sequence comparisons.

DNA sequences have been compared in a 4,400-bp region for Escherichia coli K12 and 36 ECOR strains. Discontinuities in degree of similarity, previously inferred, are confirmed in detail. Three clonal frames are described on the basis of the present local high-resolution data, as well as previous analyses of restriction fragment length polymorphism (RFLP) and of multilocus enzyme electrophoresis (MLEE) covering small regions more widely dispersed on the chromosome. These three approaches show important consistency. The data illustrate the fact that, in the limited context of intraspecific genomic sequence variation, clonality and homology are synonymous. Two estimable quantitative properties are defined: recency of common ancestry (the reciprocal of the log10 of the number of generations since the most recent common ancestor), and the number of nucleotide pairs over which a given recency of common ancestry applies. In principle, these parameters are measures of the degree and physical extent of homology. The small size of apparent recombinational replacements, together with the observation that they occasionally occur in discontinuous series, raises the question of whether they result from the superimposition of replacements of much larger size (as expected from an elementary interpretation of conjugation and transduction in experimental E. coli systems) or via an alternative mechanism. Length polymorphisms of several sorts are described.

Base Sequence↗

Evolution of antibiotic resistance genes: the DNA sequence of a kanamycin resistance gene from Staphylococcus aureus.

The kanamycin resistance gene from Staphylococcus aureus has been sequenced and its structure compared with similar genes isolated from Streptomyces fradiae and from two transposons, Tn5 and Tn903, originally isolated from Klebsiella pneumoniae and Salmonella typhimurium, respectively. The genes are all homologous but, since their common ancestor, have undergone extensive divergence, with more than 43% divergence between the closest pair. The phylogeny of the genes cannot be made congruent to the phylogeny of the taxa from which they were isolated without requiring rather improbable differences in rates. One is therefore led to conclude that there have been multiple occurrences of gene transfer between these species. Thus, although they are homologous, they are neither orthologous nor paralogous. It is suggested that homologous genes of this type be called xenologous.

Amino Acid Sequence↗

Rapid evolution of a heteroplasmic repetitive sequence in the mitochondrial DNA control region of carnivores.

We describe a repetitive DNA region at the 3' end of the mitochondrial DNA (mtDNA) control region and compare it in 21 carnivore species representing eight carnivore families. The sequence and organization of the repetitive motifs can differ extensively between arrays; however, all motifs appear to be derived from the core motif "ACGT." Sequence data and Southern blot analysis demonstrate extensive heteroplasmy. The general form of the array is similar between heteroplasmic variants within an individual and between individuals within a species (varying primarily in the length of the array, though two clones from the northern elephant seal are exceptional). Within certain families, notably ursids, the array structure is also similar between species. Similarity between species was not apparent in other carnivore families, such as the mustelids, suggesting rapid changes in the organization and sequence of some arrays. The pattern of change seen within and between species suggests that a dominant mechanism involved in the evolution of these arrays is DNA slippage. A comparative analysis shows that the motifs that are being reiterated or deleted vary within and between arrays, suggesting a varying rate of DNA turnover. We discuss the evolutionary implications of the observed patterns of variation and extreme levels of heteroplasmy.

Animals↗

Microsatellite evolution inferred from human-chimpanzee genomic sequence alignments.

Most studies of microsatellite evolution utilize long, highly mutable loci, which are unrepresentative of the majority of simple repeats in the human genome. Here we use an unbiased sample of 2,467 microsatellite loci derived from alignments of 5.1 Mb of genomic sequence from human and chimpanzee to investigate the mutation process of tandemly repetitive DNA. The results indicate that the process of microsatellite evolution is highly heterogeneous, exhibiting differences between loci of different lengths and motif sizes and between species. We find a highly significant tendency for human dinucleotide repeats to be longer than their orthologues in chimpanzees, whereas the opposite trend is observed in mononucleotide repeat arrays. Furthermore, the rate of divergence between orthologues is significantly higher at longer loci, which also show significantly greater mutability per repeat number. These observations have important consequences for understanding the molecular mechanisms of microsatellite mutation and for the development of improved measures of genetic distance.

Animals↗

Evolution of gene position: chromosomal arrangement and sequence comparison of the Drosophila melanogaster and Drosophila virilis sina and Rh4 genes.

The seven in absentia (sina) gene of Drosophila encodes a nuclear protein required for normal eye development. In Drosophila melanogaster, the sina gene is located within an intron of the Rh4 opsin gene. We examine here the nucleotide sequences and chromosomal arrangements of these genes in Drosophila virilis. An interspecies comparison between D. melanogaster and D. virilis reveals that the protein-coding sequences of the sina and Rh4 genes are highly conserved, but the relative chromosomal position and structural arrangement of these genes differ between the two species. In particular, the sina and Rh4 genes are widely separated in D. virilis, and there is no intron in the Rh4 gene. Our results suggest that the Rh4 gene was translocated to another chromosomal location by a retrotransposition event.

Animals↗

Reconsidering the evolution of eukaryotic selenoproteins: a novel nonmammalian family with scattered phylogenetic distribution.

While the genome sequence and gene content are available for an increasing number of organisms, eukaryotic selenoproteins remain poorly characterized. The dual role of the UGA codon confounds the identification of novel selenoprotein genes. Here, we describe a comparative genomics approach that relies on the genome-wide prediction of genes with in-frame TGA codons, and the subsequent comparison of predictions from different genomes, wherein conservation in regions flanking the TGA codon suggests selenocysteine coding function. Application of this method to human and fugu genomes identified a novel selenoprotein family, named SelU, in the puffer fish. The selenocysteine-containing form also occurred in other fish, chicken, sea urchin, green algae and diatoms. In contrast, mammals, worms and land plants contained cysteine homologues. We demonstrated selenium incorporation into chicken SelU and characterized the SelU expression pattern in zebrafish embryos. Our data indicate a scattered evolutionary distribution of selenoproteins in eukaryotes, and suggest that, contrary to the picture emerging from data available so far, other taxa-specific selenoproteins probably exist.

Amino Acid Sequence↗

The human glucocerebrosidase gene and pseudogene: structure and evolution.

We report the sequence of the entire human gene encoding beta-glucocerebrosidase and that of the associated pseudogene. The gene contains 11 exons extending from base pair 355 to base pair 7232 in the overall sequence. The gene promoter contains TATA- and CAT-like boxes upstream of the major 5' end of the glucocerebrosidase RNA. The two TATA boxes lie between nucleotides (-23)-(-27) and (-33)-(-39) and the two possible CAT boxes reside between nucleotides (-90)-(-94) and (-96)-(-99) in relation to the major 5' end of the mRNA. The functionality of the promoter region was monitored by coupling it to the bacterial gene coding for chloramphenicol acetyltransferase (CAT) and assaying the expression of the enzyme in cells transfected with this vector. The glucocerebrosidase promoter not only directs synthesis of the bacterial enzyme but also exhibits the same pattern of tissue-specific expression as that of the endogenous gene. An apparently tightly linked pseudogene is approximately 96% homologous to the functional gene. However, introns 2, 4, 6, and 7 have large "deletions" consisting of Alu sequences 313, 626, 320, and 277 bp in length, respectively. It is entirely possible that the ancestral gene lacks these sequences and that they have been inserted into the introns of the functioning gene. There is also a 55-bp deletion from a part of exon 9 flanked by a short inverted repeat. The sequence data should facilitate development of methods for diagnosis of Gaucher disease at the molecular level.

Base Sequence↗

The complete nucleotide sequence of mouse immunoglobin gamma 2a gene and evolution of heavy chain genes: further evidence for intervening sequence-mediated domain transfer.

We have determined the complete nucleotide sequence (1990 base pairs) of mouse immunoglobulin gamma 2a gene, and compared it with the sequences of other gamma subclass genes so far sequenced, i.e. gamma 1 and gamma 2b genes. Divergence of the nucleotide sequence between a compared pair of the gamma genes varies extensively among different segments of the gene. For example, comparison of the gamma 2a and gamma 2b genes has revealed a remarkable homology in a long continuous segment (about 900 bases) that covers from the 3' portion of the first intervening sequence to the third intervening sequence. However, there is no particular segment of the gamma gene that is conserved universally among the three gamma genes. These findings suggest that, during their evolution, segments of the gamma genes had been scrambled between different subclass genes through recombinations within intervening sequences, thus providing further evidence for the intervening sequence-mediated domain transfer hypothesis. We have discussed several possible phylogenic trees which can explain the difference of divergence in various segments of the gamma genes.

Animals↗

Expression patterns of TEL genes in Poaceae suggest a conserved association with cell differentiation.

Poaceae species present a conserved distichous phyllotaxy (leaf position along the stem) and share common properties with respect to leaf initiation. The goal of this work was to determine if these common traits imply common genes. Therefore, homologues of the maize TERMINAL EAR1 gene in Poaceae were studied. This gene encodes an RNA-binding motif (RRM) protein, that is suggested to regulate leaf initiation. Using degenerate primers, one unique tel (terminal ear1-like) gene from seven Poaceae members, covering almost all the phylogenetic tree of the family, was identified by PCR. These genes present a very high degree of similarity, a much conserved exon-intron structure, and the three RRMs and TEL characteristic motifs. The evolution of tel sequences in Poaceae strongly correlates with the known phylogenetic tree of this family. RT-PCR gene expression analyses show conserved tel expression in the shoot apex in all species, suggesting functional orthology between these genes. In addition, in situ hybridization experiments with specific antisense probes show tel transcript accumulation in all differentiating cells of the leaf, from the recruitment of leaf founder cells to leaf margins cells. Tel expression is not restricted to initiating leaves as it is also found in pro-vascular tissues, root meristems, and immature inflorescences. Therefore, these results suggest that TEL is not only associated with leaf initiation but more generally with cell differentiation in Poaceae.

Amino Acid Sequence↗

Structural classification of thioredoxin-like fold proteins.

Protein structure classification is necessary to comprehend the rapidly growing structural data for better understanding of protein evolution and sequence-structure-function relationships. Thioredoxins are important proteins that ubiquitously regulate cellular redox status and various other crucial functions. We define the thioredoxin-like fold using the structure consensus of thioredoxin homologs and consider all circular permutations of the fold. The search for thioredoxin-like fold proteins in the PDB database identified 723 protein domains. These domains are grouped into eleven evolutionary families based on combined sequence, structural, and functional evidence. Analysis of the protein-ligand structure complexes reveals two major active site locations for the thioredoxin-like proteins. Comparison to existing structure classifications reveals that our thioredoxin-like fold group is broader and more inclusive, unifying proteins from five SCOP folds, five CATH topologies and seven DALI domain dictionary globular folding topologies. Considering these structurally similar domains together sheds new light on the relationships between sequence, structure, function and evolution of thioredoxins.

Amino Acid Motifs↗