PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

The evolution of word composition in metazoan promoter sequence.

The field of molecular evolution provides many examples of the principle that molecular differences between species contain information about evolutionary history. One surprising case can be found in the frequency of short words in DNA: more closely related species have more similar word compositions. Interest in this has often focused on its utility in deducing phylogenetic relationships. However, it is also of interest because of the opportunity it provides for studying the evolution of genome function. Word-frequency differences between species change too slowly to be purely the result of random mutational drift. Rather, their slow pattern of change reflects the direct or indirect action of purifying selection and the presence of functional constraints. Many such constraints are likely to exist, and an important challenge is to distinguish them. Here we develop a method to do so by isolating the effects acting at different word sizes. We apply our method to 2-, 4-, and 8-base-pair (bp) words across several classes of noncoding sequence. Our major result is that similarities in 8-bp word frequencies scale with evolutionary time for regions immediately upstream of genes. This association is present although weaker in intronic sequence, but cannot be detected in intergenic sequence using our method. In contrast, 2-bp and 4-bp word frequencies scale with time in all classes of noncoding sequence. These results suggest that different genomic processes are involved at different word sizes. The pattern in 2-bp and 4-bp words may be due to evolutionary changes in processes such as DNA replication and repair, as has been suggested before. The pattern in 8-bp words may reflect evolutionary changes in gene-regulatory machinery, such as changes in the frequencies of transcription-factor binding sites, or in the affinity of transcription factors for particular sequences.

Amino Acids↗

Pectin degrading glycoside hydrolases of family 28: sequence-structural features, specificities and evolution.

Family 28 belongs to the largest families of glycoside hydrolases. It covers several enzyme specificities of bacterial, fungal, plant and insect origins. This study deals with all available amino acid sequences of family 28 members. First, it focuses on the detailed analysis of 115 sequences of polygalacturonases yielding their evolutionary tree. The large data set allowed modification of some of the existing family 28 sequence characteristics and to draw the sequence features specific for bacterial and fungal exopolygalacturonases discriminating them from the endopolygalacturonases. The evolutionary tree reflects both the taxonomy and specificity so that bacterial, fungal and plant enzymes form their own clusters, the endo- and exo-mode of action being respected, too. The only insect (animal) representative is most related to fungal endopolygalacturonases. The present study brings further: (i) the analysis of available rhamnogalacturonase sequences; (ii) the elucidation of relatedness between the recently added member, the endo-xylogalacturonan hydrolase and the rest of the family; and (iii) revealing the sequence features characteristic of the individual enzyme specificities and the evolutionary relationships within the entire family 28. The disulfides common for the individual enzyme groups were also proposed. With regard to functionally important residues of polygalacturonases, xylogalacturonan hydrolase possesses all of them, while the rhamnogalacturonases, known to lack the histidine residue (His223; Aspergillus niger polygalacturonase II numbering), have a further tyrosine (Tyr291) replaced by a conserved tryptophan. Evolutionarily, the xylogalacturonan hydrolase is most related to fungal exopolygalacturonases and the rhamnogalacturonases form their own cluster on the adjacent branch.

Amino Acid Sequence↗

Identification of essential sequence motifs in the node/notochord enhancer of Foxa2 (Hnf3beta) gene that are conserved across vertebrate species.

The expression of a winged-helix transcription factor, Foxa2/HNF3beta, is essential for development of the node and the notochord. We examined the node/notochord enhancer of mouse Foxa2 for sequence motifs conserved across vertebrate species. We cloned Foxa2 genes from chicken and fish, and identified the respective node/notochord enhancers that were active in transgenic mouse embryos. Comparison of the sequences of the enhancers revealed three evolutionally conserved sequence motifs, CS1, CS2 and CS3. Mutational analysis of the mouse enhancer indicated that CS3 is indispensable for gene expression in the node and the notochord, while CS1 and CS2 are required to augment enhancer activity. These motifs do not correspond to the consensus binding sequences of transcription factors known to be involved in node/notochord development.

Amino Acid Motifs↗

Sequence diversity and molecular evolution of the heat-modifiable outer membrane protein gene (ompA) of Mannheimia(Pasteurella) haemolytica, Mannheimia glucosida, and Pasteurella trehalosi.

The OmpA (or heat-modifiable) protein is a major structural component of the outer membranes of gram-negative bacteria. The protein contains eight membrane-traversing beta-strands and four surface-exposed loops. The genetic diversity and molecular evolution of OmpA were investigated in 31 Mannheimia (Pasteurella) haemolytica, 6 Mannheimia glucosida, and 4 Pasteurella trehalosi strains by comparative nucleotide sequence analysis. The OmpA proteins of M. haemolytica and M. glucosida contain four hypervariable domains located at the distal ends of the surface-exposed loops. The hypervariable domains of OmpA proteins from bovine and ovine M. haemolytica isolates are very different but are highly conserved among strains from each of these two host species. Fourteen different alleles representing four distinct phylogenetic classes, classes I to IV, were identified in M. haemolytica and M. glucosida. Class I, II, and IV alleles were associated with bovine M. haemolytica, ovine M. haemolytica, and M. glucosida strains, respectively, whereas class III alleles were present in certain M. haemolytica and M. glucosida isolates. Class I and II alleles were associated with divergent lineages of bovine and ovine M. haemolytica strains, respectively, indicating a history of horizontal DNA transfer and assortative (entire gene) recombination. Class III alleles have mosaic structures and were derived by horizontal DNA transfer and intragenic recombination. Our findings suggest that OmpA is under strong selective pressure from the host species and that it plays an important role in host adaptation. It is proposed that the OmpA protein of M. haemolytica acts as a ligand and is involved in binding to specific host cell receptor molecules in cattle and sheep. P. trehalosi expresses two OmpA homologs that are encoded by different tandemly arranged ompA genes. The P. trehalosi ompA genes are highly diverged from those of M. haemolytica and M. glucosida, and evidence is presented to suggest that at least one of these genes was acquired by horizontal DNA transfer.

Adaptation, Biological↗

Concurrent neutral evolution of mRNA secondary structures and encoded proteins.

Messenger RNA sequences often have to preserve functional secondary structure elements in addition to coding for proteins. We present a statistical analysis of retroviral mRNA which supports the hypothesis that the natural genetic code is adapted to such complementary coding. These sequences are still able to explore efficiently the space of possible proteins by point mutations. This is borne out by the observation that, in stem regions of retroviral mRNA foldings, silent mutations on one strand are preferentially accompanied by conservative mutations on the other. Distances between amino acids based on physicochemical properties are used to quantify the conservation of protein function under the constraint of maintained RNA secondary structure. We find that preservation of RNA secondary structure by compensatory mutations is evolutionary compatible with the efficient search for new variants on the protein level.

Base Sequence↗

Evolution of alternative splicing: deletions, insertions and origin of functional parts of proteins from intron sequences.

Alternative splicing is thought to be a major source of functional diversity in animal proteins. We analyzed the evolutionary conservation of proteins encoded by alternatively spliced genes and predicted the ancestral state for 73 cases of alternative splicing (25 insertions and 48 deletions). The amino acid sequences of most of the inserts in proteins produced by alternative splicing are as conserved as the surrounding sequences. Thus, alternative splicing often creates novel isoforms by the insertion of new, functional protein sequences that probably originated from noncoding sequences of introns.

Alternative Splicing↗

Evolution of tandemly repeated sequences: What happens at the end of an array?

Tandemly repeated sequences are a major component of the eukaryotic genome. Although the general characteristics of tandem repeats have been well documented, the processes involved in their origin and maintenance remain unknown. In this study, a region on the paternal sex ratio (PSR) chromosome was analyzed to investigate the mechanisms of tandem repeat evolution. The region contains a junction between a tandem array of PSR2 repeats and a copy of the retrotransposon NATE, with other dispersed repeats (putative mobile elements) on the other side of the element. Little similarity was detected between the sequence of PSR2 and the region of NATE flanking the array, indicating that the PSR2 repeat did not originate from the underlying NATE sequence. However, a short region of sequence similarity (11/15 bp) and an inverted region of sequence identity (8 bp) are present on either side of the junction. These short sequences may have facilitated nonhomologous recombination between NATE and PSR2, resulting in the formation of the junction. Adjacent to the junction, the three most terminal repeats in the PSR2 array exhibited a higher sequence divergence relative to internal repeats, which is consistent with a theoretical prediction of the unequal exchange model for tandem repeat evolution. Other NATE insertion sites were characterized which show proximity to both tandem repeats and complex DNAs containing additional dispersed repeats. An "accretion model" is proposed to account for this association by the accumulation of mobile elements at the ends of tandem arrays and into "islands" within arrays. Mobile elements inserting into arrays will tend to migrate into islands and to array ends, due to the turnover in the number of intervening repeats.

Base Sequence↗

Evolution of N-terminal sequences of the vertebrate HOXA13 protein.

While the the role of the homeodomain in HOX function has been evaluated extensively, little attention has been given to the non-homeodomain portions of the HOX proteins. To investigate the evolution of the HOXA13 protein and to identify conserved residues in the N-terminal region of the protein with potential functional significance, N-terminal Hoxa13 coding sequences were PCR-amplified from fish, amphibian, reptile, chicken, and marsupial and eutherian mammal genomic DNA. Compared with fish HOXA13, the mammalian protein has increased in size by 35% primarily owing to the accumulation of alanine repeats and flanking segments rich in proline, glycine, or serine within the first 215 amino acids. Certain residues and amino acid motifs were strongly conserved, and several HOXA13 N-terminal domains were also shared in the paralogous HOXB 13 and HOXD13 genes; however, other conserved regions appear to be unique to HOXA13. Two domains highly conserved in HOXA13 orthologs are shared with Drosophila AbdB and other vertebrate AbdB-like proteins. Marsupial and eutherian mammalian HOXA13 proteins have three large homopolymeric alanine repeats of 14, 12, and 17-18 residues that are absent in reptiles, birds, and fish. Thus, the repeats arose after the divergence of reptiles from the lineage that would give rise to the mammals. In contrast, other short homopolymeric alanine repeats in mammalian HOXA13 have remained virtually the same length, suggesting that forces driving or limiting repeat expansion are context dependent. Consecutive stretches of identical third-base usage in alanine codons within the large repeats were found, supporting replication slippage as a mechanism for their generation. However, numerous species-specific base substitutions affecting third-base alanine repeat codon positions were observed, particularly in the largest repeat. Therefore, if the large alanine repeats were present prior to eutherian mammal development as is suggested by the opossum data, then a dynamic process of recurring replication slippage and point mutation within alanine repeat codons must be considered to reconcile these observations. This model might also explain why the alanine repeats are flanked by proline, serine, and glycine-rich sequences, and it reveals a biological mechanism that promotes increases in protein size and, potentially, acquisition of new functions.

Alanine↗

AdoMet radical proteins--from structure to evolution--alignment of divergent protein sequences reveals strong secondary structure element conservation.

Eighteen subclasses of S-adenosyl-l-methionine (AdoMet) radical proteins have been aligned in the first bioinformatics study of the AdoMet radical superfamily to utilize crystallographic information. The recently resolved X-ray structure of biotin synthase (BioB) was used to guide the multiple sequence alignment, and the recently resolved X-ray structure of coproporphyrinogen III oxidase (HemN) was used as the control. Despite the low 9% sequence identity between BioB and HemN, the multiple sequence alignment correctly predicted all but one of the core helices in HemN, and correctly predicted the residues in the enzyme active site. This alignment further suggests that the AdoMet radical proteins may have evolved from half-barrel structures (alphabeta)4 to three-quarter-barrel structures (alphabeta)6 to full-barrel structures (alphabeta)8. It predicts that anaerobic ribonucleotide reductase (RNR) activase, an ancient enzyme that, it has been suggested, serves as a link between the RNA and DNA worlds, will have a half-barrel structure, whereas the three-quarter barrel, exemplified by HemN, will be the most common architecture for AdoMet radical enzymes, and fewer members of the superfamily will join BioB in using a complete (alphabeta)8 TIM-barrel fold to perform radical chemistry. These differences in barrel architecture also explain how AdoMet radical enzymes can act on substrates that range in size from 10 atoms to 608 residue proteins.

Amino Acid Sequence↗

Polymorphism, shared functions and convergent evolution of genes with sequences coding for polyalanine domains.

Mutations causing expansions of polyalanine domains are responsible for nine hereditary diseases. Other GC-rich sequences coding for some polyalanine domains were found to be polymorphic in human. These observations prompted us to identify all sequences in the human genome coding for polyalanine stretches longer than four alanines and establish their degree of polymorphism. We identified 494 annotated human proteins containing 604 polyalanine domains. Thirty-two percent (31/98) of tested sequences coding for more than seven alanines were polymorphic. The length of the polyalanine-coding sequence and its GCG or GCC repeat content are the major predictors of polymorphism. GCG codons are over-represented in human polyalanine coding sequences. Our data suggest that GCG and GCC codons play a key role in polyalanine-coding sequence appearance and polymorphism. The grouping by shared function of polyalanine-containing proteins in Homo sapiens, Drosophila melanogaster and Caenorhabditis elegans shows that the majority are involved in transcriptional regulation. Phylogenetic analyses of HOX, GATA and EVX protein families demonstrate that polyalanine domains arose independently in different members of these families, suggesting that convergent molecular evolution may have played a role. Finally polyalanine domains in vertebrates are conserved between mammals and are rarer and shorter in Gallus gallus and Danio rerio. Together our results show that the polymorphic nature of sequences coding for polyalanine domains makes them prime candidates for mutations in hereditary diseases and suggests that they have appeared in many different protein families through convergent evolution.

Amino Acid Sequence↗

Sequence heterochrony and the evolution of development.

One of the most persistent questions in comparative developmental biology concerns whether there are general rules by which ontogeny and phylogeny are related. Answering this question requires conceptual and analytic approaches that allow biologists to examine a wide range of developmental events in well-structured phylogenetic contexts. For evolutionary biologists, one of the most dominant approaches to comparative developmental biology has centered around the concept of heterochrony. However, in recent years the focus of studies of heterochrony largely has been limited to one aspect, changes in size and shape. I argue that this focus has restricted the kinds of questions that have been asked about the patterns of developmental change in phylogeny, which has narrowed our ability to address some of the most fundamental questions about development and evolution. Here I contrast the approaches of growth heterochrony with a broader view of heterochrony that concentrates on changes in developmental sequence. I discuss a general approach to sequence heterochrony and summarize newly emerging methods to analyze a variety of kinds of developmental change in explicit phylogenetic contexts. Finally, I summarize a series of studies on the evolution of development in mammals that use these new approaches.

Animals↗

Molecular evolution of centromere-associated nucleotide sequences in two species of canids.

The major centromeric satellite nt sequences present in the domestic dog (Canis familiaris) and in the grey fox (Urocyon cineroargenteus) have been examined. The dog satellite monomer is 737 bp long and contains 51% G + C; the grey fox satellite monomer is 880 b long and contains 54% G + C. The two satellites share three regions of 78, 92 and 314 bp with 70-80% sequence similarity. Sequence data from 16 monomers of dog satellite and 19 monomers of grey fox satellite demonstrate that the substitution spectra are different in the two canid species. For example, substitutions involving G or C residues are much more common in the grey fox satellite than in the domestic dog satellite despite their similar G + C contents.

Animals↗

Strategies for the in vitro evolution of protein function: enzyme evolution by random recombination of improved sequences.

Sets of genes improved by directed evolution can be recombined in vitro to produce further improvements in protein function. Recombination is particularly useful when improved sequences are available; costs of generating such sequences, however, must be weighed against the costs of further evolution by sequential random mutagenesis. Four genes encoding para-nitrobenzyl (pNB) esterase variants exhibiting enhanced activity were recombined in two cycles of high-fidelity DNA shuffling and screening. Genes encoding enzymes exhibiting further improvements in activity were analyzed in order to elucidate evolutionary processes at the DNA level and begin to provide an experimental basis for choosing in vitro evolution strategies and setting key parameters for recombination. DNA sequencing of improved variants from the two rounds of DNA shuffling confirmed important features of the recombination process: rapid fixation and accumulation of beneficial mutations from multiple parent sequences as well as removal of silent and deleterious mutations. The five to sixfold further enhancement of total activity towards the para-nitrophenyl (pNP) ester of loracarbef was obtained through recombination of mutations from several parent sequences as well as new point mutations. Computer simulations of recombination and screening illustrate the trade-offs between recombining fewer parent sequences (in order to reduce screening requirements) and lowering the potential for further evolution. Search strategies which may substantially reduce screening requirements in certain situations are described.

Carboxylic Ester Hydrolases↗

The comparative amino acid sequences, substrate specificities and gene or cDNA nucleotide sequences of some prokaryote and eukaryote amidinotransferases: implications for evolution.

The amino acid sequences of the amidinotransferases and the nucleotide sequences of their genes or cDNA from four Streptomyces species (seven genes) and from the kidneys of rat, pig, human and human pancreas were compared. The overall amino acid and nucleotide sequences of the prokaryotes and eukaryotes were very similar and further, three regions were identified that were highly identical. Evidence is presented that there is virtually zero chance that the overall and high identity regions of the amino acid sequence similarities and the overall nucleotide sequence similarities between Streptomyces and mammals represent random match. Both rat and lamprey amidinotransferases were able to use inosamine phosphate, the amidine group acceptor of Streptomyces. We have concluded that the structure and function of the amidinotransferases and their genes has been highly conserved through evolution from prokaryotes to eukaryotes. The evolution has occurred with: (1) a high degree of retention of nucleotide and amino acid sequences; (2) a high degree of retention of the primitive Streptomyces guanine + cytosine (G + C) third codon position composition in certain high identity regions of the eukaryote cDNA; (3) a decrease in the specificities for the amidine group acceptors; and (4) most of the mutations silent in the regions suggested to code for active sites in the enzymes.

Amidinotransferases↗

The evolution and recognition of protein sequence repeats.

Many proteins sequences contain motifs which display similarity. The similarities between the repeats are a result of gene duplication and/or gene fusion. The evolutionary role of repeats within protein sequences is considered and some repeat examples are given ranging from tandem repeats to multiple types of repeats which are sequentially interspersed. Existing computer methods to delineate repeats in individual protein sequences are discussed and a novel sensitive repeat recognition method is introduced.

Algorithms↗

Phylogenetic analysis of tribe Salsoleae (Chenopodiaceae) based on ribosomal ITS sequences: implications for the evolution of photosynthesis types.

Diversity in photosynthetic pathways in the angiosperm family Chenopodiaceae is expressed in both biochemical and anatomical characters. To understand the evolution of photosynthetic diversity, we reconstructed the phylogeny of representative species of tribe Salsoleae of subfamily Salsoloideae, a group that exhibits in microcosm the patterns of photosynthetic variation present in the family as a whole, and examined the distribution of photosynthetic characters on the resulting phylogenetic tree. Phylogenetic relationships were inferred from parsimony analysis of nucleotide sequences of the internal transcribed spacer regions (ITS) of the 18S-26S nuclear ribosomal DNA of 34 species of Salsola and related genera (Halothamnus, Climacoptera, Girgensohnia, Halocharis, and Haloxylon) and representative outgroups from tribes Camphorosmeae (Camphorosma lessingii, Kochia prostrata, and K. scoparia) and Atripliceae (Atriplex spongiosa). A highly resolved strict consensus tree largely agrees with photosynthetic type and anatomy of leaves and cotyledons. The sequence data provide strong support for the origin and evolution of two main lineages of plants in tribe Salsoleae, with NAD-ME and NADP-ME C(4) photosynthesis, respectively. These groups have different C(4) photosynthetic types in leaves and different structural and photosynthetic characteristics in cotyledons. Phylogenetic relationships inferred from ITS sequences generally agree with classifications based on morphological data, but deviations from the existing taxonomy were also observed. The NAD-ME C(4) lineage contains species classified in sections Caroxylon, Malpigipila, Cardiandra, Belanthera, and Coccosalsola, and the NADP-ME lineage comprises species from sections Coccosalsola and Salsola. Reconstruction of photosynthetic characters on the ITS phylogeny indicates separate NAD-ME and NADP-ME lineages and suggests two reversions to C(3) photosynthesis. Reconstruction of geographic distributions suggests Salsoleae originated and diversified in central Asia and subsequently dispersed to Africa, Europe, and Mongolia. Inferred patterns and processes of photosynthetic evolution in Salsoleae should further our understanding of biochemical and anatomical evolution in Chenopodiaceae as a whole.

Journal Article↗

Involvement of two different urf-s related mitochondrial sequences in the molecular evolution of the CMS-specific S-Pcf locus in petunia.

In petunia, a mitochondrial (mt) locus, S-Pcf, has been found to be strongly associated with cytoplasmic male sterility (CMS). The S-Pcf locus consists of three open reading frames (ORF) that are co-transcribed. The first ORF, Pcf, contains parts of the atp9 and coxII genes and an unidentified reading frame, urf-s. The second and third ORFs contain NADH dehydrogenase subunit 3 (nad3) and ribosomal protein S12 (rps12) sequences, respectively. The nad3 and rps12 sequences included in the S-Pcf locus are identical to the corresponding sequences on the mt genome of fertile petunia. In both CMS and fertile petunia, only a single copy of nad3 and rps12 had been detected on the physical map of the main mt genome. The origin of the urf-s sequence and the molecular events leading to the formation of the chimeric S-Pcf locus are not known. This paper presents evidence indicating that two different mt sequences, related to urf-s and found in fertile petunia lines (orf-h and Rf-1), might have been involved in the molecular evolution of the S-Pcf locus. Southern analysis of mtDNA derived from both fertile and sterile petunia plants suggests that one of these urf-s related sequences (showing 100% homology to urf-s and termed orf-h) is located on a sublimon. An additional, low-homology urf-s related sequence (Rf-1) is shown to be located on the main mt genome 5' to the nad3 gene. It is, thus, suggested that the sequence of events leading to the generation of the S-Pcf locus might have involved introduction of the orf-h sequence, via homologous recombination, into the main mt genome 5' to nad3 at the region where the Rf-1 sequence is located.

Biological Evolution↗