PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Essential role of duplications of short motif sequences in the genomic evolution of Bombyx mori.

The Bombyx fibroin gene has a discrete mosaic structure of various repetitive sequences, which may have evolved through various repeating arrangements. Detailed sequence analysis of the fibroin gene containing coding and noncoding regions revealed that the whole sequence could be arranged as an array of short repetitive sequences. A portion of the intron of the fibroin gene is one of interspersed repetitive elements. We cloned a 1.5-kb DNA fragment of the Bombyx genome that contains interspersed elements homologous to the intron sequence. Sequence comparison between the intron and the 1.5-kb fragment shows that partial duplication has frequently occurred in evolutionary progress, and the resultant repetitive blocks of short motif sequences are abundant in the genome. These facts suggest that tandem duplication of the short motif sequence is an important rearrangement in genomic evolution of the fibroin gene.

Animals↗

Comparative genome sequencing of Drosophila pseudoobscura: chromosomal, gene, and cis-element evolution.

We have sequenced the genome of a second Drosophila species, Drosophila pseudoobscura, and compared this to the genome sequence of Drosophila melanogaster, a primary model organism. Throughout evolution the vast majority of Drosophila genes have remained on the same chromosome arm, but within each arm gene order has been extensively reshuffled, leading to a minimum of 921 syntenic blocks shared between the species. A repetitive sequence is found in the D. pseudoobscura genome at many junctions between adjacent syntenic blocks. Analysis of this novel repetitive element family suggests that recombination between offset elements may have given rise to many paracentric inversions, thereby contributing to the shuffling of gene order in the D. pseudoobscura lineage. Based on sequence similarity and synteny, 10,516 putative orthologs have been identified as a core gene set conserved over 25-55 million years (Myr) since the pseudoobscura/melanogaster divergence. Genes expressed in the testes had higher amino acid sequence divergence than the genome-wide average, consistent with the rapid evolution of sex-specific proteins. Cis-regulatory sequences are more conserved than random and nearby sequences between the species--but the difference is slight, suggesting that the evolution of cis-regulatory elements is flexible. Overall, a pattern of repeat-mediated chromosomal rearrangement, and high coadaptation of both male genes and cis-regulatory sequences emerges as important themes of genome divergence between these species of Drosophila.

Animals↗

The evolution of the protein synthesis system, II. From chemical evolution to biological evolution.

The sequence of events previously proposed for modern protein synthesis is reviewed. It begins with an abiological synthesis of a template, and evolves through two model autocatalytic systems to a primitive cell that has a rudimentary biological protein synthesis system. A possible scheme for the origin of tRNA's is described so as to fill the gap between the model and the modern system. Fragments of genes that existed in and around the primitive system are proposed to be precursors of tRNA's. Since these fragments must have been undesirable components for the system, the origin and evolution of tRNA's may be regarded as an excellent answer by the primitive system to adverse circumstances.

Animals↗

Identification of functional domains in the plasma apolipoproteins by analysis of inter-species sequence variability.

Molecular evolution theory posits that sequence motifs essential for protein function are constrained by selective pressure from changing over long stretches of evolutionary time. Thus, analysis of inter-species amino acid sequence variability, by identifying highly conserved intervals, should predict the location of domains critical for protein function. We have analyzed the amino acid sequences of the mammalian apolipoproteins A-I, A-IV, C-I, C-II, C-III, D, and E with a computer algorithm that calculates numerical residue variability scores. The application of a median sieve filter to the data facilitated identification of the exact boundaries of highly conserved domains, which coincided with the location of known structural features and functional domains in this family of proteins. The analysis also identified highly conserved intervals in every apolipoprotein whose function is unknown at present, but which are candidates for regions with specific functional roles.

Amino Acid Sequence↗

The salmon gene encoding apolipoprotein A-I: cDNA sequence, tissue expression and evolution.

A cDNA encoding an apolipoprotein (Apo) has been isolated from the Atlantic salmon (Salmo salar) and sequenced. It encodes a peptide of 258 amino acids (aa), including a signal peptide of 18 aa, with 5'- and 3'-untranslated regions of the mRNA of 12 and 329 nucleotides, respectively. The protein has structural features in common with other Apo's of human and avian origin, including conserved sequences in the signal peptide and a series of internal repeats of 22 aa. The sequence has been identified as salmon Apo A-I (sApoA-I), and has 23% aa identity with human ApoA-I. Northern-blot analysis using the sApoA-I cDNA probe against total RNA prepared from several salmon tissues detects the expression of this gene in liver, intestine and muscle. A phylogenetic analysis reveals that the mammalian ApoA-I, ApoA-IV and Apo-E aa sequences are more closely related to each other than any of them are to sApoA-I. This suggests that the duplication events, from which A-I, A-IV and E arose, occurred after the divergence of the tetrapod and teleost ancestors.

Amino Acid Sequence↗

Evolution of the monotremes. The sequences of the protamine P1 genes of platypus and echidna.

The protamine P1 genes from two monotremes, platypus (Ornithorhynchus anatinus) and echidna (Tachyglossus aculeatus) were isolated after polymerase-chain-reaction amplification then cloned and sequenced. The two protamine P1 genes are of 290 bp and 311 bp for platypus and echidna, respectively, and are clearly orthologous to the published sequences of protamine P1 genes of eutherian mammals and birds. Both genes contain an intron, like the mammals and marsupials and unlike the bird P1 genes that are intronless. The deduced protein sequences from the coding areas of the platypus and echidna protamine P1 genes do not contain any cysteine residues. This absence of cysteine residues leaves the sperm nuclei susceptible to disruption in vitro by exposure to increasing ionic strength and is a characteristic of fish, birds and marsupials. In contrast, the P1 protamines of placental mammals invariably contain 6-9 cysteine residues that, as a result of the formation of intermolecular and intramolecular disulfide bridges, significantly increase the stability of the sperm nuclei that can only be disrupted following disulfide-bond cleavage. Phylogenetic analysis of the protamine P1 gene sequences indicates that the monotremes occupy a position half-way between the eutherian mammals and birds. From the DNA sequences we estimate the time of divergence of the platypus and the echidna to be around 22 million years ago. This date agrees very well with the published estimates of divergence based on other criteria.

Amino Acid Sequence↗

Enterobacterial repetitive intergenic consensus (ERIC) sequences in Escherichia coli: Evolution and implications for ERIC-PCR.

Enterobacterial repetitive intergenic consensus (ERIC) sequences are 127-bp imperfect palindromes that occur in multiple copies in the genomes of enteric bacteria and vibrios. Here we investigate the distribution of these elements in the complete genome sequences of nine Escherichia coli (including Shigella species) strains. There is a significant tendency for copies to be adjacent to more highly expressed genes. There is considerable variation among strains with respect to the presence of an element in any particular intergenic region, but some copies appear to have been conserved since before the divergence of E. coli and Salmonella enterica. In comparisons of orthologous copies between these species, ERIC sequences are surprisingly conserved, implying that they have acquired some function, perhaps related to mRNA stability. The relationships among copies within E. coli are consistent with a master copy mode of generation. Insertion of new copies seems to occur at, and involve duplication of, the dinucleotide TA. Two classes of inserts of about 70 bp each occur at different specific sites within ERIC sequences; these inserts evolve independently of the ERIC sequences. The small number of ERIC sequences in E. coli genomes indicates that a widely used bacterial fingerprinting method using primers based on ERIC sequences (ERIC-PCR) does not rely on the presence of ERIC sequences.

Base Sequence↗

cDNA sequence, protein structure, and evolution of the single hemocyanin from Aplysia californica, an opisthobranch gastropod.

By protein immunobiochemistry and cDNA sequencing, we have found only a single hemocyanin polypeptide in an opisthobranch gastropod, the sea hare Aplysia californica, which contrasts with previously studied prosobranch gastropods, which express two distinct isoforms of this extracellular respiratory protein. We have cloned and sequenced the cDNA encoding the complete polypeptide of Aplysia californica hemocyanin (AcH). The cDNA comprises 11,433 bp, encompassing a 5'UTR of 77 bp, a 3'UTR of 1057 bp, and an open reading frame for a signal peptide of 20 amino acids plus a polypeptide of 3412 amino acids (Mr ca. 387 kDa). This polypeptide is the subunit of the cylindrical native hemocyanin (Mr ca. 8 MDa). It comprises eight different functional units (FUs: a, b, c, d, e, f, g, h) that have been identified immunobiochemically after limited proteolysis of AcH purified from the hemolymph. Each FU shows a highly conserved copper-A and copper-B site for reversible oxygen binding. FU AcH-h carries a specific C-terminal extension of ca. 100 amino acids that include two cysteines that may be utilized for disulfide bridge formation. Potential N-glycosylation sites are present in six FUs but lacking in AcH-b and AcH-c. On the basis of multiple sequence alignments, phylogenetic trees and a statistically firm molecular clock were calculated. The latter suggests that the last common ancestor of Haliotis and Aplysia lived 373+/-47 million years ago, in convincing agreement with fossil records from the early Devonian. However, the gene duplication yielding the two distinct hemocyanin isoforms found today in Haliotis tuberculata occurred 343+/-43 million years ago.

Amino Acid Sequence↗

Environmental effect on the relative contribution of the charge-transfer mechanisms within a short DNA sequence.

Time evolution of the charge-transfer site population is studied in a short DNA sequence to determine the type of governing charge-transfer mechanism. The system consists of a 5'-GAGGG-3' nucleobase sequence coupled with a dissipative bath that represents the DNA phosphate backbone and solvents. Relative contribution of transfer mechanisms to the whole charge-transfer process has been obtained using the on-the-fly filtered propagator functional path integral method with the density matrix decomposition. Partial density matrixes of the incoherent hopping and coherent superexchange pathways as well as the full reduced density matrix have been evaluated and discussed for both debye and ohmic baths. It was found that the relative contribution of the transfer mechanisms is rather sensitive to the frequency-dependent environmental description.

Base Sequence↗

Sequence diversity and molecular evolution of the merozoite surface antigen 2 of Plasmodium falciparum.

Eleven new alleles of the Plasmodium falciparum merozoite surface antigen 2 (MSA2) from Papua New Guinea were analyzed by direct sequencing of polymerase chain reaction (PCR) products. We have used the sequence information to trace the molecular evolution of MSA2. The repeats of ten alleles belonging to the 3D7 allelic family differed considerably in size, nucleotide sequence, and repeat copy number. In the repeat region of these new alleles, codon usage was extremely biased with an exclusive use of NNT codons. Another new allele sequenced belonged to the FC27 family and confirmed the family-specific conserved structure of 96 and 36 bp repeats. In order to assess sequence microheterogeneity within samples defined as the same genotype by restriction fragment length polymorphism (RFLP), we have analyzed single-strand conformation polymorphism (SSCP) of different samples of the most frequent allele (D10 of the FC27 family) in the study population. No sequence heterogeneity could be detected within the repeat region. Based on analysis of the repeat regions in both allelic families, we discuss the hypothesis of a different evolutionary strategy being represented by each of the allelic families. Kew words: Merozoite surface antigen 2 - Nucleotide sequence comparisons - Molecular evolution

Amino Acid Sequence↗

Organization and evolution of repeated DNA sequences in closely related plant genomes.

In common with many other eukaryotic species, the genomes of species in the genus Allium contain a high proportion of repeated DNA sequences, which may be implicated in the considerable differences in genome size that are seen between even very closely related species. The gross organization of repetitive sequences within the genome of Allium sativum and of some other related species has been investigated using DNA/DNA hybridization studies. Such studies show that there has been much modulation in the amounts of different repeated DNA families during the evolution of the genus Allium; these repetitive elements are interspersed in all species with sequences of low repetition. The organization and distribution of one particular repetitive family within the genus has been examined using a cloned hybridization probe. Hybridization of this probe to DNA from related genomes reveals that this element is present in all Allium species examined, but with large-scale modulation of its abundance, and some considerable changes in its sequence environment. The evolution of such genome-specific arrangements of common repetitive elements and the possible mechanisms by which they might be maintained are discussed.

Autoradiography↗

Evolution of maize inferred from sequence diversity of an Adh2 gene segment from archaeological specimens.

A segment of the nuclear gene encoding alcohol dehydrogenase 2 (Adh2) was amplified and sequenced from extracts of archaeological maize specimens up to 4700 years old and from contemporary samples. Sequence diversity in ancient maize equals that of contemporary maize. Some ancient Adh2 alleles are identical or closely related to contemporary alleles. The data suggest that the gene pool of maize is millions of years old and that domestic races of maize stem from several wild ancestral populations.

Alcohol Dehydrogenase↗

Targeting presequence acquisition after mitochondrial gene transfer to the nucleus occurs by duplication of existing targeting signals.

We have cloned a gene for mitochondrial ribosomal protein S11 (RPS11), which is encoded in lower plants by the mitochondrial genome, in higher plants by the nuclear genome, demonstrating genetic information transfer from the mitochondrial genome to the nucleus during flowering plant evolution. The sequence s11-1 encodes an N-terminal extension as well as an organelle-derived RPS11 region. Surprisingly, the N-terminal region has high amino acid sequence similarity with the presequence of the beta-subunit of ATP synthase from plant mitochondria, suggesting a common lineage of the presequences. The deduced N-terminal region of s11-2, a second nuclear-encoded homolog of rps11, shows high sequence similarity with the putative presequence of cytochrome oxidase subunit Vb. The sharing of the N-terminal region together with its 5' flanking untranslated nucleotide sequence in different proteins strongly suggests an involvement of duplication/recombination for targeting signal acquisition after gene migration. A remnant of ancestral rps11 sequence, transcribed and subjected to RNA editing, is found in the mitochondrial genome, indicating that inactivation of mitochondrial rps11 gene expression was initiated at the translational level prior to termination of transcription.

Adenosine Triphosphatases↗

Mulan: multiple-sequence local alignment and visualization for studying function and evolution.

Multiple-sequence alignment analysis is a powerful approach for understanding phylogenetic relationships, annotating genes, and detecting functional regulatory elements. With a growing number of partly or fully sequenced vertebrate genomes, effective tools for performing multiple comparisons are required to accurately and efficiently assist biological discoveries. Here we introduce Mulan (http://mulan.dcode.org/), a novel method and a network server for comparing multiple draft and finished-quality sequences to identify functional elements conserved over evolutionary time. Mulan brings together several novel algorithms: the TBA multi-aligner program for rapid identification of local sequence conservation, and the multiTF program for detecting evolutionarily conserved transcription factor binding sites in multiple alignments. In addition, Mulan supports two-way communication with the GALA database; alignments of multiple species dynamically generated in GALA can be viewed in Mulan, and conserved transcription factor binding sites identified with Mulan/multiTF can be integrated and overlaid with extensive genome annotation data using GALA. Local multiple alignments computed by Mulan ensure reliable representation of short- and large-scale genomic rearrangements in distant organisms. Mulan allows for interactive modification of critical conservation parameters to differentially predict conserved regions in comparisons of both closely and distantly related species. We illustrate the uses and applications of the Mulan tool through multispecies comparisons of the GATA3 gene locus and the identification of elements that are conserved in a different way in avians than in other genomes, allowing speculation on the evolution of birds. Source code for the aligners and the aligner-evaluation software can be freely downloaded from http://www.bx.psu.edu/miller_lab/.

Animals↗

Examination of protein sequence homologies. VI. The evolution of Escherichia coli L7/L12 equivalent ribosomal proteins ('A' proteins), and the tertiary structure.

Sequence homologies among 23 complete and two partial sequences of ribosomal 'A' proteins from eukaryotes, metabacteria, eubacteria and chloroplasts, equivalent to Escherichia coli L7/L12, were examined using a correlation method that evaluates sequence similarity quantitatively. Examination of 325 comparison matrices prepared for possible combinations of the sequences indicates that 'A' protein sequences can be classified into two types: one is the "prototype" from eubacteria and chloroplasts, and the other is the "transposition type" from eukaryotes and metabacteria, which must have resulted from the internal transposition of the prototype sequence. The transposition type of eukaryotes can further be classified into P1 and P2 lines. Sequences of the P1 line are closer to those of metabacteria than to those of the P2 line. Eleven gaps, as deletion or insertion sites of amino acid residues, are necessary for an alignment of all the sequences. According to the crystallographic data for the C-terminal fragment (CTF) from E. coli L7, all the gaps involved in the CTF are located between segments that correspond to structural and functional elements such as alpha helix, beta strand, turning loop or hinge part. The existence of specific "preservation units" in these molecules is suggested. In contrast, the transposition site is located at the center of an alpha helix element that is involved in a folding domain, indicating that the transposition event was extremely drastic.

Amino Acid Sequence↗

Evolution of a B2 tagged sequence from a long-range repeat family in the genus Mus.

A long-range repeat family of more than 50 kb repeat size is clustered in Chromosomes (Chr) 1 of Mus musculus and M. spretus. In M. musculus this long-range repeat family shows considerable variation of copy-number frequency and contains coding regions for at least two genes. In an intron of a gene, which is part of the repeat, a B2 small interspersed repetitive element (SINE) is inserted at identical positions. The B2 element is present in all copies of the long-range repeat family; it was presumably a component of the ancestral single-copy precursor sequence that gave rise by amplification to the repeat family. Copies of the long-range repeat family vary with respect to the number of TAAA tandem repeats in the A-rich 3' end region of the B2 element. As inferred from polymerase chain reaction (PCR) data, presence and frequency of repeat number variants in the (TAAA)n block are strain and species specific. The B2 element and its flanking regions were sequenced from two copies of the long-range repeat family. Sequence divergence between the two copies (only non-CG base substitutions and deletions/insertions) was determined to be 2.6%. Based on the drift rate in human Alu elements and a correction for the higher drift rates in rodents, an estimate for the divergence time of 1.7 million years was calculated. Since the long-range repeat family is present in M. musculus and M. spretus, it must have evolved by amplification before the separation of the two species about 1-4 million years ago.

Animals↗

Rapid evolution of cis-regulatory sequences via local point mutations.

Although the evolution of protein-coding sequences within genomes is well understood, the same cannot be said of the cis-regulatory regions that control transcription. Yet, changes in gene expression are likely to constitute an important component of phenotypic evolution. We simulated the evolution of new transcription factor binding sites via local point mutations. The results indicate that new binding sites appear and become fixed within populations on microevolutionary timescales under an assumption of neutral evolution. Even combinations of two new binding sites evolve very quickly. We predict that local point mutations continually generate considerable genetic variation that is capable of altering gene expression.

Animals↗