PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Exon discovery by genomic sequence alignment.

MOTIVATION: During evolution, functional regions in genomic sequences tend to be more highly conserved than randomly mutating 'junk DNA' so local sequence similarity often indicates biological functionality. This fact can be used to identify functional elements in large eukaryotic DNA sequences by cross-species sequence comparison. In recent years, several gene-prediction methods have been proposed that work by comparing anonymous genomic sequences, for example from human and mouse. The main advantage of these methods is that they are based on simple and generally applicable measures of (local) sequence similarity; unlike standard gene-finding approaches they do not depend on species-specific training data or on the presence of cognate genes in data bases. As all comparative sequence-analysis methods, the new comparative gene-finding approaches critically rely on the quality of the underlying sequence alignments. RESULTS: Herein, we describe a new implementation of the sequence-alignment program DIALIGN that has been developed for alignment of large genomic sequences. We compare our method to the alignment programs PipMaker, WABA and BLAST and we show that local similarities identified by these programs are highly correlated to protein-coding regions. In our test runs, PipMaker was the most sensitive method while DIALIGN was most specific. AVAILABILITY: The program is downloadable from the DIALIGN home page at http://bibiserv.techfak.uni-bielefeld.de/dialign/.

Animals↗

Evolution of the ribosomal RNA internal transcribed spacer one (ITS-1) in cichlid fishes of the Lake Victoria region.

The nucleotide sequences of the first internal transcribed spacer (ITS-1) of the ribosomal RNA gene cluster have been determined for 11 species of closely related endemic cichlid fishes of the Lake Victoria region (LVR) and 6 related East African cichlids. The ITS-1 sequences confirmed independently derived basal phylogenies, but provide limited insight within this species flock. The line leading to Pseudocrenilabrus multicolor arose early, close to the divergence event that separated the tilapiine and haplochromine tribes of the "African Group" of the family Cichlidae. In this phylogeny, Astatoreochromis alluaudi and the riverine Astatotilapia burtoni are sister taxa, which together are a sister group to a monophyletic assemblage including both Lake Victoria and Lake Edward taxa. The ITS-1 data support the monophyly of haplochromine genera across lakes. Since Lake Victoria is believed to have been dry between 14, 500 and 12,400 BPE, the modern assemblage must have been derived from reinvasion by the products of earlier cladogenesis events. Thus, although the regional superflock is monophyletic, the haplochromines of Lake Victoria itself did not evolve in situ from a single ancestor.

Africa, Eastern↗

Tempo and mode of Ty element evolution in Saccharomyces cerevisiae.

The Saccharomyces cerevisiae genome contains five families of long terminal repeat (LTR) retrotransposons, Ty1-Ty5. The sequencing of the S. cerevisiae genome provides an unprecedented opportunity to examine the patterns of molecular variation existing among the entire genomic complement of Ty retrotransposons. We report the results of an analysis of the nucleotide and amino acid sequence variation within and between the five Ty element families of the S. cerevisiae genome. Our results indicate that individual Ty element families tend to be highly homogenous in both sequence and size variation. Comparisons of within-element 5' and 3' LTR sequences indicate that the vast majority of Ty elements have recently transposed. Furthermore, intrafamily Ty sequence comparisons reveal the action of negative selection on Ty element coding sequences. These results taken together suggest that there is a high level of genomic turnover of S. cerevisiae Ty elements, which is presumably in response to selective pressure to escape host-mediated repression and elimination mechanisms.

Amino Acid Sequence↗

A second uniquely human mutation affecting sialic acid biology.

Siglecs are immunoglobulin superfamily member lectins that selectively recognize different types and linkages of sialic acids, which are major components of cell surface and secreted glycoconjugates. We report here a human Siglec-like molecule (Siglec-L1) that lacks a conserved arginine residue known to be essential for optimal sialic acid recognition by previously known Siglecs. Loss of the arginine from an ancestral molecule was caused by a single nucleotide substitution that occurred after the common ancestor of humans with the great apes but before the origin of modern humans. The chimpanzee Siglec-L1 ortholog remains fully functional and preferentially recognizes N-glycolylneuraminic acid, which is a common sialic acid in great apes and other mammals. Reintroducing the ancestral arginine into the human molecule regenerates the same properties. Thus, the single base pair mutation that replaced the arginine on human Siglec-L1 is likely to be evolutionarily related to the previously reported loss of N-glycolylneuraminic acid expression in the human lineage. Siglec-L1 and its chimpanzee Siglec ortholog also have a different expression pattern from previously reported Siglecs because they are found on the lumenal edge of epithelial cell surfaces. Notably, the human genome contains several Siglec-like pseudogenes that have independent mutations that would have replaced the arginine residue required for optimal sialic acid recognition. Thus, additional changes in the biology of sialic acids may have taken place during human evolution.

Amino Acid Sequence↗

Interspersed sequence organization and developmental representation of cloned poly(A) RNAs from sea urchin eggs.

A random primed complementary DNA (cDNA) clone library constructed from total maternal poly(A) RNA of sea urchin eggs was screened with two cloned genomic repetitive sequence probes. Sets of cDNA clones reacting with each of these repetitive sequences were recovered. Most of the cloned transcripts included both single copy and repeat sequence elements. Except for the shared repeat sequence element, both the repetitive and single copy regions of the members of each set of clones failed to crossreact. Single copy probes linked to the repeats on the cloned maternal RNAs are represented in an asymmetric manner. It follows that many different genomic members of a given dispersed repeat sequence family are represented in the maternal RNA. RNA gel blots carried out with several repeat probes display about 10 to 20 prominent maternal poly(A) RNAs containing transcripts of each repetitive sequence family. The interspersed maternal transcripts are 3000 to 15,000 bases in length. Maternal transcripts reacting with single copy probes derived from the cloned cDNAs persist during embryonic development, and in some cases appear to be augmented by similar, newly synthesized embryo transcripts. Two examples were found in which additional transcripts of different length appear at specific developmental stages. The transcribed single copy regions are highly polymorphic in the genomes of different individual sea urchins, and comparisons of closely related sea urchin species showed that both the prevalence and length of specific maternal transcripts change rapidly during evolution. Nucleotide sequences of two homologous repeat elements occurring on different cloned transcripts displayed translation stop codons in every possible reading frame. These repeat sequences display structural features suggesting that there has been evolutionary transposition into transcription units active during oogenesis. The repeat elements and their flanking single copy regions reside either in very long 3' or 5'-terminal sequences, or in unprocessed intervening sequences in the maternal poly(A) RNA. These findings lead us to the proposal that the majority of the cytoplasmic poly(A) RNA in echinoderm eggs and early embryos is similar in form to RNAs that occur in the nucleus rather than to the messenger RNA of later cells.

Animals↗

Signatures of adaptive evolution within human non-coding sequence.

The human genome is often portrayed as consisting of three sequence types, each distinguished by their mode of evolution. Purifying selection is estimated to act on 2.5-5.0% of the genome, whereas virtually all remaining sequence is considered to have evolved neutrally and to be devoid of functionality. The third mode of evolution, positive selection of advantageous changes, is considered rare. Such instances have been inferred only for a handful of sites, and these lie almost exclusively within protein-coding genes. Nevertheless, the majority of positively selected sequence is expected to lie within the wealth of functional 'dark matter' present outside of the coding sequence. Here, we review the evolutionary evidence for the majority of human-conserved DNA lying outside of the protein-coding sequence. We argue that within this non-coding fraction lies at least 1 Mb of functional sequence that has accumulated many beneficial nucleotide replacements. Illuminating the functions of this adaptive dark matter will lead to a better understanding of the sequence changes that have shaped the innovative biology of our species.

Evolution, Molecular↗

Sequences of four mouse histone H3 genes: implications for evolution of mouse histone genes.

The sequences of four histone H3 genes coding for the replication variant proteins H3.1 and H3.2 have been determined. Three of these genes, two coding for H3.1 proteins and one for an H3.2 protein, are located on chromosome 13 and expressed at low levels. The fourth gene, encoding an H3.2 protein, is located on chromosome 3 and expressed at a high level. The coding regions of the three genes on chromosome 13 are more similar to each other than to the H3 gene on chromosome 3, and equally divergent from it, suggesting that either gene duplication or gene conversion has occurred since the genes were dispersed onto two chromosomes. A 14-base sequence including the CCAAT sequence and located 5' to the genes on chromosome 13 has been conserved. The histone H3 gene on chromosome 3 has multiple potential binding sites for the Sp1 transcription factor. The coding regions show greater than 95% conservation among the four genes. This is due to the strict pattern of codon usage and the presence of two long (greater than 60 base) regions of completely conserved nucleic acid sequence. These conserved regions in the coding sequence may have an important functional role at the mRNA or DNA level.

Amino Acid Sequence↗

Nucleotide sequence divergence and functional constraint in mRNA evolution.

Comparison of about 50 pairs of homologous nucleotide sequences for different genes revealed that the substitutions between synonymous codons occurred at much higher rates than did amino acid substitutions. Furthermore, five pairs of mRNA sequences for different genes were compared in species that had diverged at the same time. The evolutionary rate of synonymous substitution was estimated to be 5.1 X 10(-9) per site per year on the average and is approximately constant among different genes. It also is suggested that this property would be suitable for a molecular clock to determine the evolutionary relationships and branching order of duplicated genes. Each functional block of the noncoding region evolves with a rate that is almost constant, regardless of the types of genes. The intervening sequence and the 5' portion of the 3' noncoding region show considerable divergence, the extent of which is almost comparable to that in the synonymous codon sites, whereas the other blocks consisting of the 5' noncoding region and the 3' portion of the 3' noncoding region are strongly conserved, showing approximatley half of the divergence of the synonymous sites. This strong sequence preservation might be due to the functional requirements for transcription and modification of mRNA.

Amino Acid Sequence↗

Phylogeny and evolution of antlered deer determined from mitochondrial DNA sequences.

Mitochondrial DNA sequences of both ribosomal RNA genes and three adjacent transfer RNA genes were obtained for the three extant subfamilies of antlered deer (Cervinae, Muntiacinae, and Odocoileinae) as well as for their antlerless sister group Hydropotinae (family Cervidae). Phylogenetic analysis of these sequences (each nearly 2.7 kilobase pairs in length) supports a cervine/muntiacine clade to the exclusion of odocoileines. These results are statistically significant, stable, and congruent with some independent data. Our mitochondrial DNA sequences, when coupled with other information, indicate that the earliest fossil antlered deer are not closely related to living muntiacines or any other contemporary subfamily. From this information, we hypothesize an Old World, Late Miocene origin of Odocoileinae.

Animals↗

Evolutionary conservation of protein regions in the protonmotive cytochrome b and their possible roles in redox catalysis.

The amino acid sequences of the protonmotive cytochrome b from seven representative and phylogenetically diverse species have been compared to identify protein regions or segments that are conserved during evolution. The sequences analyzed included both prokaryotic and eukaryotic examples as well as mitochondrial cytochrome b and chloroplast b6 proteins. The principal conclusion from these analyses is that there are five protein regions--each comprising about 20 amino acid residues--that are consistently conserved during evolution. These domains are evident despite the low density of invariant residues. The two most highly conserved regions, spanning approximately consensus residues 130-150 and 270-290, are located in extramembrane loops and are hypothesized to constitute part of the Qo reaction center. The intramembrane, hydrophobic protein regions containing the heme-ligating histidines are also conserved during evolution. It was found, however, that the conservation of the protein segments extramembrane to the histidine residues ligating the low potential b566 heme group showed a higher degree of sequence conservation. The location of these conserved regions suggests that these extramembrane segments are also involved in forming the Qo reaction center. A protein segment putatively constituting a portion of the Qi reaction center, located approximately in the region spanned by consensus residues 20-40, is conserved in species as divergent as mouse and Rhodobacter. This region of the protein shows substantially less sequence conservation in the chloroplast cytochrome b6. The catalytic role of these conserved regions is strongly supported by locations of residues that are altered in mutants resistant to inhibitors of cytochrome b electron transport.

Amino Acid Sequence↗

Molecular identification of Sphingomonas sp. A1 alginate lyase (A1-IV') as a member of novel polysaccharide lyase family 15 and implications in alginate lyase evolution.

Sphingomonas sp. A1 (strain A1) produces three endotypes (A1-I [65 kDa], A1-II [25 kDa], and A1-III [40 kDa]) and an exotype (A1-IV [86 kDa]) alginate lyases in cytoplasm. These four enzymes cooperatively depolymerize alginate into constituent monosaccharides. In addition to the genes for these lyases, novel genes encoding hypothetical proteins homologous with A1-IV were found in the genomes of many bacteria including strain A1. One such protein, A1-IV' (90 kDa) of strain A1, was overexpressed in Escherichia coli cells, purified, and characterized. A1-IV' catalyzed the cleavage of glycosidic bonds in alginate through a beta-elimination reaction and released unsaturated di- and trisaccharides as main products, thus indicating that the enzyme is an endotype alginate lyase. A1-IV', which differed from A1-IV in some enzymatic properties, was not expressed in strain A1, suggesting that A1-IV' has no significant role in alginate metabolism. A1-IV' and other A1-IV homologs facilitate the creation of novel polysaccharide lyase family 15 based on their primary structures, implying the evolution route of alginate lyases in family PL-15.

Amino Acid Sequence↗

Exploring protein sequence space using knowledge-based potentials.

Knowledge-based potentials can be used to decide whether an amino acid sequence is likely to fold into a prescribed native protein structure. We use this idea to survey the sequence-structure relations in protein space. In particular, we test the following two propositions which were found to be important for efficient evolution: the sequences folding into a particular native fold form extensive neutral networks that percolate through sequence space. The neutral networks of any two native folds approach each other to within a few point mutations. Computer simulations using two very different potential functions, M. Sippl's PROSA pair potential and a neural network based potential, are used to verify these claims.

Amino Acid Sequence↗

Molecular evolution from abiotic scratch.

Recent papers on the emerging new theory of protein evolution are reviewed. Reconstruction of codon chronology, analysis of loop fold structure of proteins, and quantitative correspondence between optimal DNA ring closure size and protein domain size allow to outline specific stages in early protein evolution, each with its own size range.

Amino Acid Sequence↗

New approaches to the analysis of palindromic sequences from the human genome: evolution and polymorphism of an intronic site at the NF1 locus.

The nature of any long palindrome that might exist in the human genome is obscured by the instability of such sequences once cloned in Escherichia coli. We describe and validate a practical alternative to the analysis of naturally-occurring palindromes based upon cloning and propagation in Saccharomyces cerevisiae. With this approach we have investigated an intronic sequence in the human Neurofibromatosis 1 (NF1) locus that is represented by multiple conflicting versions in GenBank. We find that the site is highly polymorphic, exhibiting different degrees of palindromy in different individuals. A side-by-side comparison of the same plasmids in E.coli versus. S.cerevisiae demonstrated that the more palindromic alleles were inevitably corrupted upon cloning in E.coli, but could be propagated intact in yeast. The high quality sequence obtained from the yeast-based approach provides insight into the various mechanisms that destabilize a palindrome in E.coli, yeast and humans, into the diversification of a highly polymorphic site within the NF1 locus during primate evolution, and into the association between palindromy and chromosomal translocation.

Alleles↗

Sequence determinants of function and evolution in serine proteases.

Serine proteases of the chymotrypsin family have maintained a common fold over an evolutionary span of more than one billion years. Notwithstanding modest changes in sequence, this class of enzymes has developed a wide variety of substrate specificities and important biological functions such as fibrinolysis, blood coagulation, and complement activation. Recently it has become apparent that the protease domain, especially its C-terminal sequence, accounts fully for this functional diversity and is the most important element in shaping serine protease evolution.

Animals↗