PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

Evolutionary constraints on yeast protein size.

BACKGROUND: Despite a strong evolutionary pressure to reduce genome size, proteins vary in length over a surprisingly wide range also in very compact genomes. Here we investigated the evolutionary forces that act on protein size in the yeast Saccharomyces cerevisiae utilizing a system-wide bioinformatics approach. Data on yeast protein size was compared to global experimental data on protein expression, phenotypic pleiotropy, protein-protein interactions, protein evolutionary rate and biochemical classification. RESULTS: Comparing the experimentally determined abundance of individual proteins, highly expressed proteins were found to be consistently smaller than lowly expressed proteins, in accordance with the biosynthetic cost minimization hypothesis. Yeast proteins able to maintain a high expression level despite a large size tended to belong to a very distinct set of protein families, notably nuclear transport and translation initiation/elongation. Large proteins have significantly more protein-protein interactions than small proteins, suggesting that a requirement for multiple interaction domains may constitute a positive selective pressure for large protein size in yeast. The higher frequency of protein-protein interactions in large proteins was not accompanied by a higher phenotypic pleiotropy. Hence, the increase in interactions may not reflect an increase in function differentiation. Proteins of different sizes also evolved at similar rates. Finally, whereas the biological process involved was found to have little influence on protein size the biochemical activity exerted by the protein represented a dominant factor. More than one third of all biochemical activity classes were enriched in one or more size intervals. CONCLUSION: In yeast, there is an inverse relationship between protein size and protein expression such that highly expressed proteins tend to be of smaller size. Also, protein size is moderately affected by protein connectivity and strongly affected by biochemical activity. Phenotypic pleiotropy does not seem to affect protein size.

Amino Acid Sequence↗

Conservation of noncoding microsatellites in plants: implication for gene regulation.

BACKGROUND: Microsatellites are extremely common in plant genomes, and in particular, they are significantly enriched in the 5' noncoding regions. Although some 5' noncoding microsatellites involved in gene regulation have been described, the general properties of microsatellites as regulatory elements are still unknown. To address the question of microsatellites associated with regulatory elements, we have analyzed the conserved noncoding microsatellite sequences (CNMSs) in the 5' noncoding regions by inter- and intragenomic phylogenetic footprinting in the Arabidopsis and Brassica genomes. RESULTS: We identified 247 Arabidopsis-Brassica orthologous and 122 Arabidopsis paralogous CNMSs, representing 491 CT/GA and CTT/GAA repeats, which accounted for 10.6% of these types located in the 500-bp regions upstream of coding sequences in the Arabidopsis genome. Among these identified CNMSs, 18 microsatellites show high conservation in the regulatory regions of both orthologous and paralogous genes, and some of them also appear in the corresponding positions of more distant homologs in Arabidopsis, as well as in other plants. A computational scan of CNMSs for known cis-regulatory elements showed that light responsive elements were clustered in the region of CT/GA repeats, as well as salicylic acid responsive elements in the (CTT)n/(GAA)n sequences. Patterns of gene expression revealed that 70-80% of CNMS (CTT)n/(GAA)n associated genes were regulated by salicylic acid, which was consistent with the prediction of regulatory elements in silico. CONCLUSION: Our analyses showed that some noncoding microsatellites were conserved in plants and appeared to be ancient. These CNMSs served as regulatory elements involved in light and salicylic acid responses. Our findings might have implications in the common features of the over-represented microsatellites for gene regulation in plant-specific pathways.

Arabidopsis↗

Modulation of floral development by a gibberellin-regulated microRNA.

Floral initiation and floral organ development are both regulated by the phytohormone gibberellin (GA). For example, in short-day photoperiods, the Arabidopsis floral transition is strongly promoted by GA-mediated activation of the floral meristem-identity gene LEAFY. In addition, anther development and pollen microsporogenesis depend on GA-mediated opposition of the function of specific members of the DELLA family of GA-response repressors. We describe the role of a microRNA (miR159) in the regulation of short-day photoperiod flowering time and of anther development. MiR159 directs the cleavage of mRNA encoding GAMYB-related proteins. These proteins are transcription factors that are thought to be involved in the GA-promoted activation of LEAFY, and in the regulation of anther development. We show that miR159 levels are regulated by GA via opposition of DELLA function, and that both the sequence of miR159 and the regulation of miR159 levels by DELLA are evolutionarily conserved. Finally, we describe the phenotypic consequences of transgenic over-expression of miR159. Increased levels of miR159 cause a reduction in LEAFY transcript levels, delay flowering in short-day photoperiods, and perturb anther development. We propose that miR159 is a phytohormonally regulated homeostatic modulator of GAMYB activity, and hence of GAMYB-dependent developmental processes.

Agrobacterium tumefaciens↗

Rearrangements in the Cf-9 disease resistance gene cluster of wild tomato have resulted in three genes that mediate Avr9 responsiveness.

Cf resistance genes in tomato confer resistance to the fungal leaf pathogen Cladosporium fulvum. Both the well-characterized resistance gene Cf-9 and the related 9DC gene confer resistance to strains of C. fulvum that secrete the Avr9 protein and originate from the wild tomato species Lycopersicon pimpinellifolium. We show that 9DC and Cf-9 are allelic, and we have isolated and sequenced the complete 9DC cluster of L. pimpinellifolium LA1301. This 9DC cluster harbors five full-length Cf homologs, including orthologs of the most distal homologs of the Cf-9 cluster and three central 9DC genes. Two 9DC genes (9DC1 and 9DC2) have an identical coding sequence, whereas 9DC3 differs at its 3' terminus. From a detailed comparison of the 9DC and Cf-9 clusters, we conclude that the Cf-9 and Hcr9-9D genes from the Cf-9 cluster are ancestral to the first 9DC gene and that the three 9DC genes were generated by subsequent intra- and intergenic unequal recombination events. Thus, the 9DC cluster has undergone substantial rearrangements in the central region, but not at the ends. Using transient transformation assays, we show that all three 9DC genes confer Avr9 responsiveness, but that 9DC2 is likely the main determinant of Avr9 recognition in LA1301.

Cladosporium↗

Mutation and evolution of microsatellite loci in Neurospora.

The patterns of mutation and evolution at 13 microsatellite loci were studied in the filamentous fungal genus Neurospora. First, a detailed investigation was performed on five microsatellite loci by sequencing each microsatellite, together with its nonrepetitive flanking regions, from a set of 147 individuals from eight species of Neurospora. To elucidate the genealogical relationships among microsatellite alleles, repeat number was mapped onto trees constructed from flanking-sequence data. This approach allowed the potentially convergent microsatellite mutations to be placed in the evolutionary context of the less rapidly evolving flanking regions, revealing the complexities of the mutational processes that have generated the allelic diversity conventionally assessed in population genetic studies. In addition to changes in repeat number, frequent substitution mutations within the microsatellites were detected, as were substitutions and insertion/deletions within the flanking regions. By comparing microsatellite and flanking-sequence divergence, clear evidence of interspecific allele length homoplasy and microsatellite mutational saturation was observed, suggesting that these loci are not appropriate for inferring phylogenetic relationships among species. In contrast, little evidence of intraspecific mutational saturation was observed, confirming the utility of these loci for population-level analyses. Frequency distributions of alleles within species were generally consistent with the stepwise mutational model. By comparing variation within species at the microsatellites and the flanking-sequence, estimated microsatellite mutation rates were approximately 2500 times greater than mutation rates of flanking DNA and were consistent with estimates from yeast and fruit flies. A positive relationship between repeat number and variance in repeat number was significant across three genealogical depths, suggesting that longer microsatellite alleles are more mutable than shorter alleles. To test if the observed patterns of microsatellite variation and mutation could be generalized, an additional eight microsatellite loci were characterized and sequenced from a subset of the same Neurospora individuals.

Base Sequence↗

The urea transporter (UT) family: bioinformatic analyses leading to structural, functional, and evolutionary predictions.

We have identified all currently sequenced members of the urea transporter (UT) family (TC #1.A.28). Homologues occur exclusively in vertebrate animals and bacteria but not in other eukaryotic kingdoms or archaea. Sequence, structural, and phylogenetic analyses reveal conserved regions and residues and suggest that a primordial 5 transmembrane helical segment (TMS)-encoding genetic element duplicated to give a 10 TMS-encoding element early during evolutionary history, at about the time when eukaryotes diverged from prokaryotes. Two well-conserved, strongly amphipathic, putative alpha-helices that precede both 5 TMS repeat elements are predicted to be of structural, functional, or biogenic significance. A second duplication event (or a gene fusion event) occurred during development of the vertebrate lineage, giving rise to 20 TMS mammalian homologues. The results suggest that vertebrates acquired UT genetic information from bacteria only once and that all current orthologues and paralogues in the animal kingdom arose from this one primordial system.

Amino Acid Sequence↗

Pronounced intraspecific haplotype divergence at the RPP5 complex disease resistance locus of Arabidopsis.

In Arabidopsis ecotype Landsberg erecta (Ler), RPP5 confers resistance to the pathogen Peronospora parasitica. RPP5 is part of a clustered multigene family encoding nucleotide binding-leucine-rich repeat (LRR) proteins. We compared 95 kb of DNA sequence carrying the Ler RPP5 haplotype with the corresponding 90 kb of Arabidopsis ecotype Columbia (Col-0). Relative to the remainder of the genome, the Ler and Col-0 RPP5 haplotypes exhibit remarkable intraspecific polymorphism. The RPP5 gene family probably evolved by extensive recombination between LRRs from an RPP5-like progenitor that carried only eight LRRs. Most members have variable LRR configurations and encode different numbers of LRRs. Although many members carry retroelement insertions or frameshift mutations, codon usage analysis suggests that regions of the genes have been subject to purifying or diversifying selection, indicating that these genes were, or are, functional. The RPP5 haplotypes thus carry dynamic gene clusters with the potential to adapt rapidly to novel pathogen variants by gene duplication and modification of recognition capacity. We propose that the extremely high level of polymorphism at this complex resistance locus is maintained by frequency-dependent selection.

Amino Acid Sequence↗

U3 snoRNA genes with and without intron in the Kluyveromyces genus: yeasts can accommodate great variations of the U3 snoRNA 3'-terminal domain.

The U3 snoRNA coding sequences from the genomic DNAs of Kluyveromyces delphensis and four variants of the Kluyveromyces marxianus species were cloned by PCR amplification. Nucleotide sequence analysis of the amplification products revealed a unique U3 snoRNA gene sequence in all the strains studied, except for K. marxianus var. fragilis. The K. marxianus U3 genes were intronless, whereas an intron similar to those of the Saccharomyces cerevisiae U3 genes was found in K. delphensis. Hence, U3 genes with and without intron are found in yeasts of the Saccharomycetoideae subfamily. The secondary structure of the K. delphensis pre-U3 snoRNA and of the K. marxianus mature snoRNAs were studied experimentally. They revealed a strong conservation in yeasts of (1) the architecture of U3 snoRNA introns, (2) the 5'-terminal domain of the mature snoRNA, and (3) the protein-anchoring regions of the U3 snoRNA 3' domain. In contrast, stem-loop structures 2, 3, and 4 of the 3' domain showed great variations in size, sequence, and structure. Using a genetic test, we show that, in spite of these variations, the Kluyveromyces U3 snoRNAs are functional in S. cerevisiae. We also show that S. cerevisiae U3A snoRNAs lacking the stem-loop structure 2 or 4 are functional. Hence, U3 snoRNA function can accommodate great variations of the RNA 3'-terminal domain.

Base Sequence↗

Molecular organization of the gene encoding Xenopus laevis transforming growth factor-beta 5.

Transforming growth factor-beta s (TGF-beta s) are multifunctional polypeptides, known to influence proliferation and differentiation of many cell types. TGF-beta 5 cDNA was cloned from Xenopus laevis and this isoform is unique to the amphibians. Here, we report the isolation and characterization of the TGF-beta 5 genomic clones to determine the structure of TGF-beta 5 gene. The gene consists of seven exons and all intron-exon boundaries follow the GT-AG consensus. The organization of TGF-beta 5 gene was identical to that of the mammalian TGF-beta isoforms, with the exception of position of the first splice junction. We determined the size of TGF-beta 5 gene to be approximately 20 kb.

Amino Acid Sequence↗

Cloning, molecular characterization, and distribution of a gene homologous to delta opioid receptor from zebrafish (Danio rerio).

A full-length cDNA, ZFOR1, has been isolated from the teleost zebrafish (Danio rerio) using a probe from rat mu opioid receptor. ZFOR1 encodes a 373 amino acid protein with seven potential transmembrane domains that shows a high degree of homology to mammalian delta opioid receptor. We have also isolated a genomic clone which contains two exons of ZFOR1, homologous to exons 2 and 3 in mouse and human delta opioid receptor. Expression of ZFOR1 appears to be restricted to nervous tissue as assessed by Northern blot. In situ hybridization in zebrafish brain with specific probes revealed several discrete areas of ZFOR1 expression; higher levels are detected in dorsal telencephalic areas, the periventricular layer of the optic tectum, and the granular layer of the cerebellum. This is the first molecular evidence of the presence of the delta opioid receptor in a non-mammalian species and suggests that it has been highly conserved throughout vertebrate evolution.

Amino Acid Sequence↗

Evolution of a perfect simple sequence repeat locus in the context of its flanking sequence.

Microsatellites, which have rapidly become the preferred markers in population genetics, reliably assign individual chinook salmon to the winter, fall, late-fall, or spring chinook runs in the Sacramento River in California's Central Valley (Banks et al. 2000. Can. J. Fish. Aquat. Sci. 57:915-927). A substantial proportion of this discriminatory power comes from Ots-2, a simple CA repeat, which is expected to evolve rapidly under the stepwise mutation model. We have sequenced a 300-bp region around this locus and typed 668 microsatellite-flanking sequence haplotypes to explore further the basis of this microsatellite divergence. Three sites of nucleotide polymorphism in the Ots-2 flanking sequence define five haplotypes that are shared by the Californian and Canadian populations. The Ots-2 microsatellite alleles are nonrandomly distributed among these five haplotypes in a pattern of gametic disequilibrium that is also shared among populations. Divergence between the winter run and other Central Valley stocks appears to be caused by a combination of surprisingly static evolution at Ots-2 within a context of more rapidly changing haplotype frequencies.

Alleles↗

Rapid evolution of the ribonuclease A superfamily: adaptive expansion of independent gene clusters in rats and mice.

The two eosinophil ribonucleases, eosinophil-derived neurotoxin (EDN/RNase 2) and eosinophil cationic protein (ECP/RNase 3), are among the most rapidly evolving coding sequences known among primates. The eight mouse genes identified as orthologs of EDN and ECP form a highly divergent, species-limited cluster. We present here the rat ribonuclease cluster, a group of eight distinct ribonuclease A superfamily genes that are more closely related to one another than they are to their murine counterparts. The existence of independent gene clusters suggests that numerous duplications and diversification events have occurred at these loci recently, sometime after the divergence of these two rodent species ( approximately 10-15 million years ago). Nonsynonymous substitutions per site (d(N)) calculated for the 64 mouse/rat gene pairs indicate that these ribonucleases are incorporating nonsilent mutations at accelerated rates, and comparisons of nonsynonymous to synonymous substitution (d(N) / d(S)) suggest that diversity in the mouse ribonuclease cluster is promoted by positive (Darwinian) selection. Although the pressures promoting similar but clearly independent styles of rapid diversification among these primate and rodent genes remain uncertain, our recent findings regarding the function of human EDN suggest a role for these ribonucleases in antiviral host defense.

Amino Acid Motifs↗

Heterogeneity and evolution rates of delta virus RNA sequences.

To investigate the geographical divergence of delta virus RNA sequences, 868 nucleotides (nt), including the delta antigen-coding region, were determined in isolates from two Japanese patients, M and S, by polymerase chain reaction and direct sequencing and compared with three previously reported nucleotide sequences. The sequence obtained for hepatitis delta virus RNA from patient M was approximately 92% identical to sequences previously obtained for two other strains of hepatitis delta virus, whereas the sequence of hepatitis delta virus RNA obtained from patient S was approximately 81% identical to the previously sequenced strains. This suggests that delta agent in Japan has a heterogeneous origin and the delta virus RNA sequence from Japanese patient S is the most divergent delta virus isolate yet analyzed. To study the evolution rate of delta virus RNA, viral isolates obtained 3 and 4 years apart from each of two patients were also sequenced. It was estimated that the substitution rate of viral RNA was 0.57 x 10(-3) nt per site per year in patient M and 0.64 x 10(-3) nt per site per year in patient S for the delta antigen gene.

Amino Acid Sequence↗

Experimental evolution of conflict mediation between genomes.

Transitions to new levels of biological complexity often require cooperation among component individuals, but individual selection among those components may favor a selfishness that thwarts the evolution of cooperation. Biological systems with elements of cooperation and conflict are especially challenging to understand because the very direction of evolution is indeterminate and cannot be predicted without knowing which types of selfish mutations and interactions can arise. Here, we investigated the evolution of two bacteriophages (f1 and IKe) experimentally forced to obey a life cycle with elements of cooperation and conflict, whose outcome could have ranged from extinction of the population (due to selection of selfish elements) to extreme cooperation. Our results show the de novo evolution of a conflict mediation system that facilitates cooperation. Specifically, the two phages evolved to copackage their genomes into one protein coat, ensuring cotransmission with each other and virtually eliminating conflict. Thereafter, IKe evolved such extreme genome reduction that it lost the ability to make its own virions independent of f1. Our results parallel a variety of conflict mediation mechanisms existing in nature: evolution of reduced genomes in symbionts, cotransmission of partners, and obligate coexistence between cooperating species.

Bacteriophages↗

Computational analysis of the evolution of the structure and function of 1-deoxy-D-xylulose-5-phosphate synthase, a key regulator of the mevalonate-independent pathway in plants.

We investigated molecular evolution of 1-deoxy-D-xylulose-5-phosphate synthase (DXS), an important regulatory enzyme of the mevalonate-independent pathway involved in terpenoid biosynthesis. Sequence alignment showed that some regions, likely to be functionally important, were highly conserved among all of the plant DXS sequences analysed. Phylogenetic trees were inferred using DXS sequences from 11 species of Angiosperms and showed the division of the sequences into two classes. Clustering of DXS sequences did not correspond to phylogenetic relationships among the plant species studied. There was no consistency in the similarity of the variable regions in the secondary structure of the DXS functional protein except for Capsicum and Lycopersicon, both members of the Solanaceae. Hydrophobicity plots for the functional region of DXS revealed great similarity in their hydrophobic structure, consistent with the phylogenetic trees inferred, and with eight prominent hydrophilic and hydrophobic peaks. We also observed a consistent set of features common to the DXS transit peptides studied. These features were the same hydrophobic slope, a hydrophobic region in residues 35-45, and, in eight of 12 sequences, a Pro-Pro-Thr sequence at the C-terminal end. The transit sequences are likely bipartite and contain features that suggest the DXS protein is not only targeted to the chloroplast, but also to the thylakoid. To our knowledge this is the first suggestion that DXS is located specifically in the chloroplast thylakoid.

Amino Acid Sequence↗

Surprising fidelity of template-directed chemical ligation of oligonucleotides.

BACKGROUND: Nucleic acid replication via oligonucleotide ligation has been shown to be extremely prone to errors. If this is the case, it is difficult to envision how the assembly and replication of short oligonucleotides could have contributed to the origin of life and to the evolution of a putative RNA world. In order to assess the fidelity of oligonucleotide replication more accurately, chemical ligation reactions were performed with constant-sequence DNA templates and random-sequence DNA pools as substrates. RESULTS: In keeping with earlier results, constant-sequence hairpin templates were not faithfully copied by random-sequence substrates. Linear templates, however, showed exceptional replication fidelity, particularly when random hexamers were ligated at 25 degrees C. Surprisingly, at low temperatures the formation of G.A base pairs was common and sometimes occurred even more readily than the formation of the corresponding Watson-Crick A-T and G-C base pairs. CONCLUSIONS: The fidelity of ligation reactions increases with temperature and decreases with the length of the random-sequence substrates. Oligonucleotides with a defined sequence can be copied faithfully in the absence of enzymes. Thus, to the extent that short oligonucleotides could readily have been generated by prebiotic mechanisms, it is possible that the earliest self-replicators arose via oligonucleotide ligation.

Base Composition↗

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms↗

Molecular and temporal characteristics of human retropseudogenes.

One of the primary forces driving genome evolution is retrotranscription. In addition to creating new genetic material from which new genes with new functions arise, retrotranscription leaves traces of its action in the form of retropseudogenes. These loci, which are intronless, retrotransposed copies of mature mRNAs from functional antecedent genes, are layered throughout genomes as a molecular fossil record of genome evolution. A survey of 138 functional source genes in the human genome has revealed more than three hundred retropseudogenes. Analysis of the characteristics of the source genes shows that, on average, their size, G/C content, and expression patterns fit the canonical features of source genes reported elsewhere. Details of insertion site duplications for these loci are consistent with a model of retropseudogene formation involving endogenous retrotranscription and enzymatic mobilization and retroposition. Retrotranscription event age estimates reveal a pattern in which the highest densities appear after major phylogenetic events in primate history and then decline. This temporal pattern suggests that the processes forging genome evolution are most active during periods of speciation and adaptive radiation and then steadily diminish until the next burst of activity.

Base Sequence↗