PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

The evolutionary analysis of the Tnt1 retrotransposon in Nicotiana species reveals the high variability of its regulatory sequences.

We studied the evolution of the tobacco Tnt1 retrotransposon by analyzing Tnt1 partial sequences containing both coding domains and U3 regulatory sequences obtained from a number of Nicotiana species. We detected three different subfamilies of Tnt1 elements, Tnt1A, Tnt1B, and Tnt1C, that differ completely in their U3 regions but share conserved flanking coding and LTR regions. U3 divergence between the three subfamilies is found in the region that contains the regulatory sequences that control the expression of the well-characterized Tnt1-94 element. This suggests that expression of the three Tnt1 subfamilies might be differently regulated. The three Tnt1 subfamilies were present in the Nicotiana genome at the time of species divergence, but have evolved independently since then in the different genomes. Each Tnt1 subfamily seems to have conserved its ability to transpose in a limited and different number of Nicotiana species. Our results illustrate the high variability of Tnt1 regulatory sequences. We propose that this high sequence variability could allow these elements to evolve regulatory mechanisms in order to optimize their coexistence with their host genome.

Base Sequence↗

Cyanobacterial evolution: results of 16S ribosomal ribonucleic acid sequence analyses.

We report here the sequences of oligonucleotides released by T1-ribonuclease digestion of the 16S ribosomal RNA's (rRNA's) of unicellular cyanobacteria Agmenellum quadruplicatum (strain BG-1) and Synechococcus 7502. We compare them with sequences previously obtained for the 16S RNA's of six other cyanobacteria and two chloroplasts, and conclude that: (i) Synechocystis-like unicells form a discrete cluster which also (and surprisingly) includes Agmenelium quadruplicatum, usually considered to be a Synechococcus; (ii) filamentous cyanobacteria of the genera Nostoc and Fischerella arose from within the Synechocystis group; (iii) phylogenetic diversity (and hence presumably evolutionary antiquity) within the Synechococcus group is very great; and (iv) red algal chloroplasts are of definite cyanobacterial origin, while Euglena chloroplasts are of separate and quite possibly noncyanobacterial origin. We also present the results of a computer-aided search among the 10 oligonucleotide 'catalogues' for families of related but nonidentical sequences. Examination of these families reinforces the above conclusions.

Base Sequence↗

The evolution of multi-isoacceptor tRNA families. Sequence of tRNA Leu CAA and tRNA Leu CAG from Anacystis nidulans.

Two leucine tRNAs from the cyanophyte Anacystis nidulans have been isolated, and their complete nucleotide sequences have been determined by combining data from oligonucleotide fingerprints and sequencing gels. The two sequences are 87 nucleotides long, have the anticodons CAA and CAG, and differ from each other at a total of 28 positions. They have been compared to other known tRNA Leu sequences and incorporated into a phylogenetic tree comprising prokaryotic and chloroplastic tRNA Leu sequences. Mutations inferred from the tree show that some parts of the tRNA molecule are highly variable (the extra arm and the acceptor stem) while others are much more conserved (the D and T arms). The topology of the tree supports the idea that blue-green algae and chloroplasts share a common prokaryotic ancestor and show a basic divergence between XAA and XAG anticodon-containing tRNAs, suggesting that these two subfamilies result from an ancient gene duplication. Finally, comparison of this phylogenetic tree with those of other multi-isoacceptor tRNA families shows no common scheme, which may be due to independent refinement of codon-reading patterns in different tRNA families.

Base Sequence↗

Genomic and genetic definition of a functional human centromere.

The definition of centromeres of human chromosomes requires a complete genomic understanding of these regions. Toward this end, we report integration of physical mapping, genetic, and functional approaches, together with sequencing of selected regions, to define the centromere of the human X chromosome and to explore the evolution of sequences responsible for chromosome segregation. The transitional region between expressed sequences on the short arm of the X and the chromosome-specific alpha satellite array DXZ1 spans about 450 kilobases and is satellite-rich. At the junction between this satellite region and canonical DXZ1 repeats, diverged repeat units provide direct evidence of unequal crossover as the homogenizing force of these arrays. Results from deletion analysis of mitotically stable chromosome rearrangements and from a human artificial chromosome assay demonstrate that DXZ1 DNA is sufficient for centromere function. Evolutionary studies indicate that, while alpha satellite DNA present throughout the pericentromeric region of the X chromosome appears to be a descendant of an ancestral primate centromere, the current functional centromere based on DXZ1 sequences is the product of the much more recent concerted evolution of this satellite DNA.

Animals↗

Phytochrome evolution: a phylogenetic tree with the first complete sequence of phytochrome from a cryptogamic plant (Selaginella martensii spring).

We have sequenced cDNA and genomic clones coding for phytochrome of the fern Selaginella. On the amino acid level, this phytochrome shares sequence homologies with phytochromes of higher plants which range between 62 (phytochrome B of Arabidopsis) and 55 (56)% [phytochrome C of Arabidopsis (Avena)]. Introns in the Selaginella gene are short and occupy positions known from phytochrome sequences of higher plants. A rooted phylogenetic tree based on mutation distances puts Selaginella phytochrome closest to the hypothetical ancestor. A similar tree arises if the tree is constructed with partial sequences (about 200 amino acids) around the chromophore attachment site. An extension of this tree by sequences of other cryptogamic plants (Mougeotia, Ceratodon, Psilotum) shows all these sequences including those of the phytochromes B and C of Arabidopsis on a branch, well separated from the branch formed by phytochromes known to accumulate in etiolated plants. The rooted phytochrome phylogenetic tree, however, is difficult to reconcile with the fossil record.

Amino Acid Sequence↗

Analysis of the organisation and localisation of the FSHD-associated tandem array in primates: implications for the origin and evolution of the 3.3 kb repeat family.

The D4Z4 locus is a polymorphic tandem repeat sequence on human chromosome 4q35. This locus is implicated in the neuromuscular disorder facioscapulohumeral muscular dystrophy (FSHD). The majority of sporadic cases of FSHD are associated with de novo DNA deletions within D4Z4. However, it is still not known how this rearrangement causes FSHD. Although the repeat contains homeobox sequences, despite exhaustive searching, no transcript from this locus has been identified. Therefore, it has been proposed that the deletion may invoke a position effect on a nearby gene. In order to try to understand the role of the D4Z4 repeat in this disease, we decided to investigate its conservation in other species. In this study, the long-range organisation and localisation of loci homologous to D4Z4 were investigated in primates using Southern blot analysis, pulsed field gel electrophoresis and fluorescence in situ hybridisation. In humans, probes to D4Z4 identify, in addition to the 4q35 locus, a closely related tandem repeat at 10qter and many related repeat loci mapping to the acrocentric chromosomes; a similar pattern was seen in all the great apes. In Old World monkeys, however, only one locus was detected in addition to that on the homologue of human chromosome 4, suggesting that the D4Z4 locus may have originated directly from the progenitor locus. The finding that tandem arrays closely related to D4Z4 have been maintained at loci homologous to human chromosome 4q35-qter in apes and Old World monkeys suggests a functionally important role for these sequences.

Animals↗

Highly condensed potato pericentromeric heterochromatin contains rDNA-related tandem repeats.

The heterochromatin in eukaryotic genomes represents gene-poor regions and contains highly repetitive DNA sequences. The origin and evolution of DNA sequences in the heterochromatic regions are poorly understood. Here we report a unique class of pericentromeric heterochromatin consisting of DNA sequences highly homologous to the intergenic spacer (IGS) of the 18S.25S ribosomal RNA genes in potato. A 5.9-kb tandem repeat, named 2D8, was isolated from a diploid potato species Solanum bulbocastanum. Sequence analysis indicates that the 2D8 repeat is related to the IGS of potato rDNA. This repeat is associated with highly condensed pericentromeric heterochromatin at several hemizygous loci. The 2D8 repeat is highly variable in structure and copy number throughout the Solanum genus, suggesting that it is evolutionarily dynamic. Additional IGS-related repetitive DNA elements were also identified in the potato genome. The possible mechanism of the origin and evolution of the IGS-related repeats is discussed. We demonstrate that potato serves as an interesting model for studying repetitive DNA families because it is propagated vegetatively, thus minimizing the meiotic mechanisms that can remove novel DNA repeats.

Centromere↗

Identity elements of Thermus thermophilus tRNA(Thr).

In this study, we identified nucleotides that specify aminoacylation of tRNA(Thr) by Thermus thermophilus threonyl-tRNA synthetase (ThrRS) using in vitro transcripts. Mutation studies showed that the first base pair in the acceptor stem as well as the second and third positions of the anticodon are major identity elements of T. thermophilus tRNA(Thr), which are essentially the same as those of Escherichia coli tRNA(Thr). The discriminator base, U73, also contributed to the specific aminoacylation, but not the second base pair in the acceptor stem. These findings are in contrast to E. coli tRNA(Thr), where the second base pair is required for threonylation, with the discriminator base, A73, playing no roles. In addition, among several mutations at the third base pair in the acceptor stem, only the G3-U70 mutant was a poor substrate for ThrRS, suggesting that the G3-U70 wobble pair, which is the identity determinant of tRNA(Ala), acts as a negative element for ThrRS. Similar results were obtained in E. coli and yeast. Thus, this manner of rejection of tRNA(Ala) is also likely to have been retained in the threonine system throughout evolution.

Anticodon↗

Evolution on the X chromosome: unusual patterns and processes.

Although the X chromosome is usually similar to the autosomes in size and cytogenetic appearance, theoretical models predict that its hemizygosity in males may cause unusual patterns of evolution. The sequencing of several genomes has indeed revealed differences between the X chromosome and the autosomes in the rates of gene divergence, patterns of gene expression and rates of gene movement between chromosomes. A better understanding of these patterns should provide valuable information on the evolution of genes located on the X chromosome. It could also suggest solutions to more general problems in molecular evolution, such as detecting selection and estimating mutational effects on fitness.

Amino Acid Sequence↗

Assessing the likelihood of recurrence during RNA evolution in vitro.

Recurrence is the possibility of resulting in the same endpoint multiple times when a living system is allowed to evolve repeatedly starting from a given initial point. This concept is of concern to both evolutionary theoreticians and molecular biologists who use nucleic acid selection techniques to mimic biotic and computorial processes in the test tube. Using the continuous in vitro evolution methodology, many replicate experimental evolutionary lineages with populations of catalytic RNA were performed to gain insight into the parameters that could affect recurrence. The likelihood that the same genotype will result in parallel trials of an evolution experiment in vitro depends on several factors, including the phenotype under selection, the size and composition of the initial diverse pool of nucleic acids used in the experiment, the degree of mutation possible during the experiment, the shape of the fitness landscape through which the population evolves, and the strategies used to invoke selection and to search the landscape, among others. By considering these factors, it can be predicted that recurrence is more likely when a small, wild-type-based starting pool is used with efficient selection and search strategies involving little online mutagenesis within a rugged adaptive landscape with a strong local optimum. The recurrence experiments performed here on the 150-nucleotide ligase ribozyme demonstrate that it repeatedly jumps from one peak in a fitness landscape to another, apparently hurdling a deep fitness valley. These predictions can and should be tested by additional multiple replicates of actual evolution experiments in the laboratory.

Base Sequence↗

Human and mouse genomic sequences reveal extensive breakpoint reuse in mammalian evolution.

The human and mouse genomic sequences provide evidence for a larger number of rearrangements than previously thought and reveal extensive reuse of breakpoints from the same short fragile regions. Breakpoint clustering in regions implicated in cancer and infertility have been reported in previous studies; we report here on breakpoint clustering in chromosome evolution. This clustering reveals limitations of the widely accepted random breakage theory that has remained unchallenged since the mid-1980s. The genome rearrangement analysis of the human and mouse genomes implies the existence of a large number of very short "hidden" synteny blocks that were invisible in the comparative mapping data and ignored in the random breakage model. These blocks are defined by closely located breakpoints and are often hard to detect. Our results suggest a model of chromosome evolution that postulates that mammalian genomes are mosaics of fragile regions with high propensity for rearrangements and solid regions with low propensity for rearrangements.

Animals↗

The molecular-cytogenetic analysis of grasses and its application to studying relationships among species of the Triticeae.

An analysis of four species from the genus Secale, including the study of different accessions, has shown that the properties of DNA clones of monomer units from three repeated sequence loci, namely, Ter, Nor, and 5S DNA, proved to be representative of the entire loci from which they were isolated. This finding in Secale species, including the discovery of a new locus for 5S DNA on chromosome 5R, has been used to interpret information on the Ter, Nor, and 5S DNA loci from 15 species in the Triticeae complex. The evolutionary relationship among species suggested by the DNA sequence data has shown many consistencies with a number of other characters such as those used in classical systematics, as well as geographical distribution data and isozyme and chromosome-pairing studies. Apparent inconsistencies such as a close relationship between the R and P genomes at the Ter loci are interpreted in terms of amplification-deletion phenomena known to occur at repetitive sequence loci. In addition, this study included species endemic to Australia and thus provided a broad time span in which to consider some features of repeated sequence family evolution, such as the conservation of certain parts of 5S DNA spacer regions.

Base Sequence↗

Avoidance of inter-repeat recombination by sequence divergence and a mechanism of neutral evolution.

Eucaryotic genomes are loaded with diverse repeated sequences and are therefore threatened by rearrangements via inter-repeat crossovers and by gene-inactivating conversions between genes and their inactive pseudogenes. Such repeated DNA sequences are usually diverged and polymorphic. Sequence divergence by well-spread point mutations is a potent inhibitor of homologous recombination due to the loss of recombination initiation sites and to the editing of recombinational intermediates by the mismatch repair system. Evidence is reviewed suggesting that a germ line process can identify duplicated sequences by homologous pairing, modify them by methylation and mutate by C----T transitions. Since this process requires a minimum contiguous homology that is larger than the average exon size, it is proposed that fragmentation by intron inserts protects the coding sequences from inactivation by homologous interactions with their pseudogene sequences.

Animals↗

Conservation and evolution of (CT)n/(GA)n microsatellite sequences at orthologous positions in diverse mammalian genomes.

The distribution and evolution of (CT)n microsatellites were examined in GenBank mammalian DNA sequences because these microsatellites are known to play important roles in the regulation of some genes in Drosophila melanogaster. A total of 236 (CT)n microsatellite loci were found in GenBank mammalian gene sequences. To determine whether (CT)n microsatellite arrays were conserved at orthologous positions in distantly related mammalian sequences, we determined whether orthologous sequences existed in GenBank for each of the 236 loci. A total of 47 sequence alignments could be made. For rodent x rodent comparisons, 7 of 8 (CT)n arrays were conserved at identical positions in each pair of orthologous sequences. Comparisons of orthologous sequences between different orders of mammals indicated that 11 of 39 (CT)n arrays occurred at orthologous positions or within 1 kb of orthologous positions in each pair of sequences. It appears that there is some level of conservation of (CT)n repeats in distantly related mammals. However, this level of conservation may not be greater than what might be expected to occur by chance. In 13 cases where (CT)n arrays were not conserved at orthologous positions, the lack of a (CT)n array in one sequence resulted from either nucleotide substitution within an array or nonexpansion of a shorter (CT)n element. In these cases, significant sequence identity could be detected throughout the entire region even though the repeat array was not detected in one of the sequences. In contrast, there was a disruption of sequence identity in the (CT)n microsatellite region that ranged from 24 to 1600 bp in 21 cases.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Ab initio folding of proteins using restraints derived from evolutionary information.

We present our predictions in the ab initio structure prediction category of CASP3. Eleven targets were folded, using a method based on a Monte Carlo search driven by secondary and tertiary restraints derived from multiple sequence alignments. Our results can be qualitatively summarized as follows: The global fold can be considered "correct" for targets 65 and 74, "almost correct" for targets 64, 75, and 77, "half-correct" for target 79, and "wrong" for targets 52, 56, 59, and 63. Target 72 has not yet been solved experimentally. On average, for small helical and alpha/beta proteins (on the order of 110 residues or smaller), the method predicted low resolution structures with a reasonably good prediction of the global topology. Most encouraging is that in some situations, such as with target 75 and, particularly, target 77, the method can predict a substantial portion of a rare or even a novel fold. However, the current method still fails on some beta proteins, proteins over the 110-residue threshold, and sequences in which only a poor multiple sequence alignment can be built. On the other hand, for small proteins, the method gives results of quality at least similar to that of threading, with the advantage of not being restricted to known folds in the protein database. Overall, these results indicate that some progress has been made on the ab initio protein folding problem. Detailed information about our results can be obtained by connecting to http:/(/)www.bioinformatics.danforthcenter.org/+ ++CASP3.

Algorithms↗