PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Plasmodium falciparum: worldwide sequence diversity and evolution of the malaria vaccine candidate merozoite surface protein-2 (MSP-2).

We examined patterns and putative mechanisms of sequence diversification in the merozoite surface protein-2 (MSP-2) of Plasmodium falciparum, a major dimorphic malaria vaccine candidate antigen, by analyzing 448 msp-2 alleles from all continents. We describe several nucleotide replacements, insertion and deletion events, frameshift mutations, and proliferations of repeat units that generate the extraordinary diversity found in msp-2 alleles. We discuss the role of positive selection exerted by naturally acquired type- and variant-specific immunity in maintaining the observed levels of polymorphism and suggest that this is the most likely explanation for the significant excess of nonsynonymous nucleotide replacements found in dimorphic msp-2 domains. Hybrid sequences created by meiotic recombination between alleles of different dimorphic types were observed in few (3.1%) isolates, mostly from Africa. We found no evidence for an extremely ancient origin of allelic dimorphism at the msp-2 locus, predating P. falciparum speciation, in contrast with recent findings for other surface malarial antigens.

Alleles↗

Understanding relationship between sequence and functional evolution in yeast proteins.

The underlying relationship between functional variables and sequence evolutionary rates is often assessed by partial correlation analysis. However, this strategy is impeded by the difficulty of conducting meaningful statistical analysis using noisy biological data. A recent study suggested that the partial correlation analysis is misleading when data is noisy and that the principal component regression analysis is a better tool to analyze biological data. In this paper, we evaluate how these two statistical tools (partial correlation and principal component regression) perform when data are noisy. Contrary to the earlier conclusion, we found that these two tools perform comparably in most cases. Furthermore, when there is more than one 'true' independent variable, partial correlation analysis delivers a better representation of the data. Employing both tools may provide a more complete and complementary representation of the real data. In this light, and with new analyses, we suggest that protein length and gene dispensability play significant, independent roles in yeast protein evolution.

Evolution, Molecular↗

Cloning, nucleotide sequence and molecular evolution of a rabbit processed metallothionein MT-2 pseudogene.

A rabbit metallothionein-2 pseudogene (MT-2 psi) has been isolated from a partial rabbit genomic library. Its unusual sequence shows evidence of complex rearrangements involving recombination and deletion events. There are no intervening sequences, 3' poly A tract or 5' regulatory DNA sequences. The pseudogene is flanked by two sets of direct repeats (CT)3 GT(CT)4 and CTGG(G)CTC. They are most probably the sites of insertion of MT-2 psi in the rabbit genome. In addition, a number of repetitive DNA sequences are observed flanking the MT-2 psi gene. These are features of a processed retrogene.

Amino Acids↗

The 66-kDa neurofilament protein (NF-66): sequence analysis and evolution.

A 2.5 kb cDNA clone encoding the mouse 66 kd neurofilament protein (NF-66) was isolated and sequenced. The deduced protein sequence contains 501 amino acid residues. Comparison of the mouse, rat and human NF-66 indicated > 90% homology in protein sequence and 85% homology in coding nucleotide sequence. A high degree of homology was observed between NF-66 and other intermediate filament proteins especially in the alpha-helical domain. Zooblot analyses suggested that the putative ancestral gene for vimentin and NF-66 was detectable in the avian. By comparison, the ancestral sequence for GFAP appeared after that for vimentin.

Amino Acid Sequence↗

Sequence simplicity and evolution of the 3' untranslated region of the histone H1o gene.

The H10 gene has a long 3' untranslated region (3'UTR) of 1,125 nucleotides in the rat and 1,310 in humans. Analysis of the sequences shows that they have features of simple DNA that suggest involvement of replication slippage in their evolution. These features include the length imbalance between the rat and human sequences; the abundance of single-base repeats, two-base runs and other simple motifs clustered along the sequence; and the presence of single-base repeat length polymorphisms in the rat and mouse sequences. Pairwise comparisons show numerous short insertions/deletions, often flanked by direct repeats. In addition, a proportion of short insertions/deletions results from length differences in conserved single-base repeats. Quantification of the sequence simplicity shows that simple sequences have been more actively incorporated in the human lineage than in the rodent lineage. The combination of insertions/deletions and nucleotide substitutions along the sequence gives rise to three main regions of homology: a highly variable central region flanked by more conserved regions nearest the coding region and the polyA addition site.

Animals↗

Sequence diversity of Pseudomonas aeruginosa: impact on population structure and genome evolution.

Comparative sequencing of Pseudomonas aeruginosa genes oriC, citS, ampC, oprI, fliC, and pilA in 19 environmental and clinical isolates revealed the sequence diversity to be about 1 order of magnitude lower than in comparable housekeeping genes of Salmonella. In contrast to the low nucleotide substitution rate, the frequency of recombination among different P. aeruginosa genotypes was high, leading to the random association of alleles. The P. aeruginosa population consists of equivalent genotypes that form a net-like population structure. However, each genotype represents a cluster of closely related strains which retain their sequence signature in the conserved gene pool and carry a set of genotype-specific DNA blocks. The codon adaptation index, a quantitative measure of synonymous codon bias of genes, was found to be consistently high in the P. aeruginosa genome irrespective of the metabolic category and the abundance of the encoded gene product. Such uniformly high codon adaptation indices of 0.55 to 0.85 fit the ubiquitous lifestyle of P. aeruginosa.

Biological Evolution↗

Evolution of a "conserved" amino acid sequence: a model study of an in silico investigation of the phylogenesis of some immune receptors.

In this paper we analyze a 55-amino acid (aa) sequence which is relatively well conserved in several seven-transmembrane receptor families (from Insects to Mammals) and in some Viruses. This sequence, which covers the second transmembrane domain, the first extracellular loop and the third transmembrane domain, appears in its complete configuration in most of the seven-transmembrane receptor families, as well as in the protein products of some viruses. Other seven-transmembrane receptors and viruses exhibit reduced configurations of the conserved sequence, lacking either aa 31 or aa 30-31. 53-aa configurations are typically found in most chemokine receptor (CKR) subfamilies, as well as in some viral protein products. However, the CCR1, CCR3, and CCR6 subfamilies comprise a 54-aa configuration and the CKR-related protein products, ChemR23 and RDC1, include the complete 55-aa sequence. For each CKR subfamily the "modal sequence" of the conserved segment was constructed by selecting the most frequently occurring aa at each position. Then, pairwise alignments were made between: (i) the modal CKR sequences, and (ii) the sequence (53-aa) of the Yaba-like disease virus - 7L protein. From the alignments two consensus matrices were derived: (i) the consensus 1 matrix with reference to the whole conserved segment, and (ii) the consensus 2 matrix with reference to aa 22-29, which appear to be the most variable segment of the sequence. Based on the obtained consensus values and with reference to this specific conserved segment, the following conclusions are proposed: (1) ChemR23 and RDC1 are probably the more primitive CKR forms; (2) CCR1 and CCR3 may be grouped in a single cluster; (3) CCRs 2, 4, and 5 are closely related to each other and may be grouped in a cluster; CCR7 is likely to be evolutionarily related to this cluster; (4) CXCRs 2, 3, and 4 and CCX CKR appear to be evolutionarily related to each other and very likely derived from an CCR6-like gene; (5) CCR2/4/5 and CCR7 may have derived either from CCR1/3-like or CCR6-like genes; (6). The Yaba-like disease virus--7L protein most likely derived, through "molecular piracy", from a CCR8-like gene. We also discuss possible, more remote, evolutionary links between CKRs, formylpeptide receptors, and possibly the highly conserved 18S rRNA genes.

Amino Acid Sequence↗

The gene encoding the beta-subunit of rat luteinizing hormone. Analysis of gene structure and evolution of nucleotide sequence.

The nucleotide sequence of the gene encoding the beta-subunit of rat luteinizing hormone (LH beta) has been determined from a genomic DNA fragment cloned in lambda phage Charon 4A. Blot hybridization of restriction enzyme digests of rat genomic DNA indicates that the gene is present in a single copy. The transcriptional unit is 0.98 kilobase in size and contains three exons interrupted by two introns of 245 and 225 base pairs (bp). The locations of the exon/intron junctions at amino acid codons -16/-15 and +41/+42 have been conserved between the rat LH beta gene and the related genes, human LH beta and human chorionic gonadotropin beta. Using S1 nuclease mapping and oligonucleotide-primed reverse transcription of ovariectomized rat pituitary mRNA, the start of transcription was determined to be 7 bp upstream from the start of translation. Characteristic promoter elements are present in the 5'-flanking region of the gene, including the Goldberg-Hogness sequence, TATAAA, 31 bp, and the consensus CAAT box sequence, 167 bp upstream from the start of transcription, respectively. Within the proximal 200 bp flanking the 5'-region of the transcriptional unit, there is strong homology between the rat and human LH beta genes, suggesting that these regions include sequences which may be important for regulation of gene expression. Isolation and characterization of the rat LH beta gene further defines the evolution of glycoprotein hormone genes and will facilitate the study of cellular and molecular mechanisms which regulate LH beta gene expression.

Amino Acid Sequence↗

How many nucleotides are required to resolve a phylogenetic problem? The use of a new statistical method applicable to available sequences.

The evolution of bootstrap proportions (BP) with sequence length was studied using a 28S ribosomal RNA data set. For different sequence lengths, informative sites were jackknifed several times. Bootstrapping was subsequently performed on each of these subsamples. For each node, BPs so obtained were plotted against sequence length, showing the evolution of the robustness with increasing number of informative sites. For robust nodes (BP of 100%), the pattern of BPs is unvarying and is described by a simple function BP = 100 (1-e-b(x-x')), where x is the number of informative sites and b and x' are two parameters estimated using a nonlinear regression procedure. When a node has a BP < 100% and the pattern of BPs fits this function, it is possible to estimate the number of informative sites required to obtain a given average BP. The method also identifies nonrobust nodes (nonascending clusters of BP dots), for which it seems to be more cost effective and fruitful to turn to other species and/or genes rather than to continue sequencing longer gene lengths from the same species to reach a BP of 95%.

Animals↗

Sequence, regulation, and evolution of the maize 22-kD alpha zein gene family.

We have isolated and sequenced all 23 members of the 22-kD alpha zein (z1C) gene family of maize. This is one of the largest plant gene families that has been sequenced from a single genetic background and includes the largest contiguous genomic DNA from maize with 346,292 bp to date. Twenty-two of the z1C members are found in a roughly tandem array on chromosome 4S forming a dense gene cluster 168,489-bp long. The twenty-third copy of the gene family is also located on chromosome 4S at a site approximately 20 cM closer to the centromere and appears to be the wild-type allele of the floury-2 (fl2) mutation. On the basis of an analysis of maize cDNA databases, only seven of these genes appear to be expressed including the fl2 allele. The expressed genes in the cluster are interspersed with nonexpressed genes. Interestingly, some of the expressed genes differ in their transcriptional regulation. Gene amplification appears to be in blocks of genes explaining the rapid and compact expansion of the cluster during the evolution of maize.

Cloning, Molecular↗

Characterization of the genomic Xist locus in rodents reveals conservation of overall gene structure and tandem repeats but rapid evolution of unique sequence.

The Xist locus plays a central role in the regulation of X chromosome inactivation in mammals, although its exact mode of action remains to be elucidated. Evolutionary studies are important in identifying conserved genomic regions and defining their possible function. Here we report cloning, sequence analysis, and detailed characterization of the Xist gene from four closely related species of common vole (field mouse), Microtus arvalis. Our analysis reveals that there is overall conservation of Xist gene structure both between different vole species and relative to mouse and human Xist/XIST. Within transcribed sequence, there is significant conservation over five short regions of unique sequence and also over Xist-specific tandem repeats. The majority of unique sequences, however, are evolving at an unexpectedly high rate. This is also evident from analysis of flanking sequences, which reveals a very high rate of rearrangement and invasion of dispersed repeats. We discuss these results in the context of Xist gene function and evolution.

3' Untranslated Regions↗

Sequence, assembly and evolution of a primordial ferredoxin from Thermotoga maritima.

A gene coding for the ferredoxin of the primordial, strictly anaerobic and hyperthermophilic bacterium Thermotoga maritima was cloned, sequenced and expressed in Escherichia coli. The ferredoxin gene encodes a polypeptide of 60 amino acids that incorporates a single 4Fe-4S cluster. T. maritima ferredoxin expressed in E. coli is a heat-stable, monomeric protein, the spectroscopic properties of which show that its 4Fe-4S cluster is correctly assembled within the mesophilic host, and that it remains stable during purification under aerobic conditions. Removal of the iron-sulfur cluster results in an apo-ferredoxin that has no detectable secondary structure. This observation indicates that in vivo formation of the ferredoxin structure is coupled to the insertion of the iron-sulfur cluster into the polypeptide chain. Sequence comparison of T. maritima ferredoxin with other 4Fe-4S ferredoxins revealed high sequence identities (75% and 50% respectively) to the ferredoxins from the hyperthermophilic members of the Archaea, Thermococcus litoralis and Pyrococcus furiosus. The high sequence similarity supports a close relationship between these extreme thermophilic organisms from different phylogenetic domains and suggests that ferredoxins with a single 4Fe-4S cluster are the primordial representatives of the whole protein family. This observation suggests a new model for the evolution of ferredoxins.

Amino Acid Sequence↗

The evolution of WT1 sequence and expression pattern in the vertebrates.

WT1 is a Wilms' tumour predisposition gene, encoding a protein with four C-terminal Kruppel-type zinc fingers, which is also a major regulator of kidney and gonadal development. To pinpoint key regulatory domains involved in development and evolution of the vertebrate genitourinary system, we have isolated WT1 orthologues from all gnathostome classes. Partial nucleotide sequence from chick, alligator, Xenopus laevis and zebrafish reveals extensive conservation. Both the zinc fingers and the transregulatory domain exhibit a high level of similarity in all the species examined. However, of the two alternatively spliced regions only one, the three amino acid KTS insertion between zinc fingers 3 and 4, is found in species other than mammals. The 17 amino acid insertion at the C-terminal end of the transregulatory domain is present only in mammals. Residues with reported human pathological mutations are also unaltered across species, underlining their structural significance. Studies in chick and alligator reveal that the mammalian intermediate mesoderm expression pattern is conserved in birds and reptiles. A wider role in mesodermal differentiation is suggested by expression in some paraxial and lateral mesoderm derivatives.

Alligators and Crocodiles↗

Molecular evolution of intergenic DNA in higher primates: pattern of DNA changes, molecular clock, and evolution of repetitive sequences.

A 3.1-kb intergenic DNA fragment located between the psi beta-globin and delta-globin genes in the beta-globin gene cluster was cloned from gorilla, orangutan, rhesus monkey, and spider monkey, and the nucleotide sequence of each fragment was determined. The phylogeny of these four sequences, together with two previously published allelic sequences from humans and one from chimpanzee, was constructed, and the accumulation of mutations in the region was analyzed. The sites of base substitutions are not evenly distributed within the region: two Alu repeats have accumulated 0.21 + 0.02 substitutions/site with 0.15 + 0.008 substitutions/site in the remainder of the fragment. The occurrence of substitutions at neighboring sites is more frequent than would be expected if they were independent. The observed excesses disappear when ancestral -CG- dinucleotide sites are excluded. The phylogenetic relationships of the sequences indicate that the human sequence shares a most recent coancestor with the chimpanzee sequence. The data also show that great apes have accumulated fewer mutations in this part of the genome than has the rhesus monkey. The relative rates of accumulation of 12 kinds of nucleotide substitution in the region during primate evolution are asymmetric in the DNA strands. From these rates of accumulation, the origin of a simple stretch of sequence near the 3' end of the 3.1-kb fragment was deduced to be a sequence comprising 50% T and 50% C on one strand. The two oppositely oriented Alu sequences in the 3.1-kb region were inserted at their present positions before the divergence of the New-World monkeys from other lineages. Our analysis shows that the nucleotide sequences of the two Alu repeats in spider monkey are unexpectedly similar both to each other and to the deduced ancestral sequence of Alu repeats. The data suggest that there has been some type of recombinational event between the spider monkey Alu repeats but that it was not a simple gene conversion.

Animals↗

The contribution of LTR retrotransposon sequences to gene evolution in Mus musculus.

Approximately 1.5% of mouse genes (Mus musculus) contain long terminal repeat retrotransposon sequences (LRS). Consistent with earlier findings in Caenorhabditis elegans, Drosophila melanogaster, and Homo sapiens, LRS are more likely to be associated with newly evolved genes. Evidence is presented that LRS are often recruited as novel exons or as spliced additions to existing exons. These novel gene configurations may be expressed initially as alternative transcripts providing an opportunity for the evolution of new gene function.

Animals↗

Fuzzy classification of nucleotide sequences and bacterial evolution.

A new method for reconstructing evolutionary relationship among bacteria by use of rRNA sequence data is proposed. The method is based on the concept of fuzzy classification of probabilities p(i), p(i/j) and p(i/j*) (i = A, G, C, U) of each sequence. The resulting partition tree shares common features of previous works but has some new peculiarities.

Animals↗