PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Sequence variation and evolution of the mitochondrial DNA control region in the musk shrew, Suncus murinus.

The complete mitochondrial DNA (mtDNA) control region was cloned and sequenced in the musk shrew, Suncus murinus, Insectivora. The general aspect was similar to that found in other mammals. We have found in two locations of this region the presence of arrays of tandem repeats like those in other shrew species. One array was located in the left domain containing the termination-associated sequences (TAS) and the length of a copy was 77 bp. The other repeats were situated upstream from the recognition site for the end of H-strand replication in the right domain and were 20 bp long. The left halves of the control region containing the former repeats were sequenced and compared in several laboratory lines and wild animals from different localities, variations in copy number of repeated sequences were found both among individuals and within an individual. A comparative study of repeated sequences provides useful indication for the origin and evolution of tandem repeated sequences. Strand slippage and mispairing during replication of mtDNA with concerted manner is currently regarded as a dominant theory to account molecular mechanism for tandemly repeated sequences, and the pattern of sequence and length variation in our study supports this theory. Our results, however, suggest that the evolution of the repeated sequences containing the TAS in the musk shrew might go through the process of two steps; at the first step one complete repeated and several incomplete repeated sequences had reproduced in common ancestor of the shrew, and the second stage step-up of complete repeated sequences occurred with concerted evolution after differentiation into continental and insular groups.

Animals↗

MHC class I genes in the owl monkey: mosaic organisation, convergence and loci diversity.

The MHC class I molecule plays an important role in immune response, pathogen recognition and response against vaccines and self- versus non-self-recognition. Studying MHC class I characteristics thus became a priority when dealing with Aotus to ensure its use as an animal model for biomedical research. Isolation, cloning and sequencing of exons 1-8 from 27 MHC class I alleles obtained from 13 individuals classified as belonging to three owl monkey species (A. nancymaae, A. nigriceps and A. vociferans) were carried out to establish similarities between Aotus MHC class I genes and those expressed by other New and Old World primates. Six Aotus MHC class I sequence groups (Ao-g1, Ao-g2, Ao-g3, Ao-g4, Ao-g5 and Ao-g6) weakly related to non-classical Catarrhini MHC were identified. An allelic lineage was also identified in one A. nancymaae and two A. vociferans monkeys, exhibiting a high degree of conservation, negative selection along the molecule and premature termination of the open reading frame at exon 5 (Ao-g5). These sequences' high conservation suggests that they more likely correspond to a soluble form of Aotus MHC class I molecules than to a new group of processed pseudogenes. Another group, named Ao-g6, exhibited a strong relationship with Catarrhini's classical MHC-B-C loci. Sequence evolution and variability analysis indicated that Aotus MHC class I molecules experience inter-locus gene conversion phenomena, contributing towards their high variability.

Amino Acid Sequence↗

Correlation between SIV Tat evolution and AIDS progression in cerebrospinal fluid of morphine-dependent and control macaques infected with SIV and SHIV.

Morphine abuse has been associated with higher virus replication and accelerated disease progression in a non-human primate model of AIDS. In our previous report, we have shown that 50% of morphine-addicted macaques progress rapidly and that 2/3 of the rapid progressors exhibit severe neuropathogenesis. In this report, we examined the sequence evolution of the SIV Tat protein, known to participate in AIDS neuropathology, in the cerebrospinal fluid (CSF) of morphine-dependent and control macaques over the first 20 weeks of infection. The CSF SIV Tat evolution was found to be inversely related with disease progression, and the highly neuropathogenic inoculum clone sequence was the prevalent CSF form in rapid progressors. Divergence from the inoculum clone was significantly greater in both morphine-dependent normal progressors and control macaques than in the morphine-dependent rapid progressors. Furthermore, we also found evidence of a trend that morphine alters the type of mutation, resulting in an enhanced ratio of transitions to transversions (Ts:Tv). Rapid disease exacerbates this trend and appears to influence the distribution of nonsynonymous changes in the first exon of SIV tat, with a clear majority of mutations occurring in the C-terminal half of the protein where the known functionally important domains reside. Thus, morphine abuse may change the nature and extent of mutations that drive viral evolution.

Acquired Immunodeficiency Syndrome↗

Nucleotide sequence, polymorphism, and evolution of ovine MHC class II DQA genes.

The nucleotide sequence of all exons and introns, excluding exon 1, of the ovine major histocompatibility complex (MhcOvar) genes analogous to the HLA-DQA1 and -DQA2 genes has been determined and the gene structure found to be similar to that reported for other species. The predicted amino acid sequences of the Ovar-DQA genes have been compared with the equivalent DQA genes in man, mouse, rat, rabbit, and cattle and used to determine the evolutionary relationships of the sheep class II genes to these other species. Northern blot analysis of sheep mRNA using exon specific probes for each of the two Ovar-DQA genes show that both genes are transcribed, whereas in humans there is no evidence that HLA-DQA2 is transcriptionally active. Restriction fragment length polymorphisms (RFLPs) have been used to define a polymorphic series of alleles in both Ovar-DQA genes and have indicated that the number of DQA genes is not constant in sheep as it is in humans, but varies with the haplotype.

Amino Acid Sequence↗

Untangling long branches: identifying conflicting phylogenetic signals using spectral analysis, neighbor-net, and consensus networks.

Long-branch attraction is a well-known source of systematic error that can mislead phylogenetic methods; it is frequently invoked post hoc, upon recovering a different tree from the one expected based on prior evidence. We demonstrate that methods that do not force the data onto a single tree, such as spectral analysis, Neighbor-Net, and consensus networks, can be used to detect conflicting signals within the data, including those caused by long-branch attraction. We illustrate this approach using a set of taxa from three unambiguously monophyletic families within the Pelecaniformes: the darters, the cormorants and shags, and the gannets and boobies. These three families are universally acknowledged as forming a monophyletic group, but the relationship between the families remains contentious. Using sequence data from three mitochondrial genes (12S, ATPase 6, and ATPase 8) we demonstrate that the relationship between these three families is difficult to resolve because they are separated by a short internal branch and there are conflicting signals due to long-branch attraction, which are confounded with nonhomogeneous sequence evolution across the different genes. Spectral analysis, Neighbor-Net, and consensus networks reveal conflicting signals regarding the placement of one of the darters, with support found for darter monophyly, but also support for a conflicting grouping with the outgroup, pelicans. Furthermore, parsimony and maximum-likelihood analyses produced different trees, with one of the two most parsimonious trees not supporting the monophyly of the darters. Monte Carlo simulations, however, were not sensitive enough to reveal long-branch attraction unless the branches are longer than those actually observed. These results indicate that spectral analysis, Neighbor-Net, and consensus networks offer a powerful approach to detecting and understanding the source of conflicting signals within phylogenetic data.

Animals↗

String analysis and energy minimization in the partition of DNA sequences.

Two approaches to the understanding of biological sequences are confronted. While the recognition of particular signals in sequences relies on complex physical interactions, the problem is often analysed in terms of the presence or absence of literal motifs (strings) in the sequence. We present here a test-case for evaluating the potential of this approach. We classify DNA sequences as positive or negative depending on whether they contain a single melted domain in the middle of the sequence, which is a global physical property. Two sets of positive "biological" sequences were generated by a computer simulation of evolutionary divergence along the branches of a phylogenetic tree, under the constraint that each intermediate sequence be positive. These two sets and a set of random positive sequences were subjected to pattern analysis. The observed local patterns were used to construct expert systems to discriminate positive from negative sequences. The experts achieved 79% to 90% success on random positive sequences and up to 99% on the biological sets, while making less than 2% errors on negative sequences. Thus, the global constraints imposed on sequences by a physical process may generate local patterns that are sufficient to predict, with a reasonable probability, the behaviour of the sequences. However, rather large sets of biological sequences are required to generate patterns free of illegitimate constraints. Furthermore, depending upon the initial sequence, the sets of sequences generated on a phylogenetic tree may be amenable or refractory to string analysis, while obeying identical physical constraints. Our study clarifies the relationship between experts' errors on positive and negative sequences, and the contributions of legitimate and illegitimate patterns to these errors. The test-case appears suitable both for further investigations of problems in the theory of sequence evolution and for further testing of pattern analysis techniques.

Base Sequence↗

Role of immunity in maternal-infant HIV-1 transmission.

Factors influencing human immunodeficiency virus type 1 (HIV-1) mother-to-child transmission include both immunological and virological parameters: higher viral loads have been associated with clinical stage of HIV-1-infected individuals as well as higher risk of mother-to-child transmission. Furthermore, we have shown that transmitting mothers more frequently harbour HIV-1 isolates with rapid/high syncytium-inducing (SI) biological phenotype than non-transmitting mothers do. Genetically homogeneous virus populations have been found in HIV-1-infected children at birth, in contrast to the heterogeneous virus populations often found in their infected mothers. This observation suggests that a few virus variants are transmitted or initially are replicating in the child. By comparing the HIV-1 gp120 V3 region of sequentially obtained samples from infected children with samples obtained from their mothers at delivery we found, however, that multiple variants of HIV-1 with different outgrowth kinetics can be transmitted. In addition, we have obtained results indicating an impaired ability of the immune response to adapt to the sequence evolution of HIV-1 in transmitting mothers, as assessed by measuring serum reactivities to peptides representing selected yet closely related V3 sequences. By analysing the presence of antibodies in maternal serum at delivery, which neutralize autologous isolates as well as other primary virus isolates, we have indications that a protective immunity in HIV-1 mother-to-child transmission might exist. Immunotherapy has been assessed in infected adult individuals by passive immunization with a variety of HIV-1-specific antibody products. Data from these studies indicated a differential response to therapy according to the stage of the disease. Active vaccine strategies, including envelope glycoproteins, pursued so far in seronegative adult subjects have shown limitations because broadly neutralizing antibodies, such as can be found in infected individuals, have not been evoked. Further investigations are therefore needed to give support for the potential use of either passive and/or active immunization for the prevention of HIV-1 mother-to-child transmission.

Female↗

Gene duplication, exon gain and neofunctionalization of OEP16-related genes in land plants.

OEP16, a channel protein of the outer membrane of chloroplasts, has been implicated in amino acid transport and in the substrate-dependent import of protochlorophyllide oxidoreductase A. Two major clades of OEP16-related sequences were identified in land plants (OEP16-L and OEP16-S), which arose by a gene duplication event predating the divergence of seed plants and bryophytes. Remarkably, in angiosperms, OEP16-S genes evolved by gaining an additional exon that extends an interhelical loop domain in the pore-forming region of the protein. We analysed the sequence, structure and expression of the corresponding Arabidopsis genes (atOEP16-S and atOEP16-L) and demonstrated that following duplication, both genes diverged in terms of expression patterns and coding sequence. AtOEP16-S, which contains multiple G-box ABA-responsive elements (ABREs) in the promoter region, is regulated by ABI3 and ABI5 and is strongly expressed during the maturation phase in seeds and pollen grains, both desiccation-tolerant tissues. In contrast, atOEP-L, which lacks promoter ABREs, is expressed predominantly in leaves, is induced strongly by low-temperature stress and shows weak induction in response to osmotic stress, salicylic acid and exogenous ABA. Our results indicate that gene duplication, exon gain and regulatory sequence evolution each played a role in the divergence of OEP16 homologues in plants.

Amino Acid Sequence↗

Molecular adaptation in plant hemoglobin, a duplicated gene involved in plant-bacteria symbiosis.

The evolutionary history of the hemoglobin gene family in angiosperms is unusual in that it involves two mechanisms known for potentially generating molecular adaptation: gene duplication and among-species interaction. In plants able to achieve symbiosis with nitrogen-fixing bacteria, class 2 hemoglobin is expressed at high concentrations in nodules and appears to be a key factor for the achievement and regulation of the symbiotic exchange. In this study, we make use of codon models of DNA sequence evolution with the goal of determining the nature of the selective forces which have driven the evolution of this gene. Our results suggest that adaptive evolution occurred during the period of time following the duplication event (functional divergence) and that a change in the selective pressures arose in class 2 hemoglobin in relation to the acquisition of a symbiotic function.

Adaptation, Biological↗

Multiple major increases and decreases in mitochondrial substitution rates in the plant family Geraniaceae.

BACKGROUND: Rates of synonymous nucleotide substitutions are, in general, exceptionally low in plant mitochondrial genomes, several times lower than in chloroplast genomes, 10-20 times lower than in plant nuclear genomes, and 50-100 times lower than in many animal mitochondrial genomes. Several cases of moderate variation in mitochondrial substitution rates have been reported in plants, but these mostly involve correlated changes in chloroplast and/or nuclear substitution rates and are therefore thought to reflect whole-organism forces rather than ones impinging directly on the mitochondrial mutation rate. Only a single case of extensive, mitochondrial-specific rate changes has been described, in the angiosperm genus Plantago. RESULTS: We explored a second potential case of highly accelerated mitochondrial sequence evolution in plants. This case was first suggested by relatively poor hybridization of mitochondrial gene probes to DNA of Pelargonium hortorum (the common geranium). We found that all eight mitochondrial genes sequenced from P. hortorum are exceptionally divergent, whereas chloroplast and nuclear divergence is unexceptional in P. hortorum. Two mitochondrial genes were sequenced from a broad range of taxa of variable relatedness to P. hortorum, and absolute rates of mitochondrial synonymous substitutions were calculated on each branch of a phylogenetic tree of these taxa. We infer one major, approximately 10-fold increase in the mitochondrial synonymous substitution rate at the base of the Pelargonium family Geraniaceae, and a subsequent approximately 10-fold rate increase early in the evolution of Pelargonium. We also infer several moderate to major rate decreases following these initial rate increases, such that the mitochondrial substitution rate has returned to normally low levels in many members of the Geraniaceae. Finally, we find unusually little RNA editing of Geraniaceae mitochondrial genes, suggesting high levels of retroprocessing in their history. CONCLUSION: The existence of major, mitochondrial-specific changes in rates of synonymous substitutions in the Geraniaceae implies major and reversible underlying changes in the mitochondrial mutation rate in this family. Together with the recent report of a similar pattern of rate heterogeneity in Plantago, these findings indicate that the mitochondrial mutation rate is a more plastic character in plants than previously realized. Many molecular factors could be responsible for these dramatic changes in the mitochondrial mutation rate, including nuclear gene mutations affecting the fidelity and efficacy of mitochondrial DNA replication and/or repair and--consistent with the lack of RNA editing--exceptionally high levels of "mutagenic" retroprocessing. That the mitochondrial mutation rate has returned to normally low levels in many Geraniaceae raises the possibility that, akin to the ephemerality of mutator strains in bacteria, selection favors a low mutation rate in plant mitochondria.

Base Sequence↗

Selective escape from CD8+ T-cell responses represents a major driving force of human immunodeficiency virus type 1 (HIV-1) sequence diversity and reveals constraints on HIV-1 evolution.

The sequence diversity of human immunodeficiency virus type 1 (HIV-1) represents a major obstacle to the development of an effective vaccine, yet the forces impacting the evolution of this pathogen remain unclear. To address this issue we assessed the relationship between genome-wide viral evolution and adaptive CD8+ T-cell responses in four clade B virus-infected patients studied longitudinally for as long as 5 years after acute infection. Of the 98 amino acid mutations identified in nonenvelope antigens, 53% were associated with detectable CD8+ T-cell responses, indicative of positive selective immune pressures. An additional 18% of amino acid mutations represented substitutions toward common clade B consensus sequence residues, nine of which were strongly associated with HLA class I alleles not expressed by the subjects and thus indicative of reversions of transmitted CD8 escape mutations. Thus, nearly two-thirds of all mutations were attributable to CD8+ T-cell selective pressures. A closer examination of CD8 escape mutations in additional persons with chronic disease indicated that not only did immune pressures frequently result in selection of identical amino acid substitutions in mutating epitopes, but mutating residues also correlated with highly polymorphic sites in both clade B and C viruses. These data indicate a dominant role for cellular immune selective pressures in driving both individual and global HIV-1 evolution. The stereotypic nature of acquired mutations provides support for biochemical constraints limiting HIV-1 evolution and for the impact of CD8 escape mutations on viral fitness.

Acute Disease↗

Eucaryotic genome evolution through the spontaneous duplication of large chromosomal segments.

There is growing evidence that duplications have played a major role in eucaryotic genome evolution. Sequencing data revealed the presence of large duplicated regions in the genomes of many eucaryotic organisms, and comparative studies have suggested that duplication of large DNA segments has been a continuing process during evolution. However, little experimental data have been produced regarding this issue. Using a gene dosage assay for growth recovery in Saccharomyces cerevisiae, we demonstrate that a majority of the revertant strains (58%) resulted from the spontaneous duplication of large DNA segments, either intra- or interchromosomally, ranging from 41 to 655 kb in size. These events result in the concomitant duplication of dozens of genes and in some cases in the formation of chimeric open reading frames at the junction of the duplicated blocks. The types of sequences at the breakpoints as well as their superposition with the replication map suggest that spontaneous large segmental duplications result from replication accidents. Aneuploidization events or suppressor mutations that do not involve large-scale rearrangements accounted for the rest of the reversion events (in 26 and 16% of the strains, respectively).

Base Sequence↗

A maximum-likelihood approach to analyzing nonoverlapping and overlapping reading frames.

A model is presented for sequence evolution on the basis of which one can analyze combinations of noncoding, singly coding, and multiply coding regions of aligned homologous DNA sequences. It is a generalization of Kimura's (J. Mol. Evol. 16:111-120, 1980) and Li et al.'s (J. Mol. Evol. 36:96-99, 1985) transition-transversion models with selection on replacement substitutions. Based on a hierarchy of hypotheses, one will be able to estimate selection factors and transition and transversion distances for different combinations of regions ranging from many regions, each with their private set of parameters, to one set of parameters for all regions. The method is demonstrated on two aligned HIV1 retroviruses.

Amino Acid Sequence↗

A genetic signature of interspecies variations in gene expression.

Phenotypic diversity is generated through changes in gene structure or gene regulation. The availability of full genomic sequences allows for the analysis of gene sequence evolution. In contrast, little is known about the principles driving the evolution of gene expression. Here we describe the differential transcriptional response of four closely related yeast species to a variety of environmental stresses. Genes containing a TATA box in their promoters show an increased interspecies variability in expression, independent of their functional association. Examining additional data sets, we find that this enhanced expression divergence of TATA-containing genes is consistent across all eukaryotes studied to date, including nematodes, fruit flies, plants and mammals. TATA-dependent regulation may enhance the sensitivity of gene expression to genetic perturbations, thus facilitating expression divergence at particular genetic loci.

DNA↗

Integrated gene and species phylogenies from unaligned whole genome protein sequences.

MOTIVATION: Most molecular phylogenies are based on sequence alignments. Consequently, they fail to account for modes of sequence evolution that involve frequent insertions or deletions. Here we present a method for generating accurate gene and species phylogenies from whole genome sequence that makes use of short character string matches not placed within explicit alignments. In this work, the singular value decomposition of a sparse tetrapeptide frequency matrix is used to represent the proteins of organisms uniquely and precisely as vectors in a high-dimensional space. Vectors of this kind can be used to calculate pairwise distance values based on the angle separating the vectors, and the resulting distance values can be used to generate phylogenetic trees. Protein trees so derived can be examined directly for homologous sequences. Alternatively, vectors defining each of the proteins within an organism can be summed to provide a vector representation of the organism, which is then used to generate species trees. RESULTS: Using a large mitochondrial genome dataset, we have produced species trees that are largely in agreement with previously published trees based on the analysis of identical datasets using different methods. These trees also agree well with currently accepted phylogenetic theory. In principle, our method could be used to compare much larger bacterial or nuclear genomes in full molecular detail, ultimately allowing accurate gene and species relationships to be derived from a comprehensive comparison of complete genomes. In contrast to phylogenetic methods based on alignments, sequences that evolve by relative insertion or deletion would tend to remain recognizably similar.

Algorithms↗

Detecting genomic features under weak selective pressure: the example of codon usage in animals and plants.

Large scale experiments of gene inactivation in yeast have shown that 50% of genes have no detectable impact on the phenotype, and similar observations have been made in other model organisms. This apparent paradox is probably due to the fact that many genes only have a marginal contribution to the fitness of organisms. Because of the size of populations and the number of generations that can be studied in laboratories, experimental approaches only permit to detect functional elements that have a strong phenotypic impact. Comparative sequence analysis can help to solve this problem: the analysis of sequences evolution permits to detect the action of selection, and hence to reveal functional features of genomes. This approach will be illustrated by the study of synonymous codon usage in animals and plants.

Animals↗

GenomeHistory: a software tool and its application to fully sequenced genomes.

We present a publicly available software tool (http://www.unm.edu/~compbio/software/GenomeHistory) that identifies all pairs of duplicate genes in a genome and then determines the degree of synonymous and non-synonymous divergence between each duplicate pair. Using this tool, we analyze the relations between (i) gene function and the propensity of a gene to duplicate and (ii) the number of genes in a gene family and the family's rate of sequence evolution. We do so for the complete genomes of four eukaryotes (fission and budding yeast, fruit fly and nematode) and one prokaryote (Escherichia coli). For some classes of genes we observe a strong relationship between gene function and a gene's propensity to undergo duplication. Most notably, ribosomal genes and transcription factors appear less likely to undergo gene duplication than other genes. In both fission and budding yeast, we see a strong positive correlation between the selective constraint on a gene and the size of the gene family of which this gene is a member. In contrast, a weakly negative such correlation is seen in multicellular eukaryotes.

Animals↗

Phylogenetic analysis using parsimony and likelihood methods.

The assumptions underlying the maximum-parsimony (MP) method of phylogenetic tree reconstruction were intuitively examined by studying the way the method works. Computer simulations were performed to corroborate the intuitive examination. Parsimony appears to involve very stringent assumptions concerning the process of sequence evolution, such as constancy of substitution rates between nucleotides, constancy of rates across nucleotide sites, and equal branch lengths in the tree. For practical data analysis, the requirement of equal branch lengths means similar substitution rates among lineages (the existence of an approximate molecular clock), relatively long interior branches, and also few species in the data. However, a small amount of evolution is neither a necessary nor a sufficient requirement of the method. The difficulties involved in the application of current statistical estimation theory to tree reconstruction were discussed, and it was suggested that the approach proposed by Felsenstein (1981, J. Mol. Evol. 17: 368-376) for topology estimation, as well as its many variations and extensions, differs fundamentally from the maximum likelihood estimation of a conventional statistical parameter. Evidence was presented showing that the Felsenstein approach does not share the asymptotic efficiency of the maximum likelihood estimator of a statistical parameter. Computer simulations were performed to study the probability that MP recovers the true tree under a hierarchy of models of nucleotide substitution; its performance relative to the likelihood method was especially noted. The results appeared to support the intuitive examination of the assumptions underlying MP. When a simple model of nucleotide substitution was assumed to generate data, the probability that MP recovers the true topology could be as high as, or even higher than, that for the likelihood method. When the assumed model became more complex and realistic, e.g., when substitution rates were allowed to differ between nucleotides or across sites, the probability that MP recovers the true topology, and especially its performance relative to that of the likelihood method, generally deteriorates. As the complexity of the process of nucleotide substitution in real sequences is well recognized, the likelihood method appears preferable to parsimony. However, the development of a statistical methodology for the efficient estimation of the tree topology remains a difficult open problem.

Animals↗