PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Serial NetEvolve: a flexible utility for generating serially-sampled sequences along a tree or recombinant network.

UNLABELLED: Serial NetEvolve is a flexible simulation program that generates DNA sequences evolved along a tree or recombinant network. It offers a user-friendly Windows graphical interface and a Windows or Linux simulator with a diverse selection of parameters to control the evolutionary model. Serial NetEvolve is a modification of the Treevolve program with the following additional features: simulation of serially-sampled data, the choice of either a clock-like or a variable rate model of sequence evolution, sampling from the internal nodes and the output of the randomly generated tree or network in our newly proposed NeTwick format. AVAILABILITY: From website http://biorg.cis.fiu.edu/SNE Contacts: giri@cis.fiu.edu SUPPLEMENTARY INFORMATION: Manual and examples available from http://biorg.cis.fiu.edu/SNE.

Algorithms↗

Molecular footprints of human immunoglobulin gene evolution: a new sequence family.

Analysis of the human VK (ref. 2) gene locus led to the detection of a new sequence family (L sequences). Its copy number is in the range of 10(2). The L sequences, which are about 500 bp long, are found as part of the 3' flanking regions of a clustered set of human VKI genes but they occur also separate from the genes. Models are discussed in which L sequences are viewed as molecular footprints of amplification and transposition processes of VK genes.

Amino Acid Sequence↗

The evolution of DNA sequences in Escherichia coli.

It is proposed that certain families of transposable elements originally evolved in plasmids and functioned in forming replicon fusions to aid in the horizontal transmission of non-conjugational plasmids. This hypothesis is supported by the finding that the transposable elements Tn3 and gamma delta are found almost exclusively in plasmids, and also by the distribution of the unrelated insertion sequences IS4 and IS5 among a reference collection of 67 natural isolates of Escherichia coli. Each insertion sequence was found to be present in only about one-third of the strains. Among the ten strains found to contain both insertion sequences, the number of copies of the elements was negatively correlated. With respect to IS5, approximately half of the strains containing a chromosomal copy of the insertion element also contained copies within the plasmid complement of the strain.

Base Sequence↗

Adaptive evolution of transcription factor binding sites.

BACKGROUND: The regulation of a gene depends on the binding of transcription factors to specific sites located in the regulatory region of the gene. The generation of these binding sites and of cooperativity between them are essential building blocks in the evolution of complex regulatory networks. We study a theoretical model for the sequence evolution of binding sites by point mutations. The approach is based on biophysical models for the binding of transcription factors to DNA. Hence we derive empirically grounded fitness landscapes, which enter a population genetics model including mutations, genetic drift, and selection. RESULTS: We show that the selection for factor binding generically leads to specific correlations between nucleotide frequencies at different positions of a binding site. We demonstrate the possibility of rapid adaptive evolution generating a new binding site for a given transcription factor by point mutations. The evolutionary time required is estimated in terms of the neutral (background) mutation rate, the selection coefficient, and the effective population size. CONCLUSIONS: The efficiency of binding site formation is seen to depend on two joint conditions: the binding site motif must be short enough and the promoter region must be long enough. These constraints on promoter architecture are indeed seen in eukaryotic systems. Furthermore, we analyse the adaptive evolution of genetic switches and of signal integration through binding cooperativity between different sites. Experimental tests of this picture involving the statistics of polymorphisms and phylogenies of sites are discussed.

Animals↗

Sequence variation and evolution of nuclear DNA in man and the primates.

Recent advances in nucleic acid technology have facilitated the detection and detailed structural analysis of a wide variety of genes in higher organisms, including those in man. This in turn has opened the way to an examination of the evolution of structural genes and their surrounding and intervening sequences. In a study of the evolution of haemoglobin genes and neighbouring sequences in man and the primates, we have investigated gene arrangement and DNA sequence divergence both within and between species ranging from Old World monkeys to man. This analysis is beginning to reveal the evolutionary constraints that have acted on this region of the genome during primate evolution. Furthermore, DNA sequence variation, both within and between species, provides, in principle, a novel and powerful method for determining interspecific phylogenetic distances and also for analysing the structure of present-day human populations. Application of this new branch of molecular biology to other areas of the human genome should prove important in unravelling the history of genetic changes that have occurred during the evolution of man.

Animals↗

Detection of weakly conserved ancestral mammalian regulatory sequences by primate comparisons.

BACKGROUND: Genomic comparisons between human and distant, non-primate mammals are commonly used to identify cis-regulatory elements based on constrained sequence evolution. However, these methods fail to detect functional elements that are too weakly conserved among mammals to distinguish them from non-functional DNA. RESULTS: To evaluate a strategy for large scale genome annotation that is complementary to the commonly used distal species comparisons, we explored the potential of deep intra-primate sequence comparisons. We sequenced the orthologs of 558 kb of human genomic sequence, covering multiple loci involved in cholesterol homeostasis, in 6 non-human primates. Our analysis identified six non-coding DNA elements displaying significant conservation among primates but undetectable in more distant comparisons. In vitro and in vivo tests revealed that at least three of these six elements have regulatory function. Notably, the mouse orthologs of these three functional human sequences had regulatory activity despite their lack of significant sequence conservation, indicating that they are ancestral mammalian cis-regulatory elements. These regulatory elements could be detected even in a smaller set of three primate species including human, rhesus and marmoset. CONCLUSION: We have demonstrated that intra-primate sequence comparisons can be used to identify functional modules in large genomic regions, including cis-regulatory elements that are not detectable through comparison with non-mammalian genomes. With the available human and rhesus genomes and that of marmoset, which is being actively sequenced, this strategy can be extended to the whole genome in the near future.

Animals↗

Substitution bias, rapid saturation, and the use of mtDNA for nematode systematics.

Only relatively recently have researchers turned to molecular methods for nematode phylogeny reconstruction. Thus, we lack the extensive literature on evolutionary patterns and phylogenetic usefulness of different DNA regions for nematodes that exists for other taxa. Here, we examine the usefulness of mtDNA for nematode phylogeny reconstruction and provide data that can be used for a priori character weighting or for parameter specification in models of sequence evolution. We estimated the substitution pattern for the mitochondrial ND4 gene from intraspecific comparisons in four species of parasitic nematodes from the family Trichostrongylidae (38-50 sequences per species). The resulting pattern suggests a strong mutational bias toward A and T, and a lower transition/transversion ratio than is typically observed in other taxa. We also present information on the relative rates of substitution at first, second, and third codon positions and on relative rates of saturation of different types of substitutions in comparisons ranging from intraspecific to interordinal. Silent sites saturate extremely quickly, presumably owing to the substitution bias and, perhaps, to an accelerated mutation rate. Results emphasize the importance of using only the most closely related sequences in order to infer patterns of substitution accurately for nematodes or for other taxa having strongly composition-biased DNA. ND4 also shows high amino acid polymorphism at both the intra- and interspecific levels, and in higher level comparisons, there is evidence of saturation at variable amino acid sites. In general, we recommend using mtDNA coding genes only for phylogenetics of relatively closely related nematode species and, even then, using only nonsynonymous substitutions and the more conserved mitochondrial genes (e.g., cytochrome oxidases). On the other hand, the high substitution rate in genes such as ND4 should make them excellent for population genetics studies, identifying cryptic species, and resolving relationships among closely related congeners when other markers show insufficient variation.

Amino Acid Sequence↗

Evolution of the secondary structures and compensatory mutations of the ribosomal RNAs of Drosophila melanogaster.

This paper examines the effects of DNA sequence evolution on RNA secondary structures and compensatory mutations. Models of the secondary structures of Drosophila melanogaster 18S ribosomal RNA (rRNA) and of the complex between 2S, 5.8S, and 28S rRNAs have been drawn on the basis of comparative and energetic criteria. The overall AU richness of the D. melanogaster rRNAs allows the resolution of some ambiguities in the structures of both large rRNAs. Comparison of the sequence of expansion segment V2 in D. melanogaster 18S rRNA with the same region in three other Drosophila species and the tsetse fly (Glossina morsitans morsitans) allows us to distinguish between two models for the secondary structure of this region. The secondary structures of the expansion segments of D. melanogaster 28S rRNA conform to a general pattern for all eukaryotes, despite having highly divergent sequences between D. melanogaster and vertebrates. The 70 novel compensatory mutations identified in the 28S rRNA show a strong (70%) bias toward A-U base pairs, suggesting that a process of biased mutation and/or biased fixation of A and T point mutations or AT-rich slippage-generated motifs has occurred during the evolution of D. melanogaster rDNA. This process has not occurred throughout the D. melanogaster genome. The processes by which compensatory pairs of mutations are generated and spread are discussed, and a model is suggested by which a second mutation is more likely to occur in a unit with a first mutation as such a unit begins to spread through the family and concomitantly through the population. Alternatively, mechanisms of proofreading in stem-loop structures at the DNA level, or between RNA and DNA, might be involved. The apparent tolerance of noncompensatory mutations in some stems which are otherwise strongly supported by comparative criteria within D. melanogaster 28S rRNA must be borne in mind when compensatory mutations are used as a criterion in secondary-structure modeling. Noncompensatory mutation may extend to the production of unstable structures where a stem is stabilized by RNA-protein or additional RNA-RNA interactions in the mature ribosome. Of motifs suggested to be involved in rRNA processing, one (CGAAAG) is strongly overrepresented in the 28S rRNA sequence. The data are discussed both in the context of the forces involved with the evolution of multigene families and in the context of molecular coevolution in the rDNA family in particular.

Animals↗

Characterization of the boundaries between adjacent rapidly and slowly evolving genomic regions in Drosophila.

The site of a dramatic change in the rate of DNA sequence evolution exists near the 68C glue gene clusters of several Drosophila species. We have previously determined the approximate location of this transition site by comparison of restriction maps of the regions flanking the 68C-like glue gene cluster of five members of the melanogaster species subgroup. In the present work we report the sequence of the transition region in three of these Drosophila species: D. melanogaster, D. yakuba, and D. erecta. Using a best-fit alignment of these sequences, we find that the site of transition from slowly to rapidly evolving sequences occurs abruptly within a region less than 50 nucleotides in length. Although frequency of nucleotide substitutions changes as much as 10-fold across this boundary, frequency of small insertion/deletion events stays nearly constant.

Animals↗

Phylogenetically enhanced statistical tools for RNA structure prediction.

MOTIVATION: Methods that predict the structure of molecules by looking for statistical correlation have been quite effective. Unfortunately, these methods often disregard phylogenetic information in the sequences they analyze. Here, we present a number of statistics for RNA molecular-structure prediction. Besides common pair-wise comparisons, we consider a few reasonable statistics for base-triple predictions, and present an elaborate analysis of these methods. All these statistics incorporate phylogenetic relationships of the sequences in the analysis to varying degrees, and the different nature of these tests gives a wide choice of statistical tools for RNA structure prediction. RESULTS: Starting from statistics that incorporate phylogenetic information only as independent sequence evolution models for each position of a multiple alignment, and extending this idea to a joint evolution model of two positions, we enhance the usual purely statistical methods (e.g. methods based on the Mutual Information statistic) with the use of phylogenetic information available in the sequences. In particular, we present a joint model based on the HKY evolution model, and consequently a X(2) test of independence for two positions. A significant part of this work is devoted to some mathematical analysis of these methods. We tested these statistics on regions of 16S and 23S rRNA, and tRNA.

Base Sequence↗

Predicting protein interaction sites from residue spatial sequence profile and evolution rate.

This paper proposes a novel method that can predict protein interaction sites in heterocomplexes using residue spatial sequence profile and evolution rate approaches. The former represents the information of multiple sequence alignments while the latter corresponds to a residue's evolutionary conservation score based on a phylogenetic tree. Three predictors using a support vector machines algorithm are constructed to predict whether a surface residue is a part of a protein-protein interface. The efficiency and the effectiveness of our proposed approach is verified by its better prediction performance compared with other models. The study is based on a non-redundant data set of heterodimers consisting of 69 protein chains.

Algorithms↗

Deep Sequencing Reveals Dual Evolution of SARS-CoV-2: Insights Into Defective Genomes From Wuhan-Hu-1 Variants to Omicron Subvariants.

SARS-CoV-2 has evolved from early variants dominating the first (B.1.5, B.1.1) and second (B.1.177) pandemic waves, which exhibited a higher frequency of minority mutants with deletions leading to Defective Viral Genomes (DVGs) in the spike region near the S1/S2 cleavage site than the Alpha, Beta, and Delta variants. The emergence of Omicron has significantly altered the dominant variant profile, with Omicron subvariants now representing 100% of circulating viruses. To monitor the evolution and adaptation of Omicron in the human population, a deep-sequencing study was performed in RNA samples of BA.1, BA.1.1, BA.2, BA.5, BQ.1.1, XBB.1.5 and BA.2.86 Omicron subvariants. The findings reveal two occurrences of similar evolutionary patterns within SARS-CoV-2 characterized by a shift from a significant to a very low production of DVGs. This event suggests that DVGs might play a role in the virus's spread and adaptation for persistence in infected humans.

SARS-CoV-2↗

Estimation of evolutionary distances between nucleotide sequences.

A formal mathematical analysis of the substitution process in nucleotide sequence evolution was done in terms of the Markov process. By using matrix algebra theory, the theoretical foundation of Barry and Hartigan's (Stat. Sci. 2:191-210, 1987) and Lanave et al.'s (J. Mol. Evol. 20:86-93, 1984) methods was provided. Extensive computer simulation was used to compare the accuracy and effectiveness of various methods for estimating the evolutionary distance between two nucleotide sequences. It was shown that the multiparameter methods of Lanave et al.'s (J. Mol. Evol. 20:86-93, 1984), Gojobori et al.'s (J. Mol. Evol. 18:414-422, 1982), and Barry and Hartigan's (Stat. Sci. 2:191-210, 1987) are preferable to others for the purpose of phylogenetic analysis when the sequences are long. However, when sequences are short and the evolutionary distance is large, Tajima and Nei's (Mol. Biol. Evol. 1:269-285, 1984) method is superior to others.

Base Sequence↗

Detection of convergent and parallel evolution at the amino acid sequence level.

Adaptive evolution at the molecular level can be studied by detecting convergent and parallel evolution at the amino acid sequence level. For a set of homologous protein sequences, the ancestral amino acids at all interior nodes of the phylogenetic tree of the proteins can be statistically inferred. The amino acid sites that have experienced convergent or parallel changes on independent evolutionary lineages can then be identified by comparing the amino acids at the beginning and end of each lineage. At present, the efficiency of the methods of ancestral sequence inference in identifying convergent and parallel changes is unknown. More seriously, when we identify convergent or parallel changes, it is unclear whether these changes are attributable to random chance. For these reasons, claims of convergent and parallel evolution at the amino acid sequence level have been disputed. We have conducted computer simulations to assess the efficiencies, of the parsimony and Bayesian methods of ancestral sequence inference in identifying convergent and parallel-change sites. Our results showed that the Bayesian method performs better than the parsimony method in identifying parallel changes, and both methods are inefficient in identifying convergent changes. However, the Bayesian method is recommended for estimating the number of convergent-change sites because it gives a conservative estimate. We have developed statistical tests for examining whether the observed numbers of convergent and parallel changes are due to random chance. As an example, we reanalyzed the stomach lysozyme sequences of foregut fermenters and found that parallel evolution is statistically significant, whereas convergent evolution is not well supported.

Amino Acid Sequence↗

Phylogenetic relationships of the liverworts (Hepaticae), a basal embryophyte lineage, inferred from nucleotide sequence data of the chloroplast gene rbcL.

Sequence data from the chloroplast-encoded gene rbcL were obtained for 24 liverworts, a basal group of embryophytes. Maximum likelihood and parsimony analyses of these data, along with data from other major green plant lineages, confirm hypotheses based on morphological data, such as the paraphyly of bryophytes, and the basal position of liverworts. Molecular data corroborate the deep separation between the complex thalloid and leafy/simple thalloid liverworts implied by morphological data, but the monophyly of liverworts could not be rejected. The effects of accounting for site-to-site rate heterogeneity in these data were examined using maximum likelihood methods. Comparison of trees obtained with and without rate heterogeneity showed that simply allowing for heterogeneity had a greater improvement on likelihood score than optimization of transition/transversion bias. Incorporation of site-to-site rate heterogeneity in the larger analysis, however, did not necessarily change which topology was favored. Properties of rbcL sequences from the two liverwort groups were compared. Significantly different substitution rates were found between leafy/simple thalloid and complex thalloid liverwort taxa, with rates of rbcL sequence evolution in leafy/simple thalloid taxa being higher and more indicative of those of vascular plants, and with those of complex thalloid taxa (such as Marchantia) being slower. Codon usage in rbcL in complex thalloid liverworts was biased toward NNU and NNA, compared to the leafy/simple thalloid liverworts. Although base composition and relative substitution rates differed between the two groups, no significant differences were detected within each of the two groups of liverworts. The signal present in first and second codon sites versus third codon sites was compared. While the third codon positions in rbcL across this taxon sampling are highly variable (with only 15 constant sites of 439), the trees obtained were in general agreement with trees from the entire data set and with trees obtained from independent sources of data. The presence of signal in third codon positions across greater than 400 MY of plant evolution means that definitions of saturation based on pair-wise comparisons of sequences inadequately assess phylogenetic signal.

Chloroplasts↗

Two phylogenetically highly distinct beta-tubulin genes of the basidiomycete Suillus bovinus.

Genes tubb1 and tubb2 which encode beta-tubulins 1 and 2, respectively, were characterised from the ectomycorrhizal basidiomycete Suillus bovinus. The two beta-tubulins are surprisingly divergent, with the lowest known sequence identity (60%) in any single fungal species. Comparative analysis showed that beta-tubulin 1 and the intron distribution within the tubb1 gene resemble the other beta-tubulins. beta-Tubulin 2, in contrast, is the most divergent fully described fungal beta-tubulin and the gene contains at least 21 introns, which is the largest amount known for any beta-tubulin gene. Despite this divergence, both genes are constitutively expressed in the functional compartments of the mycorrhizosphere and in pure cultures. Transcription of tubb1 is about 2.4 times higher than that of tubb2; and this difference is also seen at the translation level. Evidence suggested that phosphorylation may be the main post-translational modification of both beta-tubulins. The putative GTP-binding site residues of beta-tubulin 1 match crystallised pig beta-tubulin residues, while five of the nine differences in beta-tubulin 2 match the pig alpha-tubulin GTP-site, suggesting the presence of adaptive sequence evolution. In a Bayesian analysis, beta-tubulin 1 joins the other basidiomycete sequences, while beta-tubulin 2 loosely associates with the group of divergent ascomycete sequences without any clear relative among the known full-length fungal beta-tubulin sequences.

Amino Acid Sequence↗

Adaptation in sexuals vs. asexuals: clonal interference and the Fisher-Muller model.

Fisher and Muller's theory that recombination speeds adaptation by eliminating competition among beneficial mutations has proved a popular explanation for the advantage of sex. Recent theoretical studies have attempted to quantify the speed of adaptation under the Fisher-Muller model, partly in an attempt to understand the role of "clonal interference" in microbial experimental evolution. We reexamine adaptation in sexuals vs. asexuals, using a model of DNA sequence evolution. In this model, a modest number of sites can mutate to beneficial alleles and the fitness effects of these mutations are unequal. We study (1) transition probabilities to different beneficial mutations; (2) waiting times to the first and the last substitutions of beneficial mutations; and (3) trajectories of mean fitness through time. We find that some of these statistics are surprisingly similar between sexuals and asexuals. These results highlight the importance of the choice of substitution model in assessing the Fisher-Muller advantage of sex.

Adaptation, Biological↗