PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

A chromosomal rearrangement hotspot can be identified from population genetic variation and is coincident with a hotspot for allelic recombination.

Insights into the origins of structural variation and the mutational mechanisms underlying genomic disorders would be greatly improved by a genomewide map of hotspots of nonallelic homologous recombination (NAHR). Moreover, our understanding of sequence variation within the duplicated sequences that are substrates for NAHR lags far behind that of sequence variation within the single-copy portion of the genome. Perhaps the best-characterized NAHR hotspot lies within the 24-kb-long Charcot-Marie-Tooth disease type 1A (CMT1A)-repeats (REPs) that sponsor deletions and duplications that cause peripheral neuropathies. We investigated structural and sequence diversity within the CMT1A-REPs, both within and between species. We discovered a high frequency of retroelement insertions, accelerated sequence evolution after duplication, extensive paralogous gene conversion, and a greater than twofold enrichment of SNPs in humans relative to the genome average. We identified an allelic recombination hotspot underlying the known NAHR hotspot, which suggests that the two processes are intimately related. Finally, we used our data to develop a novel method for inferring the location of an NAHR hotspot from sequence variation within segmental duplications and applied it to identify a putative NAHR hotspot within the LCR22 repeats that sponsor velocardiofacial syndrome deletions. We propose that a large-scale project to map sequence variation within segmental duplications would reveal a wealth of novel chromosomal-rearrangement hotspots.

Alleles↗

Primary structures of dehydrogenases. Evolutionary characteristics related to functional aspects; models for isozyme developments and ancestral connections.

This chapter describes known characteristics of evolutionary changes in individual dehydrogenases, as well as possible relationships among this group of enzymes. Data from primary structures are correlated with those from other observations. Variations in the amino acid sequences demonstrate functional properties, and can be interpreted in relation to conformational aspects, subunit arrangements and enzyme stabilities. Different types of isozyme developments have occurred and show functional fixations at various levels. They define isozyme patterns of general significance in protein evolution. Sequence similarities may be found between different segments. They are analyzed in relation to known conformations, subunit sizes, species divergence and genetic mechanisms. A wide-ranging evolutionary model is discussed relating dehydrogenases and some other oligomeric enzymes to a distant, frequently remodelled ancestral building unit of repetitive occurrence.

Alcohol Oxidoreductases↗

Reduced adaptation of a non-recombining neo-Y chromosome.

Sex chromosomes are generally believed to have descended from a pair of homologous autosomes. Suppression of recombination between the ancestral sex chromosomes led to the genetic degeneration of the Y chromosome. In response, the X chromosome may become dosage-compensated. Most proposed mechanisms for the degeneration of Y chromosomes involve the rapid fixation of deleterious mutations on the Y. Alternatively, Y-chromosome degeneration might be a response to a slower rate of adaptive evolution, caused by its lack of recombination. Here we report patterns of DNA polymorphism and divergence at four genes located on the neo-sex chromosomes of Drosophila miranda. We show that a higher rate of protein sequence evolution of the neo-X-linked copy of Cyclin B relative to the neo-Y copy is driven by positive selection, which is consistent with the adaptive hypothesis for the evolution of the Y chromosome. In contrast, the neo-Y-linked copies of even-skipped and roundabout show an elevated rate of protein evolution relative to their neo-X homologues, probably reflecting the reduced effectiveness of selection against deleterious mutations in a non-recombining genome. Our results provide evidence for the importance of sexual recombination for increasing and maintaining the level of adaptation of a population.

Adaptation, Biological↗

Comparative Genomics of Sex-Determination-Related Genes Reveals Shared Evolutionary Patterns Between Bivalves and Mammals, but Not Fruit Flies.

The molecular basis of sex determination (SD), while being extensively studied in model organisms, remains poorly understood in many animal groups. Bivalves, a diverse class of molluscs with a variety of reproductive modes, represent an ideal yet challenging clade for investigating SD and the evolution of sexual systems. However, the absence of a comprehensive framework has limited progress in this field, particularly regarding the study of sex-determination-related genes (SRGs). In this study, we performed a genome-wide sequence evolutionary analysis of the Dmrt, Sox and Fox gene families in more than 40 bivalve species. For the first time, we provide an extensive and phylogenetically aware dataset of these SRGs, and we find support for the hypothesis that Dmrt-1L and Sox-H may act as primary sex-determining genes by showing their high levels of sequence diversity within the bivalve genomic context. To validate our findings, we studied the same gene families in two well-characterised systems, mammals and fruit flies (genus Drosophila). In the former, we found that the male sex-determining gene Sry exhibits a pattern of amino acid sequence diversity similar to that of Dmrt-1L and Sox-H in bivalves, consistent with its role as master SD regulator. In contrast, no such pattern was observed among genes of the fruit fly SD cascade, which is controlled by a chromosomic mechanism. Overall, our findings highlight similarities in the sequence evolution of some mammal and bivalve SRGs, possibly driven by a comparable architecture of SD cascades. This work underscores once again the importance of employing a comparative approach when investigating understudied and non-model systems.

Animals↗

Alpha1 and alpha2 domains of Aotus MHC class I and Catarrhini MHC class Ia share similar characteristics.

Functional and structural analyses of major histocompatibility complex (MHC) class I molecules of the Aotus genus are necessary to validate it as a solid animal model for biomedical research. We thus isolated, cloned and sequenced exons 2 and 3 from three Aotus species (A. nancymaae, A. nigriceps and A. vociferans). We found 24 sequences, which divided into two different groups (Ao-g1 and Ao-g2). A further sequence was identified as a processed pseudogene (Aona-PS2). Both sequence evolution and variability analyses showed that Ao-g1 and Ao-g2 display similar characteristics to Catarrhini's classical loci, such as positive selection pressure at the peptide binding region (PBR) high variability and a trans-specific evolution pattern.

Amino Acid Sequence↗

Genomic heterogeneity of background substitutional patterns in Drosophila melanogaster.

Mutation is the underlying force that provides the variation upon which evolutionary forces can act. It is important to understand how mutation rates vary within genomes and how the probabilities of fixation of new mutations vary as well. If substitutional processes across the genome are heterogeneous, then examining patterns of coding sequence evolution without taking these underlying variations into account may be misleading. Here we present the first rigorous test of substitution rate heterogeneity in the Drosophila melanogaster genome using almost 1500 nonfunctional fragments of the transposable element DNAREP1_DM. Not only do our analyses suggest that substitutional patterns in heterochromatic and euchromatic sequences are different, but also they provide support in favor of a recombination-associated substitutional bias toward G and C in this species. The magnitude of this bias is entirely sufficient to explain recombination-associated patterns of codon usage on the autosomes of the D. melanogaster genome. We also document a bias toward lower GC content in the pattern of small insertions and deletions (indels). In addition, the GC content of noncoding DNA in Drosophila is higher than would be predicted on the basis of the pattern of nucleotide substitutions and small indels. However, we argue that the fast turnover of noncoding sequences in Drosophila makes it difficult to assess the importance of the GC biases in nucleotide substitutions and small indels in shaping the base composition of noncoding sequences.

Animals↗

Compensatory relationship between splice sites and exonic splicing signals depending on the length of vertebrate introns.

BACKGROUND: The signals that determine the specificity and efficiency of splicing are multiple and complex, and are not fully understood. Among other factors, the relative contributions of different mechanisms appear to depend on intron size inasmuch as long introns might hinder the activity of the spliceosome through interference with the proper positioning of the intron-exon junctions. Indeed, it has been shown that the information content of splice sites positively correlates with intron length in the nematode, Drosophila, and fungi. We explored the connections between the length of vertebrate introns, the strength of splice sites, exonic splicing signals, and evolution of flanking exons. RESULTS: A compensatory relationship is shown to exist between different types of signals, namely, the splice sites and the exonic splicing enhancers (ESEs). In the range of relatively short introns (approximately, < 1.5 kilobases in length), the enhancement of the splicing signals for longer introns was manifest in the increased concentration of ESEs. In contrast, for longer introns, this effect was not detectable, and instead, an increase in the strength of the donor and acceptor splice sites was observed. Conceivably, accumulation of A-rich ESE motifs beyond a certain limit is incompatible with functional constraints operating at the level of protein sequence evolution, which leads to compensation in the form of evolution of the splice sites themselves toward greater strength. In addition, however, a correlation between sequence conservation in the exon ends and intron length, particularly, in synonymous positions, was observed throughout the entire length range of introns. Thus, splicing signals other than the currently defined ESEs, i.e., potential new classes of ESEs, might exist in exon sequences, particularly, those that flank long introns. CONCLUSION: Several weak but statistically significant correlations were observed between vertebrate intron length, splice site strength, and potential exonic splicing signals. Taken together, these findings attest to a compensatory relationship between splice sites and exonic splicing signals, depending on intron length.

Animals↗

Evolution of the human ASPM gene, a major determinant of brain size.

The size of human brain tripled over a period of approximately 2 million years (MY) that ended 0.2-0.4 MY ago. This evolutionary expansion is believed to be important to the emergence of human language and other high-order cognitive functions, yet its genetic basis remains unknown. An evolutionary analysis of genes controlling brain development may shed light on it. ASPM (abnormal spindle-like microcephaly associated) is one of such genes, as nonsense mutations lead to primary microcephaly, a human disease characterized by a 70% reduction in brain size. Here I provide evidence suggesting that human ASPM went through an episode of accelerated sequence evolution by positive Darwinian selection after the split of humans and chimpanzees but before the separation of modern non-Africans from Africans. Because positive selection acts on a gene only when the gene function is altered and the organismal fitness is increased, my results suggest that adaptive functional modifications occurred in human ASPM and that it may be a major genetic component underlying the evolution of the human brain.

Anthropometry↗

Molecular evolution of the hepatitis B virus genome.

The hepatitis B virus (HBV) has a circular DNA genome of about 3,200 base pairs. Economical use of the genome with overlapping reading frames may have led to severe constraints on nucleotide substitutions along the genome and to highly variable rates of substitution among nucleotide sites. Nucleotide sequences from 13 complete HBV genomes were compared to examine such variability of substitution rates among sites and to examine the phylogenetic relationships among the HBV variants. The maximum likelihood method was employed to fit models of DNA sequence evolution that can account for the complexity of the pattern of nucleotide substitution. Comparison of the models suggests that the rates of substitution are different in different genes and codon positions; for example, the third codon position changes at a rate over ten times higher than the second position. Furthermore, substantial variation of substitution rates was detected even after the effects of genes and codon positions were corrected; that is, rates are different at different sites of the same gene or at the same codon position. Such rates after the correction were also found to be positively correlated at adjacent sites, which indicated the existence of conserved and variable domains in the proteins encoded by the viral genome. A multiparameter model validates the earlier finding that the variation in nucleotide conservation is not random around the HBV genome. The test for the existence of a molecular clock suggests that substitution rates are more or less constant among lineages. The phylogenetic relationships among the viral variants were examined. Although the data do not seem to contain sufficient information to resolve the details of the phylogeny, it appears quite certain that the serotypes of the viral variants do not reflect their genetic relatedness.

Codon↗

Comparative structural modeling and inference of conserved protein classes in Drosophila seminal fluid.

The constituents of seminal fluid are a complex mixture of proteins and other molecules, most of whose functions have yet to be determined and many of which are rapidly evolving. As a step in elucidating the roles of these proteins and exposing potential functional similarities hidden by their rapid evolution, we performed comparative structural modeling on 28 of 52 predicted seminal proteins produced in the Drosophila melanogaster male accessory gland. Each model was characterized by defining residues likely to be important for structure and function. Comparisons of known protein structures with predicted accessory gland proteins (Acps) revealed similarities undetectable by primary sequence alignments. The structures predict that Acps fall into several categories: regulators of proteolysis, lipid modifiers, immunity/protection, sperm-binding proteins, and peptide hormones. The comparative structural modeling approach indicates that major functional classes of mammalian and Drosophila seminal fluid proteins are conserved, despite differences in reproductive strategies. This is particularly striking in the face of the rapid protein sequence evolution that characterizes many reproductive proteins, including Drosophila and mammalian seminal proteins.

Amino Acid Sequence↗

Variance to mean ratio, R(t), for poisson processes on phylogenetic trees.

The ratio of expected variance to mean, R(t), of numbers of DNA base substitutions for contemporary sequences related by a "star" phylogeny is widely seen as a measure of the adherence of the sequences' evolution to a Poisson process with a molecular clock, as predicted by the "neutral theory" of molecular evolution under certain conditions. A number of estimators of R(t) have been proposed, all predicted to have mean 1 and distributions based on the chi 2. Various genes have previously been analyzed and found to have values of R(t) far in excess of 1, calling into question important aspects of the neutral theory. In this paper, I use Monte Carlo simulation to show that the previously suggested means and distributions of estimators of R(t) are highly inaccurate. The analysis is applied to star phylogenies and to general phylogenetic trees, and well-known gene sequences are reanalyzed. For star phylogenies the results show that Kimura's estimators ("The Neutral Theory of Molecular Evolution," Cambridge Univ. Press, Cambridge, 1983) are unsatisfactory for statistical testing of R(t), but confirm the accuracy of Bulmer's correction factor (Genetics 123: 615-619, 1989). For all three nonstar phylogenies studied, attained values of all three estimators of R(t), although larger than 1, are within their true confidence limits under simple Poisson process models. This shows that lineage effects can be responsible for high estimates of R(t), restoring some limited confidence in the molecular clock and showing that the distinction between lineage and molecular clock effects is vital.(ABSTRACT TRUNCATED AT 250 WORDS)

Analysis of Variance↗

Patterns of molecular evolution among paralogous floral homeotic genes.

The plant MADS-box regulatory gene family includes several loci that control different aspects of inflorescence and floral development. Orthologs to the Arabidopsis thaliana MADS-box floral meristem genes APETALA1 and CAULIFLOWER and the floral organ identity genes APETALA3 and PISTILLATA were isolated from the congeneric species Arabidopsis lyrata. Analysis of these loci between these two Arabidopsis species, as well as three other more distantly related taxa, reveal contrasting dynamics of molecular evolution between these paralogous floral regulatory genes. Among the four loci, the CAL locus evolves at a significantly faster rate, which may be associated with the evolution of genetic redundancy between CAL and AP1. Moreover, there are significant differences in the distribution of replacement and synonymous substitutions between the functional gene domains of different floral homeotic loci. These results indicate that divergence in developmental function among paralogous members of regulatory gene families is accompanied by changes in rate and pattern of sequence evolution among loci.

Arabidopsis↗

Dynamically heterogenous partitions and phylogenetic inference: an evaluation of analytical strategies with cytochrome b and ND6 gene sequences in cranes.

ki ctes over whether molecular sequence data should be partitioned for phylogenetic analysis often confound two types of heterogeneity among partitions. We distinguish historical heterogeneity (i.e., different partitions have different evolutionary relationships) from dynamic heterogeneity (i.e., different partitions show different patterns of sequence evolution) and explore the impact of the latter on phylogenetic accuracy and precision with a two-gene, mitochondrial data set for cranes. The well-established phylogeny of cranes allows us to contrast tree-based estimates of relevant parameter values with estimates based on pairwise comparisons and to ascertain the effects of incorporating different amounts of process information into phylogenetic estimates. We show that codon positions in the cytochrome b and NADH dehydrogenase subunit 6 genes are dynamically heterogenous under both Poisson and invariable-sites + gamma-rates versions of the F84 model and that heterogeneity includes variation in base composition and transition bias as well as substitution rate. Estimates of transition-bias and relative-rate parameters from pairwise sequence comparisons were comparable to those obtained as tree-based maximum likelihood estimates. Neither rate-category nor mixed-model partitioning strategies resulted in a loss of phylogenetic precision relative to unpartitioned analyses. We suggest that weighted-average distances provide a computationally feasible alternative to direct maximum likelihood estimates of phylogeny for mixed-model analyses of large, dynamically heterogenous data sets.

Algorithms↗

The influence of neighboring-nucleotide composition on single nucleotide polymorphisms (SNPs) in the mouse genome and its comparison with human SNPs.

We analyzed the neighboring-nucleotide composition of 433,192 biallelic substitutions, representing the largest public collection of SNPs across the mouse genome. Large neighboring-nucleotide biases relative to the genome- or chromosome-specific average were observed at the immediate adjacent sites and small biases extended farther from the substitution site. For all substitutions, the biases for A, C, G, and T were 0.21, 2.63, 0.71, and -3.55%, respectively, on the immediate adjacent 5' site and -3.67, 0.75, 2.69, and 0.23%, respectively, on the immediate adjacent 3' side. Further examination of the six categories of substitution revealed that the neighboring-nucleotide patterns for transitions were strongly influenced by the hypermutability of dinucleotide CpG and the neighboring effects on transversions were complex. Probability of a transversion increased with increasing A + T content of the two immediate adjacent sites, which was similarly observed in the human and Arabidopsis genomes. Overall, the bias patterns for the neighboring nucleotides in the mouse and human genomes were essentially the same; however, the extent of the biases was notably less in mice. Our results provide the first comprehensive view of the neighboring-nucleotide effects in the mouse genome and are important for understanding the mutational mechanisms and sequence evolution in the mammalian genomes.

Animals↗

Accelerated evolution and Muller's rachet in endosymbiotic bacteria.

Many bacteria live only within animal cells and infect hosts through cytoplasmic inheritance. These endosymbiotic lineages show distinctive population structure, with small population size and effectively no recombination. As a result, endosymbionts are expected to accumulate mildly deleterious mutations. If these constitute a substantial proportion of new mutations, endosymbionts will show (i) faster sequence evolution and (ii) a possible shift in base composition reflecting mutational bias. Analyses of 16S rDNA of five independently derived endosymbiont clades show, in every case, faster evolution in endosymbionts than in free-living relatives. For aphid endosymbionts (genus Buchnera), coding genes exhibit accelerated evolution and unusually low ratios of synonymous to nonsynonymous substitutions compared to ratios for the same genes for enterics. This concentration of the rate increase in nonsynonymous substitutions is expected under the hypothesis of increased fixation of deleterious mutations. Polypeptides for all Buchnera genes analyzed have accumulated amino acids with codon families rich in A+T, supporting the hypothesis that substitutions are deleterious in terms of polypeptide function. These observations are best explained as the result of Muller's ratchet within small asexual populations, combined with mutational bias. In light of this explanation, two observations reported earlier for Buchnera, the apparent loss of a repair gene and the overproduction of a chaperonin, may reflect compensatory evolution. An alternative hypothesis, involving selection on genomic base composition, is contradicted by the observation that the speedup is concentrated at nonsynonymous sites.

Animals↗

Long-term excretion of vaccine-derived poliovirus by a healthy child.

A child was found to be excreting type 1 vaccine-derived poliovirus (VDPV) with a 1.1% sequence drift from Sabin type 1 vaccine strain in the VP1 coding region 6 months after he was immunized with oral live polio vaccine. Seventeen type 1 poliovirus isolates were recovered from stools taken from this child during the following 4 months. Contrary to expectation, the child was not deficient in humoral immunity and showed high levels of serum neutralization against poliovirus. Selected virus isolates were characterized in terms of their antigenic properties, virulence in transgenic mice, sensitivity for growth at high temperatures, and differences in nucleotide sequence from the Sabin type 1 strain. The VDPV isolates showed mutations at key nucleotide positions that correlated with the observed reversion to biological properties typical of wild polioviruses. A number of capsid mutations mapped at known antigenic sites leading to changes in the viral antigenic structure. Estimates of sequence evolution based on the accumulation of nucleotide changes in the VP1 coding region detected a "defective" molecular clock running at an apparent faster speed of 2.05% nucleotide changes per year versus 1% shown in previous studies. Remarkably, when compared to several type 1 VDPV strains of different origins, isolates from this child showed a much higher proportion of nonsynonymous versus synonymous nucleotide changes in the capsid coding region. This anomaly could explain the high VP1 sequence drift found and the ability of these virus strains to replicate in the gut for a longer period than expected.

Animals↗

ChloroplastDB: the Chloroplast Genome Database.

The Chloroplast Genome Database (ChloroplastDB) is an interactive, web-based database for fully sequenced plastid genomes, containing genomic, protein, DNA and RNA sequences, gene locations, RNA-editing sites, putative protein families and alignments (http://chloroplast.cbio.psu.edu/). With recent technical advances, the rate of generating new organelle genomes has increased dramatically. However, the established ontology for chloroplast genes and gene features has not been uniformly applied to all chloroplast genomes available in the sequence databases. For example, annotations for some published genome sequences have not evolved with gene naming conventions. ChloroplastDB provides unified annotations, gene name search, BLAST and download functions for chloroplast encoded genes and genomic sequences. A user can retrieve all orthologous sequences with one search regardless of gene names in GenBank. This feature alone greatly facilitates comparative research on sequence evolution including changes in gene content, codon usage, gene structure and post-transcriptional modifications such as RNA editing. Orthologous protein sets are classified by TribeMCL and each set is assigned a standard gene name. Over the next few years, as the number of sequenced chloroplast genomes increases rapidly, the tools available in ChloroplastDB will allow researchers to easily identify and compile target data for comparative analysis of chloroplast genes and genomes.

Chloroplasts↗

Geographical variation in selection, from phenotypes to molecules.

Molecular technologies now allow researchers to isolate quantitative trait loci (QTLs) and measure patterns of gene sequence variation within chromosomal regions containing important polymorphisms. I develop a simulation model to investigate gene sequence evolution within genomic regions that harbor QTLs. The QTLs influence a trait experiencing geographical variation in selection, which is common in nature and produces obvious differentiation at the phenotypic level. Counter to expectations, the simulations suggest that selection can substantially affect quantitative genetic variation without altering the amount and pattern of molecular variation at sites closely linked to the QTLs. Even with large samples of gene sequences, the likelihood of rejecting neutrality is often low. The exception is situations where strong selection is combined with low migration among demes, conditions that may be common in many plant species. The results have implications for gene sequence surveys and, perhaps more generally, for interpreting the apparently weak connection between levels of molecular and quantitative trait variation within species.

Achillea↗