PubMed Health⌕ Search

Biomedical subjects

Dan Graur

Publications and source records attributed to Dan Graur.

At least 19 recordsLinked to original sources

A comparative analysis of numt evolution in human and chimpanzee.

Mitochondrial DNA sequences are frequently transferred into the nuclear genome, giving rise to numts (nuclear DNA sequences of mitochondrial origin). So far, the evolutionary history of numts has largely been studied by using single genomes. Here, we present the first attempt to study numt evolution in a comparative manner by using a pairwise genomic alignment. The total number of numts was estimated to be 452 in human and 469 in chimpanzee. numts that were found in both genomes at identical loci were deemed to be orthologous; 391 numts (>80%) were classified as such. The preponderance of orthologous numts is due to the very short divergence time between the 2 hominoids. The rest of numts were deemed to be nonorthologous. Nonorthologous numts were subdivided into 1) ancestral numts that have lost an ortholog in one species through deletion (12 in human and 11 in chimpanzee), 2) new numts acquired by the insertion of a mitochondrial sequence after the divergence of the 2 species (34 in human and 46 in chimpanzee), and 3) paralogous numts created by the tandem duplication of a preexisting numt (2 in human). This approach also enabled us to reconstruct the numt repertoire in the common ancestor of humans and chimpanzees (409 numts). Our comparative approach is also useful in identifying the exact boundaries of numts.

Animals↗

A computational tool for the genomic identification of regions of unusual compositional properties and its utilization in the detection of horizontally transferred sequences.

Similarity Plot (S-plot) is a Windows-based application for large-scale comparisons and 2-dimensional visualization of compositional similarities between genomic sequences. This application combines 2 approaches widely used in genomics: window analysis of statistical characteristics along genomes and dot-plot visual representation. S-plot is effective in identifying highly similar regions between genomes as well as regions with unusual compositional properties (RUCPs) within a single genome, which may be indicative of horizontal gene transfer or of locus-specific selective forces. We use S-plot to identify regions that may have originated through horizontal gene transfer through a 2-step approach, by first comparing a genomic sequence to itself and, subsequently, comparing it to the genomic sequence of a closely related taxon. Moreover, by comparing these suspect sequences to one another, we can estimate a minimum number of sources for these putative xenologous sequences. We illustrate the uses of S-plot in a comparison involving Escherichia coli K12 and E. coli O157:H7. In O157:H7, we found 145 regions that have most probably originated through horizontal gene transfer. By using S-plot to compare each of these regions with 277 completely sequenced prokaryotic genomes, 1 sequence was found to have similar compositional properties to the Yersinia pseudotuberculosis genome, indicating a transfer from a Yersinia or Yersinia relative. Based upon our analysis of RUCPs in O157:H7, we infer that there were at least 53 sources of horizontally transferred sequences.

DNA, Bacterial↗

Inferring the pattern of spontaneous mutation from the pattern of substitution in unitary pseudogenes of Mycobacterium leprae and a comparison of mutation patterns among distantly related organisms.

The pattern of spontaneous mutation can be inferred from the pattern of substitution in pseudogenes, which are known to be under very weak or no selective constraint. We modified an existing method (Gojobori T, et al., J Mol Evol 18:360, 1982) to infer the pattern of mutation in bacteria by using 569 pseudogenes from Mycobacterium leprae. In Gojobori et al.'s method, the pattern is inferred by using comparisons involving a pseudogene, a conspecific functional paralog, and an outgroup functional ortholog. Because pseudogenes in M. leprae are unitary, we replaced the missing paralogs by functional orthologs from M. tuberculosis. Functional orthologs from Streptomyces coelicolor served as outgroups. We compiled a database consisting of 69,378 inferred mutations. Transitional mutations were found to constitute more than 56% of all mutations. The transitional bias was mainly due to C-->T and G-->A, which were also the most frequent mutations on the leading strand and the only ones that were significantly more frequent than the random expectation. The least frequent mutations on the leading strand were A-->T and T-->A, each with a relative frequency of less than 3%. The mutation pattern was found to differ between the leading and the lagging strands. This asymmetry is thought to be the cause for the typical chirochoric structure of bacterial genomes. The physical distance of the pseudogene from the origin of replication (ori) was found to have almost no effect on the pattern of mutation. A surprising similarity was found between the mutation pattern in M. leprae and previously inferred patterns for such distant taxa as human and Drosophila. The mutation pattern on the leading strand of M. leprae was also found to share some common features with the pattern inferred for the heavy strand of the human mitochondrial genome. These findings indicate that taxon-specific factors may only play secondary roles in determining patterns of mutation.

Animals↗

The "domino theory" of gene death: gradual and mass gene extinction events in three lineages of obligate symbiotic bacterial pathogens.

During the adaptation of an organism to a parasitic lifestyle, various gene functions may be rendered superfluous due to the fact that the host may supply these needs. As a consequence, obligate symbiotic bacterial pathogens tend to undergo reductive genomic evolution through gene death (nonfunctionalization or pseudogenization) and deletion. Here, we examine the evolutionary sequence of gene-death events during the process of genome miniaturization in three bacterial species that have experienced extensive genome reduction: Mycobacterium leprae, Shigella flexneri, and Salmonella typhi. We infer that in all three lineages, the distribution of functional categories is similar in pseudogenes and genes but different from that of absent genes. Based on an analysis of evolutionary distances, we propose a two-step "domino effect" model for reductive genome evolution. The process starts with a gradual gene-by-gene-death sequence of events. Eventually, a crucial gene within a complex pathway or network is rendered nonfunctional triggering a "mass gene extinction" of the dependent genes. In contrast to published reports according to which genes belonging to certain functional categories are prone to nonfunctionalization more frequently and earlier than genes belonging to other functional categories, we could discern no characteristic regularity in the temporal order of function loss.

Bacteria↗

The "inverse relationship between evolutionary rate and age of mammalian genes" is an artifact of increased genetic distance with rate of evolution and time of divergence.

It has recently been claimed that older genes tend to evolve more slowly than newer ones (Alba and Castresana 2005). By simulation of genes of equal age, we show that the inverse correlation between age and rate is an artifact caused by our inability to detect homology when evolutionary distances are large. Since evolutionary distance increases with time of divergence and rate of evolution, homologs of fast-evolving genes are frequently undetected in distantly related taxa and are, hence, misclassified as "new." This misclassification causes the mean genetic distance of 'new' genes to be overestimated and the mean genetic distance of "old" genes to be underestimated.

Animals↗

In search of the vertebrate phylotypic stage: a molecular examination of the developmental hourglass model and von Baer's third law.

In 1828, Karl von Baer proposed a set of four evolutionary "laws" pertaining to embryological development. According to von Baer's third law, young embryos from different species are relatively undifferentiated and resemble one another but as development proceeds, distinguishing features of the species begin to appear and embryos of different species progressively diverge from one another. An expansion of this law, called "the hourglass model," has been proposed independently by Denis Duboule and Rudolf Raff in the 1990s. According to the hourglass model, ontogeny is characterized by a starting point at which different taxa differ markedly from one another, followed by a stage of reduced intertaxonomic variability (the phylotypic stage), and ending in a von-Baer-like progressive divergence among the taxa. A possible "translation" of the hourglass model into molecular terminology would suggest that orthologs expressed in stages described by the tapered part of the hourglass should resemble one another more than orthologs expressed in the expansive parts that precede or succeed the phylotypic stage. We tested this hypothesis using 1,585 mouse genes expressed during 26 embryonic stages, and their human orthologs. Evolutionary divergence was estimated at different embryonic stages by calculating pairwise distances between corresponding orthologous proteins from mouse and human. Two independent datasets were used. One dataset contained genes that are expressed solely in a single developmental stage; the second was made of genes expressed at different developmental stages. In the second dataset the genes were classified according to their earliest stage of expression. We fitted second order polynomials to the two datasets. The two polynomials displayed minima as expected from the hourglass model. The molecular results suggest, albeit weakly, that a phylotypic stage (or period) indeed exists. Its temporal location, sometimes between the first-somites stage and the formation of the posterior neuropore, was in approximate agreement with the morphologically defined phylotypic stage. The molecular evidence for the later parts of the hourglass model, i.e., for von Baer's third law, was stronger than that for the earlier parts.

Animals↗

GC composition of the human genome: in search of isochores.

The isochore theory, proposed nearly three decades ago, depicts the mammalian genome as a mosaic of long, fairly homogeneous genomic regions that are characterized by their guanine and cytosine (GC) content. The human genome, for instance, was claimed to consist of five distinct isochore families: L1, L2, H1, H2, and H3, with GC contents of <37%, 37%-42%, 42%-47%, 47%-52%, and >52%, respectively. In this paper, we address the question of the validity of the isochore theory through a rigorous sequence-based analysis of the human genome. Toward this end, we adopt a set of six attributes that are generally claimed to characterize isochores and statistically test their veracity against the available draft sequence of the complete human genome. By the selection criteria used in this study: distinctiveness, homogeneity, and minimal length of 300 kb, we identify 1,857 genomic segments that warrant the label "isochore." These putative isochores are nonuniformly scattered throughout the genome and cover about 41% of the human genome. We found that a four-family model of putative isochores is the most parsimonious multi-Gaussian model that can be fitted to the empirical data. These families, however, are GC poor, with mean GC contents of 35%, 38%, 41%, and 48% and do not resemble the five isochore families in the literature. Moreover, due to large overlaps among the families, it is impossible to classify genomic segments into isochore families reliably, according to compositional properties alone. These findings undermine the utility of the isochore theory and seem to indicate that the theory may have reached the limits of its usefulness as a description of genomic compositional structures.

Base Composition↗

Evolutionary conservation of bacterial operons: does transcriptional connectivity matter?

In the literature, it has been frequently suggested that the connectivity of a protein, i.e., the number of proteins with which it interacts, is inversely correlated with the rate of evolution. We attempted to extrapolate from proteins to operons by testing the hypothesis that operons with high transcriptional connectivity, i.e., operons that are controlled through interactions with many transcription factors, are evolutionarily more conserved at the structure and sequence levels than low-connectivity operons. With Escherichia coli used as reference, two structural- and two sequence-conservation measures were determined for 82 groups of homologous operons from 30 completely-sequenced bacterial genomes. In E. coli, large operons tend to be regulated by more transcription factors than either smaller operons or single genes. Large E. coli operons that are regulated by single transcription factors were found to be regulated by activators more frequently than by repressors. Levels of sequence conservation and structural conservation of operons were found to be independent of each other, i.e., structurally conserved operons may be divergent in sequence, and vice versa. Transcriptional connectivity was found to influence neither sequence nor structural conservation of operons. Although this finding seems to contradict the situation in genes, a critical review of the literature indicates that although gene connectivity is frequently touted as a factor in determining rates of evolution, only a very small fraction of the variability in degrees of evolutionary conservation is explainable by this factor.

Bacteria↗

The comparative method rules! Codon volatility cannot detect positive Darwinian selection using a single genome sequence.

All established methods for detecting positive selection at the molecular level rely on comparisons between nucleotide sequences. An exceptional method that purports to detect selection on the basis of a single genomic sequence has recently been proposed. This method uses a measure called "codon volatility," defined for each codon as the ratio between the number of nonsynonymous codons that differ from the codon under study at a single nucleotide position and the number of sense codons that differ from the codon under study at a single nucleotide position. Here, we examine various properties of codon volatility and its derivatives and use simulation of evolutionary processes to determine whether they can be used to detect selective pressures. Codons for only four amino acids (glycine, leucine, arginine, and serine) show any variation in codon volatility. Thus, codon volatility is mainly a proxy for amino acid usage, rather than for codon usage, with 65% of all synonymous changes and 27% of all nonsynonymous changes being undetectable by this measure. Genes identified by the volatility method as being subject to positive selection tend to have idiosyncratic amino acid compositions (e.g., they are glycine rich or arginine poor). An additional property of codon volatility is the near zero variance of its mean expectation, which translates into overestimated statistical significance estimates, especially in the absence of corrections for multiple comparisons. A comparison with measures of selection inferred through comparative methodology reveals no relationship between the results of the two methods. Finally, we show that codon volatility can increase in the absence of positive Darwinian selection; that is, increased codon volatility is not indicative of positive selection.

Animals↗

Comparison of site-specific rate-inference methods for protein sequences: empirical Bayesian methods are superior.

The degree to which an amino acid site is free to vary is strongly dependent on its structural and functional importance. An amino acid that plays an essential role is unlikely to change over evolutionary time. Hence, the evolutionary rate at an amino acid site is indicative of how conserved this site is and, in turn, allows evaluation of its importance in maintaining the structure/function of the protein. When using probabilistic methods for site-specific rate inference, few alternatives are possible. In this study we use simulations to compare the maximum-likelihood and Bayesian paradigms. We study the dependence of inference accuracy on such parameters as number of sequences, branch lengths, the shape of the rate distribution, and sequence length. We also study the possibility of simultaneously estimating branch lengths and site-specific rates. Our results show that a Bayesian approach is superior to maximum-likelihood under a wide range of conditions, indicating that the prior that is incorporated into the Bayesian computation significantly improves performance. We show that when branch lengths are unknown, it is better first to estimate branch lengths and then to estimate site-specific rates. This procedure was found to be superior to estimating both the branch lengths and site-specific rates simultaneously. Finally, we illustrate the difference between maximum-likelihood and Bayesian methods when analyzing site-conservation for the apoptosis regulator protein Bcl-x(L).

Animals↗

Minimal conditions for exonization of intronic sequences: 5' splice site formation in alu exons.

Alu exonization, which is an evolutionary pathway that creates primate-specific transcriptomic diversity, is a powerful tool for studying alternative-splicing regulation. Through bioinformatic analyses combined with experimental methodology, we identified the mutational changes needed to create functional 5' splice sites in Alu. We revealed a complex mechanism by which the sequence composition of the 5' splice site and its base pairing with the small nuclear RNA U1 govern alternative splicing. We show that in Alu-derived GC introns the strength of the base pairing between U1 snRNA and the 5' splice site controls the skipping/inclusion ratio of alternative splicing. Based on these findings, we identified 7810 Alus within the human genome that are prone to exonization. Mutations in these Alus may cause genetic disorders or contribute to human-specific protein diversity.

5' Untranslated Regions↗

AluGene: a database of Alu elements incorporated within protein-coding genes.

Alu elements are short interspersed elements (SINEs) approximately 300 nucleotides in length. More than 1 million Alus are found in the human genome. Despite their being genetically functionless, recent findings suggest that Alu elements may have a broad evolutionary impact by affecting gene structures, protein sequences, splicing motifs and expression patterns. Because of these effects, compiling a genomic database of Alu sequences that reside within protein-coding genes seemed a useful enterprise. Presently, such data are limited since the structural and positional information on genes and Alu sequences are scattered throughout incompatible and unconnected databases. AluGene (http://Alugene.tau.ac.il/) provides easy access to a complete Alu map of the human genome, as well as Alu-associated information. The Alu elements are annotated with respect to coding region and exon/intron location. This design facilitates queries on Alu sequences, locations, as well as motifs and compositional properties via a one-stop search page.

Alu Elements↗

Evolution of multicellularity in Metazoa: comparative analysis of the subcellular localization of proteins in Saccharomyces, Drosophila and Caenorhabditis.

A comparison of the subcellular assignments of proteins between the unicellular Saccharomyces cerevisiae and the multicellular Drosophila melanogaster and Caenorhabditis elegans was performed using a computational tool for the prediction of subcellular localization. Nine subcellular compartments were studied: (1) extracellular domain, (2) cell membrane, (3) cytoplasm, (4) endoplasmic reticulum, (5) Golgi apparatus, (6) lysosome, (7) peroxisome, (8) mitochondria, and (9) nucleus. The transition to multicellularity was found to be characterized by an increase in the total number of proteins encoded by the genome. Interestingly, this increase is distributed unevenly among the subcellular compartments. That is, a disproportionate increase in the number of proteins in the extracellular domain, the cell membrane, and the cytoplasm is observed in multicellular organisms, while no such increase is seen in other subcellular compartments. A possible explanation involves signal transduction. In terms of protein numbers, signal transduction pathways may be roughly described as a pyramid with an expansive base in the extracellular domain (the numerous extracellular signal proteins), progressively narrowing at the cell membrane and cytoplasmic levels, and ending in a narrow tip consisting of only a handful of transcription modulators in the nucleus. Our observations suggest that extracellular signaling interactions among metazoan cells account for the uneven increase in the numbers of proteins among subcellular compartments during the transition to multicellularity.

Animals↗

Reading the entrails of chickens: molecular timescales of evolution and the illusion of precision.

For almost a decade now, a team of molecular evolutionists has produced a plethora of seemingly precise molecular clock estimates for divergence events ranging from the speciation of cats and dogs to lineage separations that might have occurred approximately 4 billion years ago. Because the appearance of accuracy has an irresistible allure, non-specialists frequently treat these estimates as factual. In this article, we show that all of these divergence-time estimates were generated through improper methodology on the basis of a single calibration point that has been unjustly denuded of error. The illusion of precision was achieved mainly through the conversion of statistical estimates (which by definition possess standard errors, ranges and confidence intervals) into errorless numbers. By employing such techniques successively, the time estimates of even the most ancient divergence events were made to look deceptively precise. For example, on the basis of just 15 genes, the arthropod-nematode divergence event was 'calculated' to have occurred 1167+/-83 million years ago (i.e. within a 95% confidence interval of approximately 350 million years). Were calibration and derivation uncertainties taken into proper consideration, the 95% confidence interval would have turned out to be at least 40 times larger ( approximately 14.2 billion years).

Animals↗

Incongruent expression profiles between human and mouse orthologous genes suggest widespread neutral evolution of transcription control.

Rapid rates of evolution can signify either a lack of selective constraint and the consequent accumulation of neutral alleles, or positive Darwinian selection driving the fixation of advantageous alleles. Based on a comparison of 1,350 orthologous gene pairs from human and mouse, we show that the evolution of gene expression profiles is so rapid that it is comparable to that of paralogous gene pairs or randomly paired genes. The expression divergence in the entire set of orthologous pairs neither strongly correlates with sequence divergence, nor focuses in any particular tissue. Moreover, comparing tissue expressions across the orthologous gene pairs, we observe that any human tissue is more similar to any other human tissue examined than to its corresponding mouse tissue. Collectively, these results indicate that, while some differences in expression profiles may be due to adaptive evolution, the levels of divergence are mostly compatible with a neutral mode of evolution, in which a mutation for ectopic expression may rise to fixation by random drift without significantly affecting the fitness. A disturbing corollary of these findings is that knowledge of where the gene is expressed may not carry information about its function.

Animals↗

Detecting excess radical replacements in phylogenetic trees.

There are a few instances in which positive Darwinian selection has been convincingly demonstrated at the molecular level. In this study, we present a novel test for detecting excess of radical amino-acid replacements. Such excess is usually indicative of positive Darwinian selection, but may also be due to relaxed functional constraints or model misspecification. In our test, each amino-acid replacement is characterized in terms of a physicochemical distance, i.e., the degree of dissimilarity between the exchanged amino-acid residues. By using phylogenetic trees based on protein sequences, our test identifies statistically significant deviations of the mean physicochemical distance from the random expectation, either along a taxonomic lineage or across a subtree. The mean inferred distance is calculated as the average physicochemical distance over all possible ancestral sequence reconstructions weighted by their likelihood. Our method substantially improves over previous approaches by taking into account the stochastic process, tree phylogeny, among-site rate variation, and alternative ancestral reconstructions. We provide a fast linear time algorithm for applying this test to all branches and all subtrees of a given phylogenetic tree. We validate this approach by applying it to two well-studied datasets: the MHC class I glycoproteins serving as a positive control, and the house-keeping gene carbonic anhydrase I serving as a negative control.

Algorithms↗

pANT: a method for the pairwise assessment of nonfunctionalization times of processed pseudogenes.

We present a method for pairwise Assessment of Nonfunctionalization Times (pANT) in processed pseudogenes. Contrary to existing methods for estimating nonfunctionalization times, pANT utilizes previously calculated probabilities of nucleotide substitution as explicit rate measurements, rather than assume that the substitution rates are the same for all nucleotides. Thus, the method allows a more accurate computation of the time that has elapsed since the nonfunctionalization of a pseudogene. Whereas existing methods require the sequence of an orthologous functional gene, which is not always at hand, pANT only uses the pairwise alignment of the gene/pseudogene pair, thus expanding the range of problems that can be tackled. To estimate evolutionary times in nonfunctional sequences, pANT measures the differences in the pairwise alignment of a gene and its paralogous processed pseudogene, using only the first and second codon positions. It assumes that, because of functional constraints, these positions in the sequence of the functional homolog have not changed since the time of nonfunctionalization of the pseudogene. Hence, the sequence of the gene may be used as the ancestor of the pseudogene. We show that the method's reliance on a detailed substitution matrix, which is derived separately for each species, makes it more accurate than existing methods. We applied pANT to the case of the unitary alpha-1,3-galactosyltransferase human pseudogene and found that our estimate of the nonfunctionalization time was in agreement with that obtained by taxonomic and paleontological considerations pertaining to the divergence between platyrrhines (New World monkeys) and cattarhines (Old World monkeys).

Animals↗