PubMed Health⌕ Search

Biomedical subjects

David Sankoff

Publications and source records attributed to David Sankoff.

15 recordsLinked to original sources

The distribution of genomic distance between random genomes.

We study the probability distribution of genomic distance d under the hypothesis of random gene order. We translate the random order assumption into a stochastic method for constructing the alternating color cycles in the decomposition of the bicolored breakpoint graph. For two random genomes of length n, we show that the expectation of n - d is O((1/2) log n).

Computational Biology↗

Stability of rearrangement measures in the comparison of genome sequences.

We present data-analytic and statistical tools for studying rates of rearrangement of whole genomes and to assess the stability of these methods with changes in the level of resolution of the genomic data. We construct datasets on the numbers of conserved syntenies and conserved segments shared by pairs of animal genomes at different levels of resolution. We fit these data to an evolutionary tree and find the rates of rearrangement on various evolutionary lineages. We document the lack of clocklike behavior of rearrangement processes, the independence of translocation and inversion rates, and the level of resolution beyond which translocations rates are lost in noise due to other processes.

Animals↗

The statistical analysis of spatially clustered genes under the maximum gap criterion.

Statistical validation of gene clusters is imperative for many important applications in comparative genomics which depend on the identification of genomic regions that are historically and/or functionally related. We develop the first rigorous statistical treatment of max-gap clusters, a cluster definition frequently used in empirical studies. We present exact expressions for the probability of observing an individual cluster of a set of marked genes in one genome, as well as upper and lower bounds on the probability of observing a cluster of h homologs in a pairwise whole-genome comparison. We demonstrate the utility of our approach by applying it to a whole-genome comparison of E. coli and B. subtilis. Code for statistical tests is available at.

Bacillus subtilis↗

Reversal distance for partially ordered genomes.

MOTIVATION: The total order of the genes or markers on a chromosome inherent in its representation as a signed per-mutation must often be weakened to a partial order in the case of real data. This is due to lack of resolution (where several genes are mapped to the same chromosomal position) to missing data from some of the datasets used to compile a gene order, and to conflicts between these datasets. The available genome rearrangement algorithms, however, require total orders as input. A more general approach is needed to handle rearrangements of gene partial orders. RESULTS: We formalize the uncertainty in gene order data by representing a chromosome from each genome as a partial order, summarized by a directed acyclic graph (DAG). The rearrangement problem is then to infer a minimal sequence of reversals for transforming any topological sort of one DAG to any one of the other DAG. Each topological sort represents a possible linearization compatible with all the datasets on the chromosome. The set of all possible topological sorts is embedded in each DAG by appropriately augmenting the edge set, so that it becomes a general directed graph (DG). The DGs representing chromosomes of two genomes are combined to produce a bicoloured graph from which we extract a maximal decomposition into alternating coloured cycles, and from which, in turn, an optimal sequence of reversals can usually be identified. We test this approach on simulated incomplete comparative maps and on cereal chromosomal maps drawn from the Gramene browser.

Algorithms↗

Genomic features in the breakpoint regions between syntenic blocks.

MOTIVATION: We study the largely unaligned regions between the syntenic blocks conserved in humans and mice, based on data extracted from the UCSC genome browser. These regions contain evolutionary breakpoints caused by inversion, translocation and other processes. RESULTS: We suggest explanations for the limited amount of genomic alignment in the neighbourhoods of breakpoints. We discount inferences of extensive breakpoint reuse as artefacts introduced during the reconstruction of syntenic blocks. We find that the number, size and distribution of small aligned fragments in the breakpoint regions depend on the origin of the neighbouring blocks and the other blocks on the same chromosome. We account for this and for the generalized loss of alignment in the regions partially by artefacts due to alignment protocols and partially by mutational processes operative only after the rearrangement event. These results are consistent with breakpoints occurring randomly over virtually the entire genome.

Algorithms↗

Improving gene network inference by comparing expression time-series across species, developmental stages or tissues.

We present a method for gene network inference and revision based on time-series data. Gene networks are modeled using linear differential equations and a generalized stepwise multiple linear regression procedure is used to recover the interaction coefficients. Our system is designed for the recovery of gene interactions concurrently in many gene regulatory networks related by a tree or a more general graph. We show how this comparative framework can facilitate the recovery of the networks and improve the quality of the solutions inferred.

Algorithms↗

Structural dynamics of eukaryotic chromosome evolution.

Large-scale genome sequencing is providing a comprehensive view of the complex evolutionary forces that have shaped the structure of eukaryotic chromosomes. Comparative sequence analyses reveal patterns of apparently random rearrangement interspersed with regions of extraordinarily rapid, localized genome evolution. Numerous subtle rearrangements near centromeres, telomeres, duplications, and interspersed repeats suggest hotspots for eukaryotic chromosome evolution. This localized chromosomal instability may play a role in rapidly evolving lineage-specific gene families and in fostering large-scale changes in gene order. Computational algorithms that take into account these dynamic forces along with traditional models of chromosomal rearrangement show promise for reconstructing the natural history of eukaryotic chromosomes.

Animals↗

Rearrangements and chromosomal evolution.

Comparisons of the genome sequences of related species suggests varying patterns of chromosomal rearrangements in different evolutionary lineages. In this review, I focus on the quantitative characterization of rearrangement processes and discuss specific inventories that have been compiled to date. Of particular interest are the statistical distribution of the lengths of inverted or locally transposed chromosome fragments (notably very short ones), inhomogeneities in susceptibility to evolutionary breakpoints in chromosomal regions, the relative importance of genome doubling in the history of multicellular eukaryotes, and of lateral transfer versus gene gain and loss in prokaryotes. These developments provide challenges to computational biologists to refine, revise and scale up mathematical models and algorithms for analyzing genome rearrangements.

Animals↗

Tests for gene clustering.

Comparing chromosomal gene order in two or more related species is an important approach to studying the forces that guide genome organization and evolution. Linked clusters of similar genes found in related genomes are often used to support arguments of evolutionary relatedness or functional selection. However, as the gene order and the gene complement of sister genomes diverge progressively due to large scale rearrangements, horizontal gene transfer, gene duplication and gene loss, it becomes increasingly difficult to determine whether observed similarities in local genomic structure are indeed remnants of common ancestral gene order, or are merely coincidences. A rigorous comparative genomics requires principled methods for distinguishing chance commonalities, within or between genomes, from genuine historical or functional relationships. In this paper, we construct tests for significant groupings against null hypotheses of random gene order, taking incomplete clusters, multiple genomes, and gene families into account. We consider both the significance of individual clusters of prespecified genes and the overall degree of clustering in whole genomes.

Algorithms↗

Chromosomal distributions of breakpoints in cancer, infertility, and evolution.

We extract 11 genome-wide sets of breakpoint positions from databases on reciprocal translocations, inversions and deletions in neoplasms, reciprocal translocations and inversions in families carrying rearrangements and the human-mouse comparative map, and for each set of positions construct breakpoint distributions for the 44 autosomal arms. We identify and interpret four main types of distribution: (i) a uniform distribution associated both with families carrying translocations or inversions, and with the comparative map, (ii) telomerically skewed distributions of translocations or inversions detected consequent to births with malformations, (iii) medially clustered distributions of translocation and deletion breakpoints in tumor karyotypes, and (iv) bimodal translocation breakpoint distributions for chromosome arms containing telomeric proto-oncogenes.

Animals↗

Genetic and molecular control of folate-homocysteine metabolism in mutant mice.

Hyperhomocysteinemia adversely affects fundamental aspects of fetal development, adulthood, and aging, but the role of elevated homocysteine levels in these birth defects and adult diseases remains unclear. Mouse models are valuable for investigating the causes and consequences of hyperhomocysteinemia. We used a phenotype-based approach to identify mouse mutants for studying the relation between single gene mutations, homocysteine levels as a measure of the status of homocysteine metabolism, and gene expression profiles as a way to assess the impact of protein deficiency in mutant mice on steady-state transcription levels of genes in the folate-homocysteine pathways. These mutants were selected based on their propensity to produce phenotypes that are reminiscent of those associated with anomalies in folate-homocysteine metabolism in humans. We report identification of new, single-gene mouse models of homocysteinemia and characterization of their molecular and physiological impact on folate-homocysteine metabolism. Mutations in several genes involved in the hedgehog and WNT signal transduction pathways, as well as a gene involved in lipid metabolism, resulted in elevated homocysteine levels and altered expression profiles of folate-homocysteine metabolism genes. These results begin to unravel the complex relations between elevation of a single amino acid in the blood and the diverse birth defects and adult diseases associated with hyperhomocysteinemia.

Animals↗

Short inversions and conserved gene cluster.

MOTIVATION: Two independent sets of recent observations on newly sequenced microbial genomes pertain to the prevalence of short inversion as a gene order rearrangement process and to the lack of conservation of gene order within conserved gene clusters. We propose a model of inversion where the key parameter is the length of the inverted fragment. RESULTS: We show that there is a qualitative difference in the pattern of evolution when the inversion length is small with respect to the cluster size and when it is large. This suggests an explanation of the lack of parallel gene order in conserved clusters and raises questions about the statistical validity of putative functionally selected gene clusters if these have only been tested against inappropriate null hypotheses.

Base Sequence↗

Chromosomal breakpoint reuse in genome sequence rearrangement.

In order to apply gene-order rearrangement algorithms to the comparison of genome sequences, Pevzner and Tesler bypass gene finding and ortholog identification and use the order of homologous blocks of unannotated sequence as input. The method excludes blocks shorter than a threshold length. Here we investigate possible biases introduced by eliminating short blocks, focusing on the notion of breakpoint reuse introduced by these authors. Analytic and simulation methods show that reuse is very sensitive to the proportion of blocks excluded. As is pertinent to the comparison of mammalian genomes, this exclusion risks randomizing the comparison partially or entirely.

Algorithms↗