PubMed Health⌕ Search

Biomedical subjects

Simon Tavaré

Publications and source records attributed to Simon Tavaré.

At least 19 recordsLinked to original sources

Genome-wide mapping of ORC and Mcm2p binding sites on tiling arrays and identification of essential ARS consensus sequences in S. cerevisiae.

BACKGROUND: Eukaryotic replication origins exhibit different initiation efficiencies and activation times within S-phase. Although local chromatin structure and function influences origin activity, the exact mechanisms remain poorly understood. A key to understanding the exact features of chromatin that impinge on replication origin function is to define the precise locations of the DNA sequences that control origin function. In S. cerevisiae, Autonomously Replicating Sequences (ARSs) contain a consensus sequence (ACS) that binds the Origin Recognition Complex (ORC) and is essential for origin function. However, an ACS is not sufficient for origin function and the majority of ACS matches do not function as ORC binding sites, complicating the specific identification of these sites. RESULTS: To identify essential origin sequences genome-wide, we utilized a tiled oligonucleotide array (NimbleGen) to map the ORC and Mcm2p binding sites at high resolution. These binding sites define a set of potential Autonomously Replicating Sequences (ARSs), which we term nimARSs. The nimARS set comprises 529 ORC and/or Mcm2p binding sites, which includes 95% of known ARSs, and experimental verification demonstrates that 94% are functional. The resolution of the analysis facilitated identification of potential ACSs (nimACSs) within 370 nimARSs. Cross-validation shows that the nimACS predictions include 58% of known ACSs, and experimental verification indicates that 82% are essential for ARS activity. CONCLUSION: These findings provide the most comprehensive, accurate, and detailed mapping of ORC binding sites to date, adding to the emerging picture of the chromatin organization of the budding yeast genome.

Algorithms↗

MMASS: an optimized array-based method for assessing CpG island methylation.

We describe an optimized microarray method for identifying genome-wide CpG island methylation called microarray-based methylation assessment of single samples (MMASS) which directly compares methylated to unmethylated sequences within a single sample. To improve previous methods we used bioinformatic analysis to predict an optimized combination of methylation-sensitive enzymes that had the highest utility for CpG-island probes and different methods to produce unmethylated representations of test DNA for more sensitive detection of differential methylation by hybridization. Subtraction or methylation-dependent digestion with McrBC was used with optimized (MMASS-v2) or previously described (MMASS-v1, MMASS-sub) methylation-sensitive enzyme combinations and compared with a published McrBC method. Comparison was performed using DNA from the cell line HCT116. We show that the distribution of methylation microarray data is inherently skewed and requires exogenous spiked controls for normalization and that analysis of digestion of methylated and unmethylated control sequences together with linear fit models of replicate data showed superior statistical power for the MMASS-v2 method. Comparison with previous methylation data for HCT116 and validation of CpG islands from PXMP4, SFRP2, DCC, RARB and TSEN2 confirmed the accuracy of MMASS-v2 results. The MMASS-v2 method offers improved sensitivity and statistical power for high-throughput microarray identification of differential methylation.

Cell Line, Tumor↗

Non-linear analysis of GeneChip arrays.

The application of microarray hybridization theory to Affymetrix GeneChip data has been a recent focus for data analysts. It has been shown that the hyperbolic Langmuir isotherm captures the shape of the signal response to concentration of Affymetrix GeneChips. We demonstrate that existing linear fit methods for extracting gene expression measures are not well adapted for the effect of saturation resulting from surface adsorption processes. In contrast to the most popular methods, we fit background and concentration parameters within a single global fitting routine instead of estimating the background before obtaining gene expression measures. We describe a non-linear multi-chip model of the perfect match signal that effectively allows for the separation of specific and non-specific components of the microarray signal and avoids saturation bias in the high-intensity range. Multimodel inference, incorporated within the fitting routine, allows a quantitative selection of the model that best describes the observed data. The performance of this method is evaluated on publicly available datasets, and comparisons to popular algorithms are presented.

Algorithms↗

Inferring population parameters from single-feature polymorphism data.

This article is concerned with a statistical modeling procedure to call single-feature polymorphisms from microarray experiments. We use this new type of polymorphism data to estimate the mutation and recombination parameters in a population. The mutation parameter can be estimated via the number of single-feature polymorphisms called in the sample. For the recombination parameter, a two-feature sampling distribution is derived in a way analogous to that for the two-locus sampling distribution with SNP data. The approximate-likelihood approach using the two-feature sampling distribution is examined and found to work well. A coalescent simulation study is used to investigate the accuracy and robustness of our method. Our approach allows the utilization of single-feature polymorphism data for inference in population genetics.

Arabidopsis↗

A unique recent origin of the allotetraploid species Arabidopsis suecica: Evidence from nuclear DNA markers.

A coalescent-based method was used to investigate the origins of the allotetraploid Arabidopsis suecica, using 52 nuclear microsatellite loci typed in eight individuals of A. suecica and 14 individuals of its maternal parent Arabidopsis thaliana, and four short fragments of genomic DNA sequenced in a sample of four individuals of A. suecica and in both its parental species A. thaliana and Arabidopsis arenosa. All loci were variable in A. thaliana but only 24 of the 52 microsatellite loci and none of the four sequence fragments were variable in A. suecica. We explore a number of possible evolutionary scenarios for A. suecica and conclude that it is likely that A. suecica has a recent, unique origin between 12,000 and 300,000 years ago. The time estimates depend strongly on what is assumed about population growth and rates of mutation. When combined with what is known about the history of glaciations, our results suggest that A. suecica originated south of its present distribution in Sweden and Finland and then migrated north, perhaps in the wake of the retreating ice.

Algorithms↗

Counting divisions in a human somatic cell tree: how, what and why?

The billions of cells within an individual can be organized by genealogy into a single somatic cell tree that starts from the zygote and ends with present day cells. In theory, this tree can be reconstructed from replication errors that surreptitiously record divisions and ancestry. Such a molecular clock approach is currently impractical because somatic mutations are rare, but more feasible measurements are possible by substituting instead the 5' to 3' order of epigenetic modifications such as CpG methylation. Epigenetic somatic errors are readily detected as age-related changes in methylation, which suggests certain adult stem cells divide frequently and "compete" for survival within niches. Potentially the genealogy of any human cell may be reconstructed without prior experimental manipulation by merely reading histories recorded in their genomes.

Biological Clocks↗

Human hair genealogies and stem cell latency.

BACKGROUND: Stem cells divide to reproduce themselves and produce differentiated progeny. A fundamental problem in human biology has been the inability to measure how often stem cells divide. Although it is impossible to observe every division directly, one method for counting divisions is to count replication errors; the greater the number of divisions, the greater the numbers of errors. Stem cells with more divisions should produce progeny with more replication errors. METHODS: To test this approach, epigenetic errors (methylation) in CpG-rich molecular clocks were measured from human hairs. Hairs exhibit growth and replacement cycles and "new" hairs physically reappear even on "old" heads. Errors may accumulate in long-lived stem cells, or in their differentiated progeny that are eventually shed. RESULTS: Average hair errors increased until two years of age, and then were constant despite decades of replacement, consistent with new hairs arising from infrequently dividing bulge stem cells. Errors were significantly more frequent in longer hairs, consistent with long-lived but eventually shed mitotic follicle cells. CONCLUSION: Constant average hair methylation regardless of age contrasts with the age-related methylation observed in human intestine, suggesting that error accumulation and therefore stem cell latency differs among tissues. Epigenetic molecular clocks imply similar mitotic ages for hairs on young and old human heads, consistent with a restart with each new hair, and with genealogies surreptitiously written within somatic cell genomes.

Biological Clocks↗

Population-based genetic epidemiologic analysis of Chlamydia trachomatis serotypes and lack of association between ompA polymorphisms and clinical phenotypes.

Chlamydia trachomatis is the leading cause of bacterial sexually transmitted diseases worldwide. Urogenital strains are classified into serotypes and genotypes based on the major outer membrane protein and its gene, ompA, respectively. Studies of the association of serotypes with clinical signs and symptoms have produced conflicting results while no studies have evaluated associations with ompA polymorphisms. We designed a population-based cross-sectional study of 344 men and women with urogenital chlamydial infections (excluding co-pathogen infections) presenting to clinics serving five U.S. cities from 1995 to 1997. Signs, symptoms and sequelae of chlamydial infection (mucopurulent cervicitis, vaginal or urethral discharge; dysuria; lower abdominal pain; abnormal vaginal bleeding; and pelvic inflammatory disease) were analyzed for associations with serotype and ompA polymorphisms. One hundred and fifty-three (44.5%) of 344 patients had symptoms consistent with urogenital chlamydial infection. Gender, reason for visit and city were significant independent predictors of symptom status. Men were 2.2 times more likely than women to report any symptoms (P=0.03) and 2.8 times more likely to report a urethral discharge than women were to report a vaginal discharge in adjusted analyses (P=0.007). Differences in serotype or ompA were not predictive except for an association between serotype F and pelvic inflammatory disease (P=0.046); however, the number of these cases was small. While there was no clinically prognostic value associated with serotype or ompA polymorphism for urogenital chlamydial infections except for serotype F, future studies might utilize multilocus genomic typing to identify chlamydial strains associated with clinical phenotypes.

Adolescent↗

Modern computational approaches for analysing molecular genetic variation data.

An explosive growth is occurring in the quantity, quality and complexity of molecular variation data that are being collected. Historically, such data have been analysed by using model-based methods. Models are useful for sharpening intuition, for explanation and for prediction: they add to our understanding of how the data were formed, and they can provide quantitative answers to questions of interest. We outline some of these model-based approaches, including the coalescent, and discuss the applicability of the computational methods that are necessary given the highly complex nature of current and future data sets.

Computational Biology↗

Genome-wide associations of gene expression variation in humans.

The exploration of quantitative variation in human populations has become one of the major priorities for medical genetics. The successful identification of variants that contribute to complex traits is highly dependent on reliable assays and genetic maps. We have performed a genome-wide quantitative trait analysis of 630 genes in 60 unrelated Utah residents with ancestry from Northern and Western Europe using the publicly available phase I data of the International HapMap project. The genes are located in regions of the human genome with elevated functional annotation and disease interest including the ENCODE regions spanning 1% of the genome, Chromosome 21 and Chromosome 20q12-13.2. We apply three different methods of multiple test correction, including Bonferroni, false discovery rate, and permutations. For the 374 expressed genes, we find many regions with statistically significant association of single nucleotide polymorphisms (SNPs) with expression variation in lymphoblastoid cell lines after correcting for multiple tests. Based on our analyses, the signal proximal (cis-) to the genes of interest is more abundant and more stable than distal and trans across statistical methodologies. Our results suggest that regulatory polymorphism is widespread in the human genome and show that the 5-kb (phase I) HapMap has sufficient density to enable linkage disequilibrium mapping in humans. Such studies will significantly enhance our ability to annotate the non-coding part of the genome and interpret functional variation. In addition, we demonstrate that the HapMap cell lines themselves may serve as a useful resource for quantitative measurements at the cellular level.

Chromosome Mapping↗

Counting human somatic cell replications: methylation mirrors endometrial stem cell divisions.

Cell proliferation may be altered in many diseases, but it is uncertain exactly how to measure total numbers of divisions. Although it is impossible to count every division directly, potentially total numbers of stem cell divisions since birth may be inferred from numbers of somatic errors. The idea is that divisions are surreptitiously recorded by random errors that occur during replication. To test this "molecular clock" hypothesis, epigenetic errors encoded in certain methylation patterns were counted in glands from 30 uteri. Endometrial divisions can differ among women because of differences in estrogen exposures or numbers of menstrual cycles. Consistent with an association between mitotic age and methylation, there was an age-related increase in methylation with stable levels after menopause, and significantly less methylation was observed in lean or older multiparous women. Methylation patterns were diverse and more consistent with niche rather than immortal stem cell lineages. There was no evidence for decreased stem cell survival with aging. An ability to count lifetime numbers of stem cell divisions covertly recorded by random replication errors provides new opportunities to link cell proliferation with aging and cancer.

Adolescent↗

Numbers of mutations to different types of colorectal cancer.

BACKGROUND: The numbers of oncogenic mutations required for transformation are uncertain but may be inferred from how cancer frequencies increase with aging. Cancers requiring more mutations will tend to appear later in life. This type of approach may be confounded by biologic heterogeneity because different cancer subtypes may require different numbers of mutations. For example, a sporadic cancer should require at least one more somatic mutation relative to its hereditary counterpart. METHODS: To better estimate numbers of mutations before transformation, 1,022 colorectal cancers were classified with respect to microsatellite instability (MSI) and germline DNA mismatch repair mutations characteristic of hereditary nonpolyposis colorectal cancer (HNPCC). MSI- cancers were also classified with respect to clinical stage. Ages at cancer and a Bayesian algorithm were used to estimate the numbers of oncogenic mutations required for transformation for each cancer subtype. RESULTS: Ages at MSI+ cancers were consistent with five or six oncogenic mutations for hereditary (HNPCC) cancers, and seven or eight mutations for its sporadic counterpart. Ages at cancer were consistent with seven mutations for sporadic MSI- cancers, and were similar (six to eight mutations) regardless of clinical cancer stage. CONCLUSION: Different biologic subtypes of colorectal cancer appear to require different numbers of oncogenic mutations before transformation. Sporadic MSI+ cancers may require more than a single additional somatic alteration compared to hereditary MSI+ cancers because the epigenetic inactivation of MLH1 commonly observed in sporadic MSI+ cancers may be a multistep process. Interestingly, estimated numbers of MSI- cancer mutations were similar (six to eight mutations) regardless of clinical cancer stage, suggesting a propensity to spread or metastasize does not require additional mutations after transformation. Estimates of oncogenic mutation numbers may help explain some of the biology underlying different cancer subtypes.

Adult↗

Estimating a nucleotide substitution rate for maize from polymorphism at a major domestication locus.

To estimate a rate for single nucleotide substitutions for maize (Zea mays ssp. mays), we have taken advantage of data from genetic and archaeological studies of the domestication of maize from its wild ancestor, teosinte (Z. mays ssp. parviglumis). Genetic studies have shown that the teosinte branched1 (tb1) gene was a major target of human selection during maize domestication, and sequence diversity in the intergenic region 5' to the tb1-coding sequence is extraordinarily low. We show that polymorphism in this region is consistent with new mutation following fixation for a small number of tb1 haplotypes during domestication. Archeological studies suggest that maize was domesticated approximately 6,250-10,000 years ago and subsequently the size of the maize population is thought to have expanded rapidly. Using the observed number of mutations within the region of selection at tb1, the approximate age of maize domestication, and approximations for the maize genealogy, we have derived estimates for the nucleotide substitution rate for the tb1 intergenic region. Using two approaches, one of which is a coalescent approach, we obtain rate estimates of approximately 2.9 x 10(-8) and 3.3 x 10(-8) substitutions per site per year. We also show that the pattern of polymorphism in the tb1 intergenic region appears to have been strongly affected by the mutagenic effect of DNA methylation. Excluding target sites of symmetric DNA methylation (CG and CNG sites) from analysis, the mutation rate estimates are reduced by approximately 50%-60%, while the rates for CG and CNG sites are nearly an order of magnitude higher. We use rate estimates from the tb1 region to estimate the timing of expansion of transposable elements in the maize genome and suggest that this expansion occurred primarily within the last million years.

Base Sequence↗

Age-related human small intestine methylation: evidence for stem cell niches.

BACKGROUND: The small intestine is constructed of many crypts and villi, and mouse studies suggest that each crypt contains multiple stem cells. Very little is known about human small intestines because mouse fate mapping strategies are impractical in humans. However, it is theoretically possible that stem cell histories are inherently written within their genomes. Genomes appear to record histories (as exemplified by use of molecular clocks), and therefore it may be possible to reconstruct somatic cell dynamics from somatic cell errors. Recent human colon studies suggest that random somatic epigenetic errors record stem cell histories (ancestry and total numbers of divisions). Potentially age-related methylation also occurs in human small intestines, which would allow characterization of their stem cells and comparisons with the colon. METHODS: Methylation patterns in individual crypts from 13 small intestines (17 to 78 years old) were measured by bisulfite sequencing. The methylation patterns were analyzed by a quantitative model to distinguish between immortal or niche stem cell lineages. RESULTS: Age-related methylation was observed in the human small intestines. Crypt methylation patterns were more consistent with stem cell niches than immortal stem cell lineages. Human large and small intestine crypt niches appeared to have similar stem cell dynamics, but relatively less methylation accumulated with age in the small intestines. There were no apparent stem cell differences between the duodenum and ileum, and stem cell survival did not appear to decline with aging. CONCLUSION: Crypt niches containing multiple stem cells appear to maintain human small intestines. Crypt niches appear similar in the colon and small intestine, and the small intestinal stem cell mitotic rate is the same as or perhaps slower than that of the colon. Although further studies are needed, age-related methylation appears to record somatic cell histories, and a somatic epigenetic molecular clock strategy may potentially be applied to other human tissues to reconstruct otherwise occult stem cell histories.

Adolescent↗

Statistical tests of the coalescent model based on the haplotype frequency distribution and the number of segregating sites.

Several tests of neutral evolution employ the observed number of segregating sites and properties of the haplotype frequency distribution as summary statistics and use simulations to obtain rejection probabilities. Here we develop a "haplotype configuration test" of neutrality (HCT) based on the full haplotype frequency distribution. To enable exact computation of rejection probabilities for small samples, we derive a recursion under the standard coalescent model for the joint distribution of the haplotype frequencies and the number of segregating sites. For larger samples, we consider simulation-based approaches. The utility of the HCT is demonstrated in simulations of alternative models and in application to data from Drosophila melanogaster.

Chromosome Segregation↗

Similar gene expression patterns characterize aging and oxidative stress in Drosophila melanogaster.

Affymetrix GeneChips were used to measure RNA abundance for approximately 13,500 Drosophila genes in young, old, and 100% oxygen-stressed flies. Data were analyzed by using a recently developed background correction algorithm and a robust multichip model-based statistical analysis that dramatically increased the ability to identify changes in gene expression. Aging and oxidative stress responses shared the up-regulation of purine biosynthesis, heat shock protein, antioxidant, and innate immune response genes. Results were confirmed by using Northerns and transgenic reporters. Immune response gene promoters linked to GFP allowed longitudinal assay of gene expression during aging in individual flies. Immune reporter expression in young flies was partially predictive of remaining life span, suggesting their potential as biomonitors of aging.

Aging↗

Pretumor progression: clonal evolution of human stem cell populations.

Multistep carcinogenesis through sequential cycles of mutation and clonal succession is usually described as tumor progression, or the clonal evolution of tumor cell populations. However, many mutations found in cancers are also compatible with normal appearing phenotypes and therefore genetic progression may precede tumor progression. To better characterize such pretumor progression (mutations in the absence of visible phenotypic changes), a quantitative model was developed that postulates most oncogenic cancer mutations first accumulate in normal appearing colon crypt niche stem cells. Each crypt contains multiple stem cells, and random niche stem cell loss with replacement eventually leads to the loss of all stem cell lineages except one. This niche succession or crypt clonal evolution is similar to the clonal succession of tumor progression except it does not require selection or change visible phenotype. Mutations may sequentially accumulate during stem cell clonal evolution either through drift (passenger mutations) or selection. To determine the feasibility of pretumor progression, mutation rates sufficient to recreate the epidemiology of colorectal cancer were estimated. Pretumor progression may completely substitute for visible tumor progression because it is theoretically possible for all cancer mutations to first accumulate in normal appearing colon with normal replication fidelity. Elevated mutation rates or tumorigenesis may be unnecessary for early progression.

Adult↗

Enhanced stem cell survival in familial adenomatous polyposis.

Individuals with heterozygous germline adenomatous polyposis coli (APC) mutations or familial adenomatous polyposis (FAP) are born with normal appearing colons but later develop hundreds to thousands of polyps. Tumor progression apparently starts after somatic loss of the normal APC allele, but germline APC mutations may potentially alter niche stem cell survival through dominant-negative interactions or haploinsufficiency. Although morphologically occult, altered stem cell turnover or clonal evolution rates may be detected by measuring the diversity of crypt sequences, with greater diversity expected with longer lived stem cell lineages. Methylation pattern diversity (numbers of unique patterns per crypt) was higher in normal appearing crypts from four of five FAP colons compared to six non-FAP colons and one attenuated FAP colon. Simulations indicate higher FAP crypt diversity is consistent with slower clonal evolution from enhanced stem cell survival, either through increased stem cell numbers or decreased stem cell lineage extinction, which is predicted to increase progression rates to cancer. Enhanced stem cell survival was associated with APC mutations that remove some but not all catenin-binding repeats. Therefore, some APC mutations may be common in colorectal cancers because they confer occult pretumor "caretaker" and "gatekeeper" defects. FAP crypts accumulate more alterations from slower stem cell clonal evolution rather than increased error rates. In non-FAP crypts, enhanced stem cell survival conferred by somatic heterozygous APC mutations would favor fixation through occult clonal niche expansions. Heterozygous APC mutations may change stem cell survival during colorectal pretumor progression.

Adenomatous Polyposis Coli↗