PubMed Health⌕ Search

Biomedical subjects

Carsten Wiuf

Publications and source records attributed to Carsten Wiuf.

At least 19 recordsLinked to original sources

Evaluating Neanderthal genetics and phylogeny.

The retrieval of Neanderthal (Homo neanderthalsensis) mitochondrial DNA is thought to be among the most significant ancient DNA contributions to date, allowing conflicting hypotheses on modern human (Homo sapiens) evolution to be tested directly. Recently, however, both the authenticity of the Neanderthal sequences and their phylogenetic position outside contemporary human diversity have been questioned. Using Bayesian inference and the largest dataset to date, we find strong support for a monophyletic Neanderthal clade outside the diversity of contemporary humans, in agreement with the expectations of the Out-of-Africa replacement model of modern human origin. From average pairwise sequence differences, we obtain support for claims that the first published Neanderthal sequence may include errors due to postmortem damage in the template molecules for PCR. In contrast, we find that recent results implying that the Neanderthal sequences are products of PCR artifacts are not well supported, suffering from inadequate experimental design and a presumably high percentage (>68%) of chimeric sequences due to "jumping PCR" events.

Animals↗

The effects of incomplete protein interaction data on structural and evolutionary inferences.

BACKGROUND: Present protein interaction network data sets include only interactions among subsets of the proteins in an organism. Previously this has been ignored, but in principle any global network analysis that only looks at partial data may be biased. Here we demonstrate the need to consider network sampling properties explicitly and from the outset in any analysis. RESULTS: Here we study how properties of the yeast protein interaction network are affected by random and non-random sampling schemes using a range of different network statistics. Effects are shown to be independent of the inherent noise in protein interaction data. The effects of the incomplete nature of network data become very noticeable, especially for so-called network motifs. We also consider the effect of incomplete network data on functional and evolutionary inferences. CONCLUSION: Crucially, when only small, partial network data sets are considered, bias is virtually inevitable. Given the scope of effects considered here, previous analyses may have to be carefully reassessed: ignoring the fact that present network data are incomplete will severely affect our ability to understand biological systems.

Evolution, Molecular↗

Consistency of estimators of population scaled parameters using composite likelihood.

Composite likelihood methods have become very popular for the analysis of large-scale genomic data sets because of the computational intractability of the basic coalescent process and its generalizations: It is virtually impossible to calculate the likelihood of an observed data set spanning a large chromosomal region without using approximate or heuristic methods. Composite likelihood methods are approximate methods and, in the present article, assume the likelihood is written as a product of likelihoods, one for each of a number of smaller regions that together make up the whole region from which data is collected. A very general framework for neutral coalescent models is presented and discussed. The framework comprises many of the most popular coalescent models that are currently used for analysis of genetic data. Assume data is collected from a series of consecutive regions of equal size. Then it is shown that the observed data forms a stationary, ergodic process. General conditions are given under which the maximum composite estimator of the parameters describing the model (e.g. mutation rates, demographic parameters and the recombination rate) is a consistent estimator as the number of regions tends to infinity.

Base Sequence↗

Recharacterization of ancient DNA miscoding lesions: insights in the era of sequencing-by-synthesis.

Although ancient DNA (aDNA) miscoding lesions have been studied since the earliest days of the field, their nature remains a source of debate. A variety of conflicting hypotheses exist about which miscoding lesions constitute true aDNA damage as opposed to PCR polymerase amplification error. Furthermore, considerable disagreement and speculation exists on which specific damage events underlie observed miscoding lesions. The root of the problem is that it has previously been difficult to assemble sufficient data to test the hypotheses, and near-impossible to accurately determine the specific strand of origin of observed damage events. With the advent of emulsion-based clonal amplification (emPCR) and the sequencing-by-synthesis technology this has changed. In this paper we demonstrate how data produced on the Roche GS20 genome sequencer can determine miscoding lesion strands of origin, and subsequently be interpreted to enable characterization of the aDNA damage behind the observed phenotypes. Through comparative analyses on 390,965 bp of modern chloroplast and 131,474 bp of ancient woolly mammoth GS20 sequence data we conclusively demonstrate that in this sample at least, a permafrost preserved specimen, Type 2 (cytosine-->thymine/guanine-->adenine) miscoding lesions represent the overwhelming majority of damage-derived miscoding lesions. Additionally, we show that an as yet unidentified guanine-->adenine analogue modification, not the conventionally argued cytosine-->uracil deamination, underpins a significant proportion of Type 2 damage. How widespread these implications are for aDNA will become apparent as future studies analyse data recovered from a wider range of substrates.

Animals↗

Genotyping and annotation of Affymetrix SNP arrays.

In this paper we develop a new method for genotyping Affymetrix single nucleotide polymorphism (SNP) array. The method is based on (i) using multiple arrays at the same time to determine the genotypes and (ii) a model that relates intensities of individual SNPs to each other. The latter point allows us to annotate SNPs that have poor performance, either because of poor experimental conditions or because for one of the alleles the probes do not behave in a dose-response manner. Generally, our method agrees well with a method developed by Affymetrix. When both methods make a call they agree in 99.25% (using standard settings) of the cases, using a sample of 113 Affymetrix 10k SNP arrays. In the majority of cases where the two methods disagree, our method makes a genotype call, whereas the method by Affymetrix makes a no call, i.e. the genotype of the SNP is not determined. By visualization it is indicated that our method is likely to be correct in majority of these cases. In addition, we demonstrate that our method produces more SNPs that are in concordance with Hardy-Weinberg equilibrium than the method by Affymetrix. Finally, we have validated our method on HapMap data and shown that the performance of our method is comparable to other methods.

Algorithms↗

Frequent occurrence of uniparental disomy in colorectal cancer.

We used SNP arrays to identify and characterize genomic alterations associated with colorectal cancer (CRC). Laser microdissected cancer cells from 15 adenocarinomas were investigated by Affymetrix Mapping 10K SNP arrays. Analysis of the data extracted from the SNP arrays revealed multiple regions with copy number alterations and loss of heterozygosity (LOH). Novel LOH areas were identified at chromosomes 13, 14 and 15. Combined analysis of the LOH and copy number data revealed genomic structures that could not have been identified analyzing either data type alone. Half of the identified LOH regions showed no evidence of a reduced copy number, indicating the presence of uniparental structures. The distribution of these structures was non-random, primarily involving 8q, 13q and 20q. This finding was supported by analysis of an independent set of array-based transcriptional profiles, consisting of 17 normal mucosa and 66 adenocarcinoma samples. The transcriptional analysis revealed an unchanged expression level in areas with intact copy number, including regions with uniparental disomy, and a reduced expression level in the LOH regions representing factual losses (including 5q, 8p and 17p). The analysis also showed that genes in regions with increased copy number (including 7p and 20q) were predominantly upregulated. Further analyses of the SNP data revealed a subset of the identified alterations to be specifically associated with TP53 inactivation (including 8q gain and 17p loss) and lymph node metastasis status (gain of 7q and 13q). Another subset of the identified alterations was shown to represent intratumor heterogeneity. In conclusion, we demonstrate that uniparental disomy is frequent in CRC, and identify genomic alterations associated with TP53 inactivation and lymph node status.

Adenocarcinoma↗

Co-clustering and visualization of gene expression data and gene ontology terms for Saccharomyces cerevisiae using self-organizing maps.

We propose a novel co-clustering algorithm that is based on self-organizing maps (SOMs). The method is applied to group yeast (Saccharomyces cerevisiae) genes according to both expression profiles and Gene Ontology (GO) annotations. The combination of multiple databases is supposed to provide a better biological definition and separation of gene clusters. We compare different levels of genome-wide co-clustering by weighting the involved sources of information differently. Clustering quality is determined by both general and SOM-specific validation measures. Co-clustering relies on a sufficient correlation between the different datasets. We investigate in various experiments how much GO information is contained in the applied gene expression dataset and vice versa. The second major contribution is a visualization technique that applies the cluster structure of SOMs for a better biological interpretation of gene (expression) clusterings. Our GO term maps reveal functional neighborhoods between clusters forming biologically meaningful functional SOM regions. To cope with the high variety and specificity of GO terms, gene and cluster annotations are mapped to a reduced vocabulary of more general GO terms. In particular, this advances the ability of SOMs to act as gene function predictors.

Artificial Intelligence↗

A likelihood approach to analysis of network data.

Biological, sociological, and technological network data are often analyzed by using simple summary statistics, such as the observed degree distribution, and nonparametric bootstrap procedures to provide an adequate null distribution for testing hypotheses about the network. In this article we present a full-likelihood approach that allows us to estimate parameters for general models of network growth that can be expressed in terms of recursion relations. To handle larger networks we have developed an importance sampling scheme that allows us to approximate the likelihood and draw inference about the network and how it has been generated, estimate the parameters in the model, and perform parametric bootstrap analysis of network data. We illustrate the power of this approach by estimating growth parameters for the Caenorhabditis elegans protein interaction network.

Animals↗

Crosslinks rather than strand breaks determine access to ancient DNA sequences from frozen sediments.

Diagenesis was studied in DNA obtained from Siberian permafrost (permanently frozen soil) ranging from 10,000 to 400,000 years in age. Despite optimal preservation conditions, we found the sedimentary DNA to be severely modified by interstrand crosslinks; single- and double-stranded breaks; and freely exposed sugar, phosphate, and hydroxyl groups. Intriguingly, interstrand crosslinks were found to accumulate approximately 100 times faster than single-stranded breaks, suggesting that crosslinking rather than depurination is the primary limiting factor for ancient DNA amplification under frozen conditions. The results question the reliability of the commonly used models relying on depurination kinetics for predicting the long-term survival of DNA under permafrost conditions and suggest that new strategies for repair of ancient DNA must be considered if the yield of amplifiable DNA from permafrost sediments is to be significantly increased. Using the obtained rate constant for interstrand crosslinks the maximal survival time of amplifiable 120-bp fragments of bacterial 16S ribosomal DNA was estimated to be approximately 400,000 years. Additionally, a clear relationship was found between DNA damage and sample age, contradicting previously raised concerns about the possible leaching of free DNA molecules between permafrost layers.

Cross-Linking Reagents↗

SOX4 expression in bladder carcinoma: clinical aspects and in vitro functional characterization.

The human transcription factor SOX4 was 5-fold up-regulated in bladder tumors compared with normal tissue based on whole-genome expression profiling of 166 clinical bladder tumor samples and 27 normal urothelium samples. Using a SOX4-specific antibody, we found that the cancer cells expressed the SOX4 protein and, thus, did an evaluation of SOX4 protein expression in 2,360 bladder tumors using a tissue microarray with clinical annotation. We found a correlation (P < 0.05) between strong SOX4 expression and increased patient survival. When overexpressed in the bladder cell line HU609, SOX4 strongly impaired cell viability and promoted apoptosis. To characterize downstream target genes and SOX4-induced pathways, we used a time-course global expression study of the overexpressed SOX4. Analysis of the microarray data showed 130 novel SOX4-related genes, some involved in signal transduction (MAP2K5), angiogenesis (NRP2), and cell cycle arrest (PIK3R3) and others with unknown functions (CGI-62). Among the genes regulated by SOX4, 25 contained at least one SOX4-binding motif in the promoter sequence, suggesting a direct binding of SOX4. The gene set identified in vitro was analyzed in the clinical bladder material and a small subset of the genes showed a high correlation to SOX4 expression. The present data suggest a role of SOX4 in the bladder cancer disease.

Apoptosis↗

Convergence properties of the degree distribution of some growing network models.

In this article we study a class of randomly grown graphs that includes some preferential attachment and uniform attachment models, as well as some evolving graph models that have been discussed previously in the literature. The degree distribution is assumed to form a Markov chain; this gives a particularly simple form for a stochastic recursion of the degree distribution. We show that for this class of models the empirical degree distribution tends almost surely and in norm to the expected degree distribution as the size of the graph grows to infinity and we provide a simple asymptotic expression for the expected degree distribution. Convergence of the empirical degree distribution has consequences for statistical analysis of network data in that it allows the full data to be summarized by the degree distribution of the nodes without losing the ability to obtain consistent estimates of parameters describing the network.

Markov Chains↗

Assessing the fidelity of ancient DNA sequences amplified from nuclear genes.

To date, the field of ancient DNA has relied almost exclusively on mitochondrial DNA (mtDNA) sequences. However, a number of recent studies have reported the successful recovery of ancient nuclear DNA (nuDNA) sequences, thereby allowing the characterization of genetic loci directly involved in phenotypic traits of extinct taxa. It is well documented that postmortem damage in ancient mtDNA can lead to the generation of artifactual sequences. However, as yet no one has thoroughly investigated the damage spectrum in ancient nuDNA. By comparing clone sequences from 23 fossil specimens, recovered from environments ranging from permafrost to desert, we demonstrate the presence of miscoding lesion damage in both the mtDNA and nuDNA, resulting in insertion of erroneous bases during amplification. Interestingly, no significant differences in the frequency of miscoding lesion damage are recorded between mtDNA and nuDNA despite great differences in cellular copy numbers. For both mtDNA and nuDNA, we find significant positive correlations between total sequence heterogeneity and the rates of type 1 transitions (adenine --> guanine and thymine --> cytosine) and type 2 transitions (cytosine --> thymine and guanine --> adenine), respectively. Type 2 transitions are by far the most dominant and increase relative to those of type 1 with damage load. The results suggest that the deamination of cytosine (and 5-methyl cytosine) to uracil (and thymine) is the main cause of miscoding lesions in both ancient mtDNA and nuDNA sequences. We argue that the problems presented by postmortem damage, as well as problems with contamination from exogenous sources of conserved nuclear genes, allelic variation, and the reliance on single nucleotide polymorphisms, call for great caution in studies relying on ancient nuDNA sequences.

Animals↗

Role of activating fibroblast growth factor receptor 3 mutations in the development of bladder tumors.

PURPOSE: Bladder tumors develop through different molecular pathways. Recent reports suggest activating mutations of the fibroblast growth factor receptor 3 (FGFR3) gene as marker for the "papillary" pathway with good prognosis, in contrast to the more malignant "carcinoma in situ" (CIS) pathway. The aim of this clinical follow-up study was to investigate the role of FGFR3 mutations in bladder cancer development in a longitudinal study. EXPERIMENTAL DESIGN: We selected 85 patients with superficial bladder tumors, stratified into early (stage T(a)/grade 1-2, n = 35) and more advanced (either stage T(1) or grade 3, n = 50) developmental stages. The patients were followed prospectively, and metachronous tumors were included. We did screening for FGFR3 and TP53 mutations by direct bidirectional sequencing and for genome-wide molecular changes with microarray technology. RESULTS: A total of 43 of 85 cases (51%) showed activating mutations of FGFR3. The mutations were associated with papillary tumors of early developmental stage. However, after stratifying for developmental stage, FGFR3-mutated tumors showed the same malignant potential as wild-type tumors. Tumors with concomitant CIS were generally FGFR3 wild type. They were characterized by different patterns of chromosomal changes and gene expression signatures compared with FGFR3-mutated tumors, indicating different molecular pathways. CONCLUSIONS: FGFR3 mutations seem to have a central role in the early development of papillary bladder tumors. These tumors follow a common molecular pathway, which is different from tumors with concomitant CIS. FGFR3 mutations do not seem to play a role in bladder cancer progression.

Carcinoma in Situ↗

Sampling properties of random graphs: the degree distribution.

We discuss two sampling schemes for selecting random subnets from a network, random sampling and connectivity dependent sampling, and investigate how the degree distribution of a node in the network is affected by the two types of sampling. Here we derive a necessary and sufficient condition that guarantees that the degree distributions of the subnet and the true network belong to the same family of probability distributions. For completely random sampling of nodes we find that this condition is satisfied by classical random graphs; for the vast majority of networks this condition will, however, not be met. We furthermore discuss the case where the probability of sampling a node depends on the degree of a node and we find that even classical random graphs are no longer closed under this sampling regime. We conclude by relating the results to real Eschericia coli protein interaction network data.

Journal Article↗

Subnets of scale-free networks are not scale-free: sampling properties of networks.

Most studies of networks have only looked at small subsets of the true network. Here, we discuss the sampling properties of a network's degree distribution under the most parsimonious sampling scheme. Only if the degree distributions of the network and randomly sampled subnets belong to the same family of probability distributions is it possible to extrapolate from subnet data to properties of the global network. We show that this condition is indeed satisfied for some important classes of networks, notably classical random graphs and exponential random graphs. For scale-free degree distributions, however, this is not the case. Thus, inferences about the scale-free nature of a network may have to be treated with some caution. The work presented here has important implications for the analysis of molecular networks as well as for graph theory and the theory of networks in general.

Journal Article↗

High-density single nucleotide polymorphism array defines novel stage and location-dependent allelic imbalances in human bladder tumors.

Bladder cancer is a common disease characterized by multiple recurrences and an invasive disease course in more than 10% of patients. It is of monoclonal or oligoclonal origin and genomic instability has been shown at certain loci. We used a 10,000 single nucleotide polymorphism (SNP) array with an average of 2,700 heterozygous SNPs to detect allelic imbalances (AI) in 37 microdissected bladder tumors from 17 patients. Eight tumors represented upstaging from Ta to T1, eight from T1 to T2+, and one from Ta to T2+. The AI was strongly stage-dependent as four chromosomal arms showed AI in > 50% of Ta samples, eight in T1, and twenty-two in T2+ samples. The tumors showed stage-dependent clonality as 61.3% of AIs were reconfirmed in later T1 tumors and 84.4% in muscle-invasive tumors. Novel unstable chromosomal areas were identified at chromosomes 6q, 10p, 16q, 20p, 20q, and 22q. The tumors separated into two distinct groups, highly stable tumors (all Ta tumors) and unstable tumors (2/3 muscle-invasive). All 11 unstable tumors had lost chromosome 17p areas and 90% chromosome 8 areas affecting Netrin-1/UNC5D/MAP2K4 genes as well as others. AI was present at the TP53 locus in 10 out of 11 unstable tumors, whereas 6 had homozygous TP53 mutations. Tumor distribution pattern reflected AI as seven out of eight patients with additional upper urinary tract tumors had genomic stable bladder tumors (P < 0.05). These data show the power of high-resolution SNP arrays for defining clinically relevant AIs.

Allelic Imbalance↗

Identification of endogenous retroviral reading frames in the human genome.

BACKGROUND: Human endogenous retroviruses (HERVs) comprise a large class of repetitive retroelements. Most HERVs are ancient and invaded our genome at least 25 million years ago, except for the evolutionary young HERV-K group. The far majority of the encoded genes are degenerate due to mutational decay and only a few non-HERV-K loci are known to retain intact reading frames. Additional intact HERV genes may exist, since retroviral reading frames have not been systematically annotated on a genome-wide scale. RESULTS: By clustering of hits from multiple BLAST searches using known retroviral sequences we have mapped 1.1% of the human genome as retrovirus related. The coding potential of all identified HERV regions were analyzed by annotating viral open reading frames (vORFs) and we report 7836 loci as verified by protein homology criteria. Among 59 intact or almost-intact viral polyproteins scattered around the human genome we have found 29 envelope genes including two novel gammaretroviral types. One encodes a protein similar to a recently discovered zebrafish retrovirus (ZFERV) while another shows partial, C-terminal, homology to Syncytin (HERV-W/FRD). CONCLUSIONS: This compilation of HERV sequences and their coding potential provide a useful tool for pursuing functional analysis such as RNA expression profiling and effects of viral proteins, which may, in turn, reveal a role for HERVs in human health and disease. All data are publicly available through a database at http://www.retrosearch.dk.

Codon↗

The probability and chromosomal extent of trans-specific polymorphism.

Balancing selection may result in trans-specific polymorphism: the maintenance of allelic classes that transcend species boundaries by virtue of being more ancient than the species themselves. At the selected site, gene genealogies are expected not to reflect the species tree. Because of linkage, the same will be true for part of the surrounding chromosomal region. Here we obtain various approximations for the distribution of the length of this region and discuss the practical implications of our results. Our main finding is that the trans-specific region surrounding a single-locus balanced polymorphism is expected to be quite short, probably too short to be readily detectable. Thus lack of obvious trans-specific polymorphism should not be taken as evidence against balancing selection. When trans-specific polymorphism is obvious, on the other hand, it may be reasonable to argue that selection must be acting on multiple sites or that recombination is suppressed in the surrounding region.

Chromosomes↗