PubMed Health⌕ Search

Biomedical subjects

Karsten Suhre

Publications and source records attributed to Karsten Suhre.

At least 19 recordsLinked to original sources

Refining the Genetic Contribution to Type 2 Diabetes Subtypes.

BACKGROUND: Type 2 diabetes (T2D) is a complex and highly heterogeneous disease driven in part by genetic predisposition and can be stratified into clinical subgroups to aid disease management. We recently grouped T2D subjects in the Qatar Biobank (QBB) cohort into Severe Insulin-Deficient Diabetes (SIDD), Severe Insulin-Resistant Diabetes (SIRD), Mild Obesity-Related Diabetes (MOD) and Mild Age-Related Diabetes (MARD) subtypes. Herein, we focused on the genetic makeup of these subtypes. METHODS: We used the QBB cohort (n = 13,808), of whom 2687 were with T2D, and comprehensively assessed polygenic risk scores (PGS) across T2D subtypes, investigated genetic loci associated with each subtype by leveraging the most recent and largest GWAS for T2D, evaluated SNP associations across T2D genetic clusters, and identified protein interaction pathways associated with these distinct T2D subtypes. RESULTS: MOD showed consistently lower PGS compared with other T2D subtypes across all tested scores. SIDD showed more associations with SNPs mapping to residual glycemic cluster compared with other T2D subtypes. The incremental analysis of PGS004838 demonstrated a high ΔAUC of 0.101 for SIDD and a moderate ΔAUC of 0.068 for SIRD, but not for MOD and MARD. Protein interaction analyses identified candidate subtype-associated gene networks linked to pathways related to glucose homeostasis in SIDD, insulin signalling and hepatic metabolism in SIRD, body fat distribution in MOD and vascular-related processes in MARD. CONCLUSION: We found heterogeneous genetic architectures across clinically defined T2D subtypes in a Middle Eastern population. Our findings provide evidence supporting differential polygenic burden, subtype genetic associations and subtype-associated biological pathways across T2D subtypes. These observations support the utility of subtype-based genetic analyses for improving biological understanding of T2D heterogeneity.

Humans↗

Determination of strongly overlapping signaling activity from microarray data.

BACKGROUND: As numerous diseases involve errors in signal transduction, modern therapeutics often target proteins involved in cellular signaling. Interpretation of the activity of signaling pathways during disease development or therapeutic intervention would assist in drug development, design of therapy, and target identification. Microarrays provide a global measure of cellular response, however linking these responses to signaling pathways requires an analytic approach tuned to the underlying biology. An ongoing issue in pattern recognition in microarrays has been how to determine the number of patterns (or clusters) to use for data interpretation, and this is a critical issue as measures of statistical significance in gene ontology or pathways rely on proper separation of genes into groups. RESULTS: Here we introduce a method relying on gene annotation coupled to decompositional analysis of global gene expression data that allows us to estimate specific activity on strongly coupled signaling pathways and, in some cases, activity of specific signaling proteins. We demonstrate the technique using the Rosetta yeast deletion mutant data set, decompositional analysis by Bayesian Decomposition, and annotation analysis using ClutrFree. We determined from measurements of gene persistence in patterns across multiple potential dimensionalities that 15 basis vectors provides the correct dimensionality for interpreting the data. Using gene ontology and data on gene regulation in the Saccharomyces Genome Database, we identified the transcriptional signatures of several cellular processes in yeast, including cell wall creation, ribosomal disruption, chemical blocking of protein synthesis, and, critically, individual signatures of the strongly coupled mating and filamentation pathways. CONCLUSION: This works demonstrates that microarray data can provide downstream indicators of pathway activity either through use of gene ontology or transcription factor databases. This can be used to investigate the specificity and success of targeted therapeutics as well as to elucidate signaling activity in normal and disease processes.

Algorithms↗

Mimivirus and the emerging concept of "giant" virus.

The recently discovered Acanthamoeba polyphaga Mimivirus is the largest known DNA virus. Its particle size (750 nm), genome length (1.2 million bp) and large gene repertoire (911 protein coding genes) blur the established boundaries between viruses and parasitic cellular organisms. In addition, the analysis of its genome sequence identified many types of genes never before encountered in a virus, including aminoacyl-tRNA synthetases and other central components of the translation machinery previously thought to be the signature of cellular organisms. In this article, we examine how the finding of such a giant virus might durably influence the way we look at microbial biodiversity, and lead us to revise the classification of microbial domains and life forms. We propose to introduce the word "girus" to recognize the intermediate status of these giant DNA viruses, the genome complexity of which makes them closer to small parasitic prokaryotes than to regular viruses.

Acanthamoeba↗

Estimation of prokaryote genomic DNA G+C content by sequencing universally conserved genes.

Determination of the DNA G+C content of prokaryotic genomes using traditional methods is time-consuming and results may vary from laboratory to laboratory, depending on the technique used. We explored the possibility of extrapolating the genomic DNA G+C content of prokaryotes from gene sequences. For this, 127 universally conserved genes were studied from 50 prokaryotic genomes in the Clusters of Orthologous Groups database. Of these, 57 genes were present as a single copy in the genomes of 157 different prokaryote species available in GenBank. There was a strong correlation [coefficient of determination (r2) >95 %] between the DNA G+C contents of 20 genes and their corresponding genomes. For each of the 157 prokaryotic genomes studied, the DNA G+C content of the 20 genes was used to determine a 'calculated' genome DNA G+C content (CGC) and this value was compared with the 'real' genome DNA G+C content (RGC). In order to select the most suitable gene for the determination of CGC values, we compared the r2 and median mol% difference between CGC and RGC as well as the sensitivity of each gene to provide CGC values for prokaryotic genomes that differ by less than 5 mol% from their RGC. The highly conserved ftsY gene (median size 1144 nucleotides), a vertically inherited member of the GTPase superfamily, showed the highest r2 value of 0.98, the smallest median mol% difference between CGC and RGC of 1.06 and a sensitivity of 100 %. Using ftsY DNA G+C content values, the CGC values of 100 genomes not included in the calculation of r2 differed by less than 5 mol% from their RGC values. These data suggest that the genomic DNA G+C content of prokaryotes may be estimated easily and reliably from the ftsY gene sequence.

Bacteria↗

Conformational flexibility of Mycobacterium tuberculosis thioredoxin reductase: crystal structure and normal-mode analysis.

The thioredoxin system exists ubiquitously and participates in essential antioxidant and redox-regulation processes via a pair of conserved cysteine residues. In Mycobacterium tuberculosis, which lacks a genuine glutathione system, the thioredoxin system provides reducing equivalents inside the cell. The three-dimensional structure of thioredoxin reductase from M. tuberculosis has been determined at 3 A resolution. TLS refinement reveals a large libration axis, showing that NADPH-binding domain has large anisotropic disorder. The relative rotation of the NADPH domain with respect to the FAD domain is necessary for the thioredoxin reduction cycle, as it brings the spatially distant reacting sites close together. Normal-mode analysis carried out based on the elastic network model shows that the motion required to bring about the functional conformational change can be accounted for by motion along one single mode. TLS refinement and normal-mode analysis thus enhance our understanding of the associated conformational changes.

Binding Sites↗

Phydbac "Gene Function Predictor": a gene annotation tool based on genomic context analysis.

BACKGROUND: The large amount of completely sequenced genomes allows genomic context analysis to predict reliable functional associations between prokaryotic proteins. Major methods rely on the fact that genes encoding physically interacting partners or members of shared metabolic pathways tend to be proximate on the genome, to evolve in a correlated manner and to be fused as a single sequence in another organism. RESULTS: The new "Gene Function Predictor", linked to the web server Phydbac proposes putative associations between Escherichia coli K-12 proteins derived from a combination of these methods. We show that associations made by this tool are more accurate than linkages found in the other established databases. Predicted assignments to GO categories, based on pre-existing functional annotations of associated proteins are also available. This new database currently holds 9,379 pairwise links at an expected success rate of at least 80%, the 6,466 functional predictions to GO terms derived from these links having a level of accuracy higher than 70%. CONCLUSION: The "Gene Function Predictor" is an automatic tool that aims to help biologists by providing them hypothetical functional predictions out of genomic context characteristics. The "Gene Function predictor" is available at http://www.igs.cnrs-mrs.fr/phydbac/indexPS.html.

Algorithms↗

Mimivirus gene promoters exhibit an unprecedented conservation among all eukaryotes.

The initial analysis of the recently sequenced genome of Acanthamoeba polyphaga Mimivirus, the largest known double-stranded DNA virus, predicted a proteome of size and complexity more akin to small parasitic bacteria than to other nucleocytoplasmic large DNA viruses and identified numerous functions never before described in a virus. It has been proposed that the Mimivirus lineage could have emerged before the individualization of cellular organisms from the three domains of life. An exhaustive in silico analysis of the noncoding moiety of all known viral genomes now uncovers the unprecedented perfect conservation of an AAAATTGA motif in close to 50% of the Mimivirus genes. This motif preferentially occurs in genes transcribed from the predicted leading strand and is associated with functions required early in the viral infectious cycle, such as transcription and protein translation. A comparison with the known promoter of unicellular eukaryotes, amoebal protists in particular, strongly suggests that the AAAATTGA motif is the structural equivalent of the TATA box core promoter element. This element is specific to the Mimivirus lineage and may correspond to an ancestral promoter structure predating the radiation of the eukaryotic kingdoms. This unprecedented conservation of core promoter regions is another exceptional feature of Mimivirus that again raises the question of its evolutionary origin.

Base Sequence↗

Mimivirus TyrRS: preliminary structural and functional characterization of the first amino-acyl tRNA synthetase found in a virus.

The amoeba-infecting Mimivirus is the largest known double-stranded DNA virus, with a 400 nm particle size, comparable to that of mycoplasma. The complete sequence of its 1.2 Mbp genome has recently been determined [Raoult et al. (2004), Science, 306, 1344-1350] and revealed numerous genes that were not expected to be found in a virus, such as genes encoding translation components, including 4-amino-acyl tRNA synthetases and homologues to various translation initiation, elongation and termination factors. A comprehensive structural and functional study of these Mimivirus gene products was initiated, as they may hold important clues about the origin of DNA viruses. Here, the first preliminary crystallographic and functional results obtained on one of these targets, Mimivirus TyrRS, are reported. Preliminary phasing was obtained using an original combination of homology modelling and normal mode analysis. Experimental evidence that Mimivirus tyrosyl tRNA synthetase recombinant gene product does indeed activate tyrosine is also presented.

Amino Acid Sequence↗

Gene and genome duplication in Acanthamoeba polyphaga Mimivirus.

Gene duplication is key to molecular evolution in all three domains of life and may be the first step in the emergence of new gene function. It is a well-recognized feature in large DNA viruses but has not been studied extensively in the largest known virus to date, the recently discovered Acanthamoeba polyphaga Mimivirus. Here, I present a systematic analysis of gene and genome duplication events in the mimivirus genome. I found that one-third of the mimivirus genes are related to at least one other gene in the mimivirus genome, either through a large segmental genome duplication event that occurred in the more remote past or through more recent gene duplication events, which often occur in tandem. This shows that gene and genome duplication played a major role in shaping the mimivirus genome. Using multiple alignments, together with remote-homology detection methods based on Hidden Markov Model comparison, I assign putative functions to some of the paralogous gene families. I suggest that a large part of the duplicated mimivirus gene families are likely to interfere with important host cell processes, such as transcription control, protein degradation, and cell regulatory processes. My findings support the view that large DNA viruses are complex evolving organisms, possibly deeply rooted within the tree of life, and oppose the paradigm that viral evolution is dominated by lateral gene acquisition, at least in regard to large DNA viruses.

Acanthamoeba↗

Bayesian decomposition analysis of bacterial phylogenomic profiles.

BACKGROUND: The past two decades have seen the appearance of new infectious diseases and the reemergence of old diseases previously thought to be under control. At the same time, the effectiveness of the existing antibacterials is rapidly decreasing due to the spread of multidrug-resistant pathogens. AIM: The aim of this study was to the identify candidate molecular targets (e.g. enzymes) within essential metabolic pathways specific to a significant subset of bacterial pathogens as the first step in the rational design of new antibacterial drugs. METHODS: We constructed a dataset of phylogenomic profiles (vectors that encode the similarity, measured by BLAST scores, of a gene across many species) for a series of 31 pathogenic bacteria of interest with 1073 genes taken from the reference organisms Escherichia coli and Mycobacterium tuberculosis. We applied Bayesian Decomposition, a matrix decomposition algorithm, to identify functional metabolic units comprising overlapping sets of genes in this dataset. RESULTS: Although no information on phylogeny was provided to the system, Bayesian Decomposition retrieved the known bacteria phylogenic relationships on the basis of the proteins necessary for survival. In addition, a set of genes required by all bacteria was identified, as well as components and enzymes specific to subsets of bacteria. CONCLUSION: The use of phylogenomic profiles and Bayesian Decomposition provide important insights for the design of new antibacterial therapeutics.

Algorithms↗

3DCoffee: combining protein sequences and structures within multiple sequence alignments.

Most bioinformatics analyses require the assembly of a multiple sequence alignment. It has long been suspected that structural information can help to improve the quality of these alignments, yet the effect of combining sequences and structures has not been evaluated systematically. We developed 3DCoffee, a novel method for combining protein sequences and structures in order to generate high-quality multiple sequence alignments. 3DCoffee is based on TCoffee version 2.00, and uses a mixture of pairwise sequence alignments and pairwise structure comparison methods to generate multiple sequence alignments. We benchmarked 3DCoffee using a subset of HOMSTRAD, the collection of reference structural alignments. We found that combining TCoffee with the threading program Fugue makes it possible to improve the accuracy of our HOMSTRAD dataset by four percentage points when using one structure only per dataset. Using two structures yields an improvement of ten percentage points. The measures carried out on HOM39, a HOMSTRAD subset composed of distantly related sequences, show a linear correlation between multiple sequence alignment accuracy and the ratio of number of provided structure to total number of sequences. Our results suggest that in the case of distantly related sequences, a single structure may not be enough for computing an accurate multiple sequence alignment.

Protein Conformation↗

Phydbac2: improved inference of gene function using interactive phylogenomic profiling and chromosomal location analysis.

Phydbac (phylogenomic display of bacterial genes) implemented a method of phylogenomic profiling using a distance measure based on normalized BLAST scores. This method was able to increase the predictive power of phylogenomic profiling by about 25% when compared to the classical approach based on Hamming distances. Here we present a major extension of Phydbac (named here Phydbac2), that extends both the concept and the functionality of the original web-service. While phylogenomic profiles remain the central focus of Phydbac2, it now integrates chromosomal proximity and gene fusion analyses as two additional non-similarity-based indicators for inferring pairwise gene functional relationships. Moreover, all presently available (January 2004) fully sequenced bacterial genomes and those of three lower eukaryotes are now included in the profiling process, thus increasing the initial number of reference genomes (71 in Phydbac) to 150 in Phydbac2. Using the KEGG metabolic pathway database as a benchmark, we show that the predictive power of Phydbac2 is improved by 27% over the previous version. This gain is accounted for on one hand, by the increased number of reference genomes (11%) and on the other hand, as a result of including chromosomal proximity into the distance measure (16%). The expanded functionality of Phydbac2 now allows the user to query more than 50 different genomes, including at least one member of each major bacterial group, most major pathogens and potential bio-terrorism agents. The search for co-evolving genes based on consensus profiles from multiple organisms, the display of Phydbac2 profiles side by side with COG information, the inclusion of KEGG metabolic pathway maps the production of chromosomal proximity maps, and the possibility of collecting and processing results from different Phydbac queries in a common shopping cart are the main new features of Phydbac2. The Phydbac2 web server is available at http://igs-server.cnrs-mrs.fr/phydbac/.

Artificial Gene Fusion↗

ElNemo: a normal mode web server for protein movement analysis and the generation of templates for molecular replacement.

Normal mode analysis (NMA) is a powerful tool for predicting the possible movements of a given macromolecule. It has been shown recently that half of the known protein movements can be modelled by using at most two low-frequency normal modes. Applications of NMA cover wide areas of structural biology, such as the study of protein conformational changes upon ligand binding, membrane channel opening and closure, potential movements of the ribosome, and viral capsid maturation. Another, newly emerging field of NMA is related to protein structure determination by X-ray crystallography, where normal mode perturbed models are used as templates for diffraction data phasing through molecular replacement (MR). Here we present ElNémo, a web interface to the Elastic Network Model that provides a fast and simple tool to compute, visualize and analyse low-frequency normal modes of large macro-molecules and to generate a large number of different starting models for use in MR. Due to the 'rotation-translation-block' (RTB) approximation implemented in ElNémo, there is virtually no upper limit to the size of the proteins that can be treated. Upon input of a protein structure in Protein Data Bank (PDB) format, ElNémo computes its 100 lowest-frequency modes and produces a comprehensive set of descriptive parameters and visualizations, such as the degree of collectivity of movement, residue mean square displacements, distance fluctuation maps, and the correlation between observed and normal-mode-derived atomic displacement parameters (B-factors). Any number of normal mode perturbed models for MR can be generated for download. If two conformations of the same (or a homologous) protein are available, ElNémo identifies the normal modes that contribute most to the corresponding protein movement. The web server can be freely accessed at http://igs-server.cnrs-mrs.fr/elnemo/index.html.

Crystallography, X-Ray↗

3DCoffee@igs: a web server for combining sequences and structures into a multiple sequence alignment.

This paper presents 3DCoffee@igs, a web-based tool dedicated to the computation of high-quality multiple sequence alignments (MSAs). 3D-Coffee makes it possible to mix protein sequences and structures in order to increase the accuracy of the alignments. Structures can be either provided as PDB identifiers or directly uploaded into the server. Given a set of sequences and structures, pairs of structures are aligned with SAP while sequence-structure pairs are aligned with Fugue. The resulting collection of pairwise alignments is then combined into an MSA with the T-Coffee algorithm. The server and its documentation are available from http://igs-server.cnrs-mrs.fr/Tcoffee/.

Internet↗

CaspR: a web server for automated molecular replacement using homology modelling.

Molecular replacement (MR) is the method of choice for X-ray crystallography structure determination when structural homologues are available in the Protein Data Bank (PDB). Although the success rate of MR decreases sharply when the sequence similarity between template and target proteins drops below 35% identical residues, it has been found that screening for MR solutions with a large number of different homology models may still produce a suitable solution where the original template failed. Here we present the web tool CaspR, implementing such a strategy in an automated manner. On input of experimental diffraction data, of the corresponding target sequence and of one or several potential templates, CaspR executes an optimized molecular replacement procedure using a combination of well-established stand-alone software tools. The protocol of model building and screening begins with the generation of multiple structure-sequence alignments produced with T-COFFEE, followed by homology model building using MODELLER, molecular replacement with AMoRe and model refinement based on CNS. As a result, CaspR provides a progress report in the form of hierarchically organized summary sheets that describe the different stages of the computation with an increasing level of detail. For the 10 highest-scoring potential solutions, pre-refined structures are made available for download in PDB format. Results already obtained with CaspR and reported on the web server suggest that such a strategy significantly increases the fraction of protein structures which may be solved by MR. Moreover, even in situations where standard MR yields a solution, pre-refined homology models produced by CaspR significantly reduce the time-consuming refinement process. We expect this automated procedure to have a significant impact on the throughput of large-scale structural genomics projects. CaspR is freely available at http://igs-server.cnrs-mrs.fr/Caspr/.

Escherichia coli Proteins↗

Set1 is required for meiotic S-phase onset, double-strand break formation and middle gene expression.

The Set1 protein of Saccharomyces cerevisiae is a histone methyltransferase (HMTase) acting on lysine 4 of histone H3. Inactivation of the SET1 gene in a diploid leads to a sporulation defect. We have studied various processes that take place during meiotic differentiation in set1delta diploid cells. The absence of Set1 leads to a delay of meiotic S-phase onset, which reflects a defect in DNA replication initiation. The timely induction of meiotic DNA replication does not require the Set1 HMTase activity, but depends on the SET domain. In addition, set1delta displays a severe impairment of the DNA double-strand break formation, which is not only the consequence of the replication delay. Transcriptional profiling experiments show that the induction of middle meiotic genes, but not of early meiotic genes, is affected by the loss of Set1. In contrast to meiotic replication, the transcriptional induction of the middle meiotic genes appears to depend on the methylation of H3-K4. Our results unveil multiple roles of Set1 in meiotic differentiation and distinguish between HMTase-dependent and -independent Set1 functions.

DNA Damage↗

On the potential of normal-mode analysis for solving difficult molecular-replacement problems.

Molecular replacement (MR) is the method of choice for X-ray crystallographic data phasing when structural data of suitable homologues are available. However, MR may fail even in cases of high sequence homology when conformational changes arising for example from ligand binding or different crystallogenic conditions come into play. In this work, the potential of normal-mode analysis as an extension to MR to allow recovery from such drawbacks is demonstrated. Three examples are presented in which screening for MR solutions with templates perturbed in the direction of one or two normal modes allows a valid MR solution to be found where MR using the original template failed to yield a model that could ultimately be refined. It has been shown recently that half of the known protein movements can be modelled by displacing the studied structure using at most two low-frequency normal modes. This suggests that normal-mode analysis has the potential to break tough MR problems in up to 50% of cases. Moreover, even in cases where an MR solution is available, this method can be used to further improve the starting model prior to refinement, eventually reducing the time spent on manual model construction (in particular for low-resolution data sets).

Carrier Proteins↗