PubMed Health⌕ Search

Biomedical subjects

Marek Kimmel

Publications and source records attributed to Marek Kimmel.

At least 19 recordsLinked to original sources

Meta-evolutionary exome analysis identifies novel type 2 diabetes mellitus genes in the UK Biobank and all of us.

Type 2 diabetes mellitus (T2DM) risk is heavily influenced by genetics, yet current association tests have explained only parts of its heritability. We developed MEVA (Meta-Evolutionary Action), a meta-analytic framework that integrates three complementary methods-EAML, Sigma-Diff, and GeneEMBED-to assess the functional burden of protein-coding variants using evolutionary data. MEVA was applied to exome data from 28,115 T2DM cases and 28,115 controls in the UK Biobank (UKB), identifying 101 genes (p&#x2009;<&#x2009;1e-5). MEVA outperformed its component methods, each of which substantially outperformed a conventional burden test (MAGMA), in recovering known T2DM genes (AUROC&#x2009;=&#x2009;0.925) and maintaining robustness in progressively smaller cohorts (AUROC&#x2009;=&#x2009;0.917). MEVA showed significant enrichment for T2DM-related loci (p&#x2009;=&#x2009;6.8e-10, p&#x2009;=&#x2009;2.0e-34), protein interactions (z&#x2009;=&#x2009;4.6, z&#x2009;=&#x2009;4.2), pathways (p&#x2009;=&#x2009;1.3e-6, z&#x2009;=&#x2009;2.0), phenotypes (p&#x2009;=&#x2009;1.3e-21, z&#x2009;=&#x2009;9.1), and literature mentions (z&#x2009;=&#x2009;7.2). Replication in 16,915 T2DM cases and 16,915 controls from All of Us (AoU) yielded 99 genes (p&#x2009;<&#x2009;1e-5), 23 of which were also recovered in the UKB cohort - far exceeding random chance. These included established genes (SLC30A8, WFS1, HNF1A) and less-characterized candidates (NRIP1, ADAM30, CALCOCO2, TUBB1, ZFP36L2, WDR90). Notably, NRIP1 loss-of-function variants were associated with increased T2DM risk in both the UKB (OR = 1.09, FDR&#x2009;=&#x2009;5.4e-4) and AoU (OR = 1.09, FDR&#x2009;=&#x2009;0.046), and TUBB1 and CALCOCO2 gain-of-function variants showed consistent risk effects (FDR&#x2009;<&#x2009;0.05). Pathway analyses revealed convergence on endoplasmic reticulum chaperone complexes (FDR&#x2009;=&#x2009;0.02) and Hippo signaling (FDR&#x2009;=&#x2009;8.5e-4). Finally, all 177 candidate genes were functionally prioritized using ten orthogonal criteria to guide experimental follow-up. These results demonstrate that combining complementary, impact-aware association tests increases sensitivity, improves replication, and expands the catalog of genetic risk factors for T2DM.

Humans↗

Reaction-diffusion approach to modeling of the spread of early tumors along linear or tubular structures.

We consider a system composed of a tubular sheet of early tumor cells, occupying the surface of a structure existing in the organism. We assume that the cells have a potential for proliferation in response to a growth factor. This model can be thought of as representing an early stage (pre-in situ) of tumor evolution. A biomedical example of such process might be the atypical adenomatous hyperplasia in the lung. Destabilization of the equilibrium in such system represents initial invasion of cancer. We are looking for a transition from a slightly perturbed equilibrium state to uncontrolled and irregular growth. We examine a mathematical model of a population of cells distributed over a linear or tubular structure. Growth of cells is regulated by a growth factor, which can diffuse over the structure. Aside from this, production of cells and of the growth factor is governed by a pair of ordinary differential equations. Equation for the cell number follows from an accepted model of cell cycle. Equation for the bounded receptor particle number follows from a time-continuous Markov process. We demonstrate existence of the solutions of the complete model, using the method of invariant rectangles. We find conditions under which diffusion causes destabilization of the spatially homogeneous steady state, leading to exponential growth and apparently chaotic spatial patterns, following a period of almost constancy. This phenomenon may serve as a mathematical explanation of "unexpected" rapid growth and invasion of temporarily stable structures composed of cancer cells.

Animals↗

Recurrent use of evolutionary importance for functional annotation of proteins based on local structural similarity.

The annotation of protein function has not kept pace with the exponential growth of raw sequence and structure data. An emerging solution to this problem is to identify 3D motifs or templates in protein structures that are necessary and sufficient determinants of function. Here, we demonstrate the recurrent use of evolutionary trace information to construct such 3D templates for enzymes, search for them in other structures, and distinguish true from spurious matches. Serine protease templates built from evolutionarily important residues distinguish between proteases and other proteins nearly as well as the classic Ser-His-Asp catalytic triad. In 53 enzymes spanning 33 distinct functions, an automated pipeline identifies functionally related proteins with an average positive predictive power of 62%, including correct matches to proteins with the same function but with low sequence identity (the average identity for some templates is only 17%). Although these template building, searching, and match classification strategies are not yet optimized, their sequential implementation demonstrates a functional annotation pipeline which does not require experimental information, but only local molecular mimicry among a small number of evolutionarily important residues.

Algorithms↗

Strength of the purifying selection against different categories of the point mutations in the coding regions of the human genome.

Using available Information on the total absolute size of the coding region of the human genome, data on codon usage and pseudogene-derived mutation rates for different single nucleotide substitutions we have estimated, for the human genome, the potential numbers of mutation events capable to produce: (1) nonsense; (2) missense (radical and conservative); (3) silent; (4) splice; and (5) protein-elongating (those changing wild-type stop codon into an amino acid encoding codon) mutations. We used the NCBI dbSNP database to retrieve data on the observed number of polymorphisms of each category. The fraction of polymorphisms in each category among all potential events in the genome depends on the strength of selection: the higher the rate of polymorphism, the weaker the selection. We used nonsense mutations as a referent group. Compared with nonsense mutations, we found that the relative selection coefficient against protein-elongating mutations was 21%, and the relative selection was 12% against missense mutations. Radical missense mutations were found to be four times more deleterious compared to conservative ones. Surprisingly, we found that silent mutations on average are not neutral; with the average harmfulness of 3% of nonsense mutations. Silent mutations may be deleterious when they affect splicing by creating cryptic donor-acceptor sites or by disturbing exonic splicing enhancers (ESESs). The average selection coefficient against splice mutations was 48% of that against nonsense mutations. Converting the relative selection coefficients into absolute ones using data on loss-of-function mutations in Saccharomyces cerevisiae and Caenorhabditis elegans, or by analysis of the expected frequency of mutations in the human genome, suggested that genetic drift could play a role in population dynamics of conservative missense and silent mutations.

Computational Biology↗

A simple model of linkage disequilibrium and genetic drift in human genomic SNPs: importance of demography and SNP age.

We propose a simple model of evolution at a pair of SNP loci, under mutation, genetic drift and recombination. The developed model allows to consider evolution of SNPs under different demographic scenarios. We applied it to SNP data containing polymorphisms spanning 19 gene regions. We initially matched the linkage disequilibrium (LD) data only, and then we reconciled both LD and heterozygosity data. The imbalance between LD and heterozygosity data, observed for some of the analyzed genomic regions, may be a signature of selection acting in these regions. However, assuming neutrality, we obtain estimates of the age of population expansion of modern humans, which are consistent with the consensus estimates. In addition, we are able to estimate the ages of the polymorphisms observed in different genomic regions and we find that they vary widely with respect to their age. Polymorphisms at loci implicated in human disease, seem to be younger than average. Our results supplement the conclusions originally obtained by Reich and co-workers for the same set of data.

Genetic Drift↗

Stochastic search gene suggestion: a Bayesian hierarchical model for gene mapping.

Mapping the genes for a complex disease, such as diabetes or rheumatoid arthritis (RA), involves finding multiple genetic loci that may contribute to the onset of the disease. Pairwise testing of the loci leads to the problem of multiple testing. Looking at haplotypes, or linear sets of loci, avoids multiple tests but results in a contingency table with sparse counts, especially when using marker loci with multiple alleles. We propose a hierarchical Bayesian model for case-parent triad data that uses a conditional logistic regression likelihood to model the probability of transmission to a diseased child. We define hierarchical prior distributions on the allele main effects to model the genetic dependencies present in the human leukocyte antigen (HLA) region of chromosome 6. First, we add a hierarchical level for model selection that accounts for both locus and allele selection. This allows us to cast the problem of identifying genetic loci relevant to the disease into a problem of Bayesian variable selection. Second, we attempt to include linkage disequilibrium as a covariance structure in the prior for model coefficients. We evaluate the performance of the procedure with some simulated examples and then apply our procedure to identifying genetic markers in the HLA region that influence risk for RA. Our software is available on the website http://www.epigenetic.org/Linkage/ssgs-public/.

Alleles↗

Stochastic regulation in early immune response.

Living cells may be considered noisy or stochastic biochemical reactors. In eukaryotic cells, in which the number of protein or mRNA molecules is relatively large, the stochastic effects originate primarily in regulation of gene activity. Transcriptional activity of a gene can be initiated by transactivator molecules binding to the specific regulatory site(s) in the target gene. The stochasticity of activator binding and dissociation is amplified by transcription and translation, since target gene activation results in a burst of mRNAs molecules, and each copy of mRNA then serves as a template for numerous protein molecules. In this article, we reformulate our model of the NF-kappaB regulatory module to analyze a single cell regulation. Ordinary differential equations, used for description of fast reaction channels of processes involving a large number of molecules, are combined with a stochastic switch to account for the activity of the genes involved. The stochasticity in gene transcription causes simulated cells to exhibit large variability. Moreover, none of them behaves like an average cell. Although the average mRNA and protein levels remain constant before tumor necrosis factor (TNF) stimulation, and stabilize after a prolonged TNF stimulation, in any single cell these levels oscillate stochastically in the absence of TNF and keep oscillating under the prolonged TNF stimulation. However, in a short period of approximately 90 min, most cells are synchronized by the TNF signal, and exhibit similar kinetics. We hypothesize that this synchronization is crucial for proper activation of early genes controlling inflammation. Our theoretical predictions of single cell kinetics are supported by recent experimental studies of oscillations in NF-kappaB signaling made on single cells.

Animals↗

Transcriptional stochasticity in gene expression.

Due to the small number of copies of molecular species involved, such as DNA, mRNA and regulatory proteins, gene expression is a stochastic phenomenon. In eukaryotic cells, the stochastic effects primarily originate in regulation of gene activity. Transcription can be initiated by a single transcription factor binding to a specific regulatory site in the target gene. Stochasticity of transcription factor binding and dissociation is then amplified by transcription and translation, since target gene activation results in a burst of mRNA molecules, and each mRNA copy serves as a template for translating numerous protein molecules. In the present paper, we explore a mathematical approach to stochastic modeling. In this approach, the ordinary differential equations with a stochastic component for mRNA and protein levels in a single cells yield a system of first-order partial differential equations (PDEs) for two-dimensional probability density functions (pdf). We consider the following examples: Regulation of a single auto-repressing gene, and regulation of a system of two mutual repressors and of an activator-repressor system. The resulting PDEs are approximated by a system of many ordinary equations, which are then numerically solved.

Animals↗

simuPOP: a forward-time population genetics simulation environment.

SUMMARY: simuPOP is a forward-time population genetics simulation environment. The core of simuPOP is a scripting language (Python) that provides a large number of objects and functions to manipulate populations, and a mechanism to evolve populations forward in time. Using this R/Splus-like environment, users can create, manipulate and evolve populations interactively, or write a script and run it as a batch file. Owing to its flexible and extensible design, simuPOP can simulate large and complex evolutionary processes with ease. At a more user-friendly level, simuPOP provides an increasing number of built-in scripts that perform simulations ranging from implementation of basic population genetics models to generating datasets under complex evolutionary scenarios. AVAILABILITY: simuPOP is freely available at http://simupop.sourceforge.net, distributed under GPL license.

Algorithms↗

Estimating the growth rates of primary lung tumours from samples with missing measurements.

A method to estimate the population variability in tumour growth rate using incomplete data was developed. We assume exponential growth and lognormal distribution for the parameter of the growth curve. Estimates of growth rate obtained based on the cases with two measurements, one of which is obtained retrospectively, are biased towards lower growth rate. For the data sets where two measurements are available for some tumours and only one measurement for others (which means that no tumour was seen in retrospect for those cases), several approaches were developed that can eliminate or substantially reduce the bias. The relative error of the best estimates, as assessed by simulation, rarely exceeds 20 per cent. We found that the results of application of our estimation procedures to chest X-ray screening data agree well with the expectations.

Cell Growth Processes↗

Analysis of differences in amino acid substitution patterns, using multilevel G-tests.

In this paper, a new algorithm is presented, which makes possible multilevel comparison of BLOSUM protein substitution matrices based on data from different groups of organisms. As an example, a comparison between substitution matrices based on data from two groups of bacterial genomes with different GC content is presented. Our approach includes evaluating the number of amino acid pairs in BLOCKS databases created separately for the two groups of bacteria using protein sequences deposited in the COG database. Differences of distributions of amino acid pair counts are tested using the chi-squared based G-test. Different analysis levels make it possible to distinguish different patterns of amino acid substitution. Application of the algorithm reveals statistically significant differences in amino acid substitution patterns between AT-rich and GC-rich groups of bacterial organisms. The differences are particularly visible in the overall substitution pattern, amino acid conservation pattern and in comparison of substitution patterns for single amino acids. The algorithm presented in this paper can be considered a novel method for multi-level comparison of amino acid substitution patterns. The presented approach is not limited to bacterial organisms and BLOSUM substitution matrices. Statistically significant differences between substitution patterns in the two groups of bacterial organisms with respect to amino acid conservation pattern can be the evidence of different rate of evolutionary change between AT-rich and GC-rich bacterial organisms.

Algorithms↗

Algorithms for structural comparison and statistical analysis of 3D protein motifs.

The comparison of structural subsites in proteins is increasingly relevant to the prediction of their biological function. To address this problem, we present the Match Augmentation algorithm (MA). Given a structural motif of interest, such as a functional site, MA searches a target protein structure for a match: the set of atoms with the greatest geometric and chemical similarity. MA is extremely efficient because it exploits the fact that the amino acids in a structural motif are not equally important to function. Using motif residues ranked on functional significance via the Evolutionary Trace (ET), MA prioritizes its search by initially forming matches with functionally significant residues, then, guided by ET, it augments this partial match stepwise until the whole motif is found. With this hierarchical strategy, MA runs considerably faster than other methods, and almost always identifies matches in homologs known to have cognate functional sites. Second, in order to interpret matches, we further introduce a statistical method using nonparametric density estimation of the frequency distribution of structural matches. Our results show that the hierarchy of functional importance within structural motifs speeds up the search within targets, and points to a new method to score their statistical significance.

Amino Acid Sequence↗

Stochastic effects of multiple regulators on expression profiles in eukaryotes.

The stochastic nature of gene regulation still remains not fully understood. In eukaryotes, the stochastic effects are primarily attributable to the binary nature of genes, which are considered either switched "on" or "off" due to the action of the transcription factors binding to the promoter. In the time period when the gene is activated, bursts of mRNA transcript are produced. In the present paper, we investigate regulation of gene expression at the single cell level. We propose a mechanism of gene regulation, which is able to explain the observed distinct transcription profiles assuming the number of co-regulatory activities, without attempting to identify the specific proteins involved. The model is motivated by our experiments on NF-kappaB-dependent genes in HeLa cells. Our experimental data shows that NF-kappaB-dependent genes can be stratified into three characteristic groups according to their expression profiles: early, intermediate and late having maximum of expression at about 1, 3 and 6 h, respectively, from the beginning of TNF stimulation. We provide a tractable analytical approach, not only in the terms of expected expression profiles and their moments, which corresponds to the measurements on the cell population, but also in the terms of single cell behavior. Comparison between these two modes of description reveals that single cells behave qualitatively different from the cell population. This analysis provides insights useful for understanding of microarray experiments.

Animals↗

Mathematical model of NF-kappaB regulatory module.

The two-feedback-loop regulatory module of nuclear factor kappaB (NF-kappaB) signaling pathway is modeled by means of ordinary differential equations. The constructed model involves two-compartment kinetics of the activators IkappaB (IKK) and NF-kappaB, the inhibitors A20 and IkappaBalpha, and their complexes. In resting cells, the unphosphorylated IkappaBalpha binds to NF-kappaB and sequesters it in an inactive form in the cytoplasm. In response to extracellular signals such as tumor necrosis factor or interleukin-1, IKK is transformed from its neutral form (IKKn) into its active form (IKKa), a form capable of phosphorylating IkappaBalpha, leading to IkappaBalpha degradation. Degradation of IkappaBalpha releases the main activator NF-kappaB, which then enters the nucleus and triggers transcription of the inhibitors and numerous other genes. The newly synthesized IkappaBalpha leads NF-kappaB out of the nucleus and sequesters it in the cytoplasm, while A20 inhibits IKK converting IKKa into the inactive form (IKKi), a form different from IKKn, no longer capable of phosphorylating IkappaBalpha. After parameter fitting, the proposed model is able to properly reproduce time behavior of all variables for which the data are available: NF-kappaB, cytoplasmic IkappaBalpha, A20 and IkappaBalpha mRNA transcripts, IKK and IKK catalytic activity in both wild-type and A20-deficient cells. The model allows detailed analysis of kinetics of the involved proteins and their complexes and gives the predictions of the possible responses of whole kinetics to the change in the level of a given activator or inhibitor.

Animals↗

Asymptotic behavior of joint distributions of characteristics of a pair of randomly chosen individuals in discrete-time Fisher-Wright models with mutations and drift.

This is a continuation of the series of articles (C.R. Rao, D.N. Shanbhag (Eds.), Handbook of Statistics 19: Stochastic Processes: Theory and Methods, Elsevier Science, Amsterdam, 2001 (Chapter 8); Math. Biosci. 175 (2002) 83; Math. Meth. Appl. Sci. 26 (2003) 1587; Adv. Appl. Probab. 36 (2004) 57) devoted to a study of the interplay between two of the main forces of population genetics, mutations and drift, in the Fisher-Wright model. We provide discrete-time versions of theorems describing asymptotic behavior of joint distributions of characteristics of a pair of individuals in this model; their continuous-time counterparts were presented in the previous papers. Furthermore, we show that imbalance index, introduced in Kimmel et al. (Genetics 148 (1998) 1921) and King et al. (Mol. Biol. Evol. 17(12) (2000) 1895) in the context of continuous-time models, may also be used in discrete-time models to detect past population growth.

Genetic Drift↗

An accurate, sensitive, and scalable method to identify functional sites in protein structures.

Functional sites determine the activity and interactions of proteins and as such constitute the targets of most drugs. However, the exponential growth of sequence and structure data far exceeds the ability of experimental techniques to identify their locations and key amino acids. To fill this gap we developed a computational Evolutionary Trace method that ranks the evolutionary importance of amino acids in protein sequences. Studies show that the best-ranked residues form fewer and larger structural clusters than expected by chance and overlap with functional sites, but until now the significance of this overlap has remained qualitative. Here, we use 86 diverse protein structures, including 20 determined by the structural genomics initiative, to show that this overlap is a recurrent and statistically significant feature. An automated ET correctly identifies seven of ten functional sites by the least favorable statistical measure, and nine of ten by the most favorable one. These results quantitatively demonstrate that a large fraction of functional sites in the proteome may be accurately identified from sequence and structure. This should help focus structure-function studies, rational drug design, protein engineering, and functional annotation to the relevant regions of a protein.

Amino Acid Motifs↗

A note on estimation of dynamics of multiple gene expression based on singular value decomposition.

Recently, data on multiple gene expression at sequential time points were analyzed, using singular value decomposition (SVD) as a means to capture dominant trends, called characteristic modes, followed by fitting of a linear discrete-time dynamical system in which the expression values at a given time point are linear combinations of the values at a previous time point. We attempt to address several aspects of the method. To obtain the model we formulate a non-linear optimization problem and present how to solve it numerically using standard MATLAB procedures. We use publicly available data to test the approach. For reader's convenience, we provide a straightforward, ready-to-use, procedure in MATLAB, which employs its standard features to analyze data of this kind. Then, we investigate the sensitivity of the method to missing measurements and its possibilities to reconstruct missing data. Also, we discuss the possible consequences of data regularization, called sometimes 'polishing', on the outcome of analysis, especially when model is to be used for prediction purposes. Summarizing we point out that approximation of multiple gene expression data preceded by SVD provides some insight into the dynamics but may also lead to unexpected difficulties, like overfitting problems.

Algorithms↗

Molecular mechanisms of autosomal recessive hypercholesterolemia.

PURPOSE OF REVIEW: Autosomal recessive hypercholesterolemia (ARH) is a rare Mendelian dyslipidemia characterized by markedly elevated plasma LDL levels, xanthomatosis, and premature coronary artery disease. LDL receptor function is normal, or only moderately impaired in fibroblasts from ARH patients, but their cultured lymphocytes show increased cell-surface LDL binding, and impaired LDL degradation, consistent with a defect in LDL receptor internalization. Recently, the disorder was shown to be caused by mutations in a phosphotyrosine binding domain protein, ARH, which is required for internalization of low density lipoproteins in the liver. This review summarizes the findings of new investigations into the pathophysiology and molecular genetics of ARH. RECENT FINDINGS: All mutations that have been characterized to date preclude the synthesis of a full-length protein. GST-pulldown experiments indicate that the phosphotyrosine binding domain of ARH interacts with the internalization sequence (NPVY) in the cytoplasmic tail of LDLR, and that conserved motifs in the C-terminal portion of the protein bind to clathrin and to the beta2-adaptin subunit of AP-2. SUMMARY: The available data suggest that ARH functions as an adaptor protein that couples LDLR to the endocytic machinery.

Adaptor Proteins, Signal Transducing↗