PubMed Health⌕ Search

Biomedical subjects

Gerton Lunter

Publications and source records attributed to Gerton Lunter.

7 recordsLinked to original sources

Genome-wide association meta-regression identifies stem cell lineage orchestration as a key driver of acne risk.

Over 85% of the population experience acne at some point in their lives, with its severity spanning a quantitative spectrum, from mild, transient outbreaks to more persistent, severe forms of the condition. Moderate to severe disease poses a substantial global burden arising from both the physical and psychological impacts of this highly visible condition. The analytical approach taken in this study aimed to address the impact of variation in the dichotomisation of acne case control status, driven by ascertainment and study design, on effect size estimates across independent genetic association studies of acne. Through a fixed intercept meta-regression framework, we combined evidence genome-wide for association with acne across studies in which case-control status had been ascertained in different settings, allowing for different severity threshold definitions. Across a combined sample of 73,997 cases and 1,103,940 controls of European, South Asian and African American ancestry we identify genetic variation at 165 genomic loci that influence acne risk. There is evidence for both shared and ancestry specific components to the genetic susceptibility to acne and for sex differences in the magnitude of effect of risk alleles at three loci. We observe that common genetic variation explains 13.4% of acne heritability on the liability scale. Consistent with the hypothesis that genetic risk primarily operates at the level of individual pilosebaceous units, a polygenic score derived from this case-control study of acne susceptibility is associated with both self-reported and clinically assessed acne severity in adolescence, further strengthening the link between genetic risk and disease severity. Prioritisation of causal genes at the identified acne risk loci, provides genetic validation of the targets of established and emerging acne therapies, including retinoid treatments. The identified acne risk loci are enriched for genes encoding downstream effectors of RXRA signalling, including SOX9 and components of the WNT and p53 pathways. Illustrating that the control of stem cell lineage plasticity and cellular fate are important mechanisms through which genetic variation influences acne susceptibility within the pilosebaceous unit.

Journal Article↗

Novel type IV secretion system involved in propagation of genomic islands.

Type IV secretion systems (T4SSs) mediate horizontal gene transfer, thus contributing to genome plasticity, evolution of infectious pathogens, and dissemination of antibiotic resistance and other virulence traits. A gene cluster of the Haemophilus influenzae genomic island ICEHin1056 has been identified as a T4SS involved in the propagation of genomic islands. This T4SS is novel and evolutionarily distant from the previously described systems. Mutation analysis showed that inactivation of key genes of this system resulted in a loss of phenotypic traits provided by a T4SS. Seven of 10 mutants with a mutation in this T4SS did not express the type IV secretion pilus. Correspondingly, disruption of the genes resulted in up to 100,000-fold reductions in conjugation frequencies compared to those of the parent strain. Moreover, the expression of this T4SS was found to be positively regulated by one of its components, the tfc24 gene. We concluded that this gene cluster represents a novel family of T4SSs involved in propagation of genomic islands.

Bacterial Proteins↗

Signatures of adaptive evolution within human non-coding sequence.

The human genome is often portrayed as consisting of three sequence types, each distinguished by their mode of evolution. Purifying selection is estimated to act on 2.5-5.0% of the genome, whereas virtually all remaining sequence is considered to have evolved neutrally and to be devoid of functionality. The third mode of evolution, positive selection of advantageous changes, is considered rare. Such instances have been inferred only for a handful of sites, and these lie almost exclusively within protein-coding genes. Nevertheless, the majority of positively selected sequence is expected to lie within the wealth of functional 'dark matter' present outside of the coding sequence. Here, we review the evolutionary evidence for the majority of human-conserved DNA lying outside of the protein-coding sequence. We argue that within this non-coding fraction lies at least 1 Mb of functional sequence that has accumulated many beneficial nucleotide replacements. Illuminating the functions of this adaptive dark matter will lead to a better understanding of the sequence changes that have shaped the innovative biology of our species.

Evolution, Molecular↗

Genome-wide identification of human functional DNA using a neutral indel model.

It has become clear that a large proportion of functional DNA in the human genome does not code for protein. Identification of this non-coding functional sequence using comparative approaches is proving difficult and has previously been thought to require deep sequencing of multiple vertebrates. Here we introduce a new model and comparative method that, instead of nucleotide substitutions, uses the evolutionary imprint of insertions and deletions (indels) to infer the past consequences of selection. The model predicts the distribution of indels under neutrality, and shows an excellent fit to human-mouse ancestral repeat data. Across the genome, many unusually long ungapped regions are detected that are unaccounted for by the neutral model, and which we predict to be highly enriched in functional DNA that has been subject to purifying selection with respect to indels. We use the model to determine the proportion under indel-purifying selection to be between 2.56% and 3.25% of human euchromatin. Since annotated protein-coding genes comprise only 1.2% of euchromatin, these results lend further weight to the proposition that more than half the functional complement of the human genome is non-protein-coding. The method is surprisingly powerful at identifying selected sequence using only two or three mammalian genomes. Applying the method to the human, mouse, and dog genomes, we identify 90 Mb of human sequence under indel-purifying selection, at a predicted 10% false-discovery rate and 75% sensitivity. As expected, most of the identified sequence represents unannotated material, while the recovered proportions of known protein-coding and microRNA genes closely match the predicted sensitivity of the method. The method's high sensitivity to functional sequence such as microRNAs suggest that as yet unannotated microRNA genes are enriched among the sequences identified. Furthermore, its independence of substitutions allowed us to identify sequence that has been subject to heterogeneous selection, that is, sequence subject to both positive selection with respect to substitutions and purifying selection with respect to indels. The ability to identify elements under heterogeneous selection enables, for the first time, the genome-wide investigation of positive selection on functional elements other than protein-coding genes.

Animals↗

Bayesian coestimation of phylogeny and sequence alignment.

BACKGROUND: Two central problems in computational biology are the determination of the alignment and phylogeny of a set of biological sequences. The traditional approach to this problem is to first build a multiple alignment of these sequences, followed by a phylogenetic reconstruction step based on this multiple alignment. However, alignment and phylogenetic inference are fundamentally interdependent, and ignoring this fact leads to biased and overconfident estimations. Whether the main interest be in sequence alignment or phylogeny, a major goal of computational biology is the co-estimation of both. RESULTS: We developed a fully Bayesian Markov chain Monte Carlo method for coestimating phylogeny and sequence alignment, under the Thorne-Kishino-Felsenstein model of substitution and single nucleotide insertion-deletion (indel) events. In our earlier work, we introduced a novel and efficient algorithm, termed the "indel peeling algorithm", which includes indels as phylogenetically informative evolutionary events, and resembles Felsenstein's peeling algorithm for substitutions on a phylogenetic tree. For a fixed alignment, our extension analytically integrates out both substitution and indel events within a proper statistical model, without the need for data augmentation at internal tree nodes, allowing for efficient sampling of tree topologies and edge lengths. To additionally sample multiple alignments, we here introduce an efficient partial Metropolized independence sampler for alignments, and combine these two algorithms into a fully Bayesian co-estimation procedure for the alignment and phylogeny problem. Our approach results in estimates for the posterior distribution of evolutionary rate parameters, for the maximum a-posteriori (MAP) phylogenetic tree, and for the posterior decoding alignment. Estimates for the evolutionary tree and multiple alignment are augmented with confidence estimates for each node height and alignment column. Our results indicate that the patterns in reliability broadly correspond to structural features of the proteins, and thus provides biologically meaningful information which is not existent in the usual point-estimate of the alignment. Our methods can handle input data of moderate size (10-20 protein sequences, each 100-200 bp), which we analyzed overnight on a standard 2 GHz personal computer. CONCLUSION: Joint analysis of multiple sequence alignment, evolutionary trees and additional evolutionary parameters can be now done within a single coherent statistical framework.

Algorithms↗

A nucleotide substitution model with nearest-neighbour interactions.

MOTIVATION: It is well known that neighbouring nucleotides in DNA sequences do not mutate independently of each other. In this paper, we introduce a context-dependent substitution model and derive an algorithm to calculate the likelihood of sequences evolving under this model. We use this algorithm to estimate neighbour-dependent substitution rates, as well as rates for dinucleotide substitutions, using a Bayesian sampling procedure. The model is irreversible, giving an arrow to time, and allowing the position of the root between a pair of sequences to be inferred without using out-groups. RESULTS: We applied the model upon aligned human-mouse non-coding data. Clear neighbour dependencies were observed, including 17-18-fold increased CpG to TpG/CpA rates compared with other substitutions. Root inference positioned the root halfway the mouse and human tips, suggesting an approximately clock-like behaviour of the irreversible part of the substitution process.

Algorithms↗