PubMed Health⌕ Search

Biomedical subjects

Peter D Stenson

Publications and source records attributed to Peter D Stenson.

9 recordsLinked to original sources

Genome-wide detection of human 5' UTR variants that impact protein translation.

The 5' untranslated region (5' UTR) of messenger RNAs (mRNAs) plays a central role in regulating protein synthesis initiation, particularly through the Kozak sequence and upstream open reading frames (uORFs). Genetic variants within these regulatory elements could affect translation, altering gene expression and contributing to clinical phenotypes in humans. We developed a computational method called 5ULTRA (5' Untranslated Region Annotation) for analysis of whole-exome sequencing and whole-genome sequencing data to detect, annotate, and prioritize 5' UTR variants with potential translation impact. 5ULTRA identifies single-nucleotide variants, indels, and splicing variants that affect uORFs by creating or disrupting start/stop codons and that alter Kozak sequence strength of either the uORFs or the main coding sequence. 5ULTRA incorporates recent uORF databases and provides comprehensive annotations. 5ULTRA implements a machine-learning score to prioritize candidate variants with predicted effects on translation and also provides specific mechanistic predictions. The score correlates strongly with experimentally measured protein-level effects of 5' UTR variants. We applied 5ULTRA to multiple genetics datasets across diverse disease contexts, identifying candidate variants including potential cancer-driving somatic mutations predicted to decrease ABI1 level or increase NRAS abundance; common variants associated with traits such as multiple sclerosis, lung function, and cardiovascular function, by altering protein levels of TAGAP, VRTN, and SPAAR, respectively; and rare germline variants in our cohort, including a splicing variant of RPSA leading to 5' UTR sequence alteration that causes congenital asplenia and a variant of TNF that could predispose to tuberculosis.

Humans↗

A systematic analysis of LINE-1 endonuclease-dependent retrotranspositional events causing human genetic disease.

Diverse long interspersed element-1 (LINE-1 or L1)-dependent mutational mechanisms have been extensively studied with respect to L1 and Alu elements engineered for retrotransposition in cultured cells and/or in genome-wide analyses. To what extent the in vitro studies can be held to accurately reflect in vivo events in the human genome, however, remains to be clarified. We have attempted to address this question by means of a systematic analysis of recent L1-mediated retrotranspositional events that have caused human genetic disease, with a view to providing a more complete picture of how L1-mediated retrotransposition impacts upon the architecture of the human genome. A total of 48 such mutations were identified, including those described as L1-mediated retrotransposons, as well as insertions reported to contain a poly(A) tail: 26 were L1 trans-driven Alu insertions, 15 were direct L1 insertions, four were L1 trans-driven SVA insertions, and three were associated with simple poly(A) insertions. The systematic study of these lesions, when combined with previous in vitro and genome-wide analyses, has strengthened several important conclusions regarding L1-mediated retrotransposition in humans: (a) approximately 25% of L1 insertions are associated with the 3' transduction of adjacent genomic sequences, (b) approximately 25% of the new L1 inserts are full-length, (c) poly(A) tail length correlates inversely with the age of the element, and (d) the length of target site duplication in vivo is rarely longer than 20 bp. Our analysis also suggests that some 10% of L1-mediated retrotranspositional events are associated with significant genomic deletions in humans. Finally, the identification of independent retrotranspositional events that have integrated at the same genomic locations provides new insight into the L1-mediated insertional process in humans.

Base Sequence↗

Meta-analysis of gross insertions causing human genetic disease: novel mutational mechanisms and the role of replication slippage.

Although gross insertions (>20 bp) comprise <1% of disease-causing mutations, they nevertheless represent an important category of pathological lesion. In an attempt to study these insertions in a systematic way, 158 gross insertions ranging in size between 21 bp and approximately 10 kb were identified using the Human Gene Mutation Database (www.hgmd.org). A careful meta-analytical study revealed extensive diversity in terms of the nature of the inserted DNA sequence and has provided new insights into the underlying mutational mechanisms. Some 70% of gross insertions were found to represent sequence duplications of different types (tandem, partial tandem, or complex). Although most of the tandem duplications were explicable by simple replication slippage, the three complex duplications appear to result from multiple slippage events. Some 11% of gross insertions were attributable to nonpolyglutamine repeat expansions (including octapeptide repeat expansions in the prion protein gene [PRNP] and polyalanine tract expansions) and evidence is presented to support the contention that these mutations are also caused by replication slippage rather than by unequal crossing over. Some 17% of gross insertions, all >or=276 bp in length, were found to be due to LINE-1 (L1) retrotransposition involving different types of element (L1 trans-driven Alu, L1 direct, and L1 trans-driven SVA). A second example of pathological mitochondrial-nuclear sequence transfer was identified in the USH1C gene but appears to arise via a novel mechanism, trans-replication slippage. Finally, evidence for another novel mechanism of human genetic disease, involving the possible capture of DNA oligonucleotides, is presented in the context of a 26-bp insertion into the ERCC6 gene.

Base Sequence↗

Complex gene rearrangements caused by serial replication slippage.

The now-classical model of replication slippage can in principle account for both simple deletions and tandem duplications associated with short direct repeats. Invariably, a single replication slippage event is invoked, irrespective of whether simple deletions or tandem duplications are involved. However, we recently identified three complex duplicational insertions that could also be accounted for by a model of serial replication slippage. We postulate that a sizeable proportion of hitherto inexplicable complex gene rearrangements may be explained by such a model. To test this idea, and to assess the generality of our initial findings, a number of complex gene rearrangements were selected from the Human Gene Mutation Database (HGMD). Some 95% (20/21) of these mutations were found to be explicable by twin or multiple rounds of replication slippage, the sole exception being a double deletion in the F9 gene that is associated with DNA sequences that appear capable of adopting non-B conformations. Of the 20 complex gene rearrangements, 19 (seven simple double deletions, one triple deletion, two double mutational events comprising a simple deletion and a simple insertion, six simple indels that may constitute a novel and non-canonical class of gene conversion, and three complex indels) were compatible with the model of serial replication slippage in cis; the remaining indel in the MECP2 gene, however, appears to have arisen via interchromosomal replication slippage in trans. Our postulate that serial replication slippage may account for a variety of complex gene rearrangements has therefore received broad support from the study of the above diverse series of mutations.

Base Sequence↗

Microdeletions and microinsertions causing human genetic disease: common mechanisms of mutagenesis and the role of local DNA sequence complexity.

In the Human Gene Mutation Database (www.hgmd.org), microdeletions and microinsertions causing inherited disease (both defined as involving < or = 20 bp of DNA) account for 8,399 (17%) and 3,345 (7%) logged mutations, in 940 and 668 genes, respectively. A positive correlation was noted between the microdeletion and microinsertion frequencies for 564 genes for which both microdeletions and microinsertions are reported in HGMD, consistent with the view that the propensity of a given gene/sequence to undergo microdeletion is related to its propensity to undergo microinsertion. While microdeletions and microinsertions of 1 bp constitute respectively 48% and 66% of the corresponding totals, the relative frequency of the remaining lesions correlates negatively with the length of the DNA sequence deleted or inserted. Many of the microdeletions and microinsertions of more than 1 bp are potentially explicable in terms of slippage mutagenesis, involving the addition or removal of one copy of a mono-, di-, or trinucleotide tandem repeat. The frequency of in-frame 3-bp and 6-bp microinsertions and microdeletions was, however, found to be significantly lower than that of mutations of other lengths, suggesting that some of these in-frame lesions may not have come to clinical attention. Various sequence motifs were found to be over-represented in the vicinity of both microinsertions and microdeletions, including the heptanucleotide CCCCCTG that shares homology with the complement of the 8-bp human minisatellite conserved sequence/chi-like element (GCWGGWGG). The previously reported indel hotspot GTAAGT and its complement ACTTAC were also found to be overrepresented in the vicinity of both microinsertions and microdeletions, thereby providing a first example of a mutational hotspot that is common to different types of gene lesion. Other motifs overrepresented in the vicinity of microdeletions and microinsertions included DNA polymerase pause sites and topoisomerase cleavage sites. Several novel microdeletion/microinsertion hotspots were noted and some of these exhibited sufficient similarity to one another to justify terming them "super-hotspot" motifs. Analysis of sequence complexity also demonstrated that a combination of slipped mispairing mediated by direct repeats, and secondary structure formation promoted by symmetric elements, can account for the majority of microdeletions and microinsertions. Thus, microinsertions and microdeletions exhibit strong similarities in terms of the characteristics of their flanking DNA sequences, implying that they are generated by very similar underlying mechanisms.

Computational Biology↗

Intrachromosomal serial replication slippage in trans gives rise to diverse genomic rearrangements involving inversions.

Serial replication slippage in cis (SRScis) provides a plausible explanation for many complex genomic rearrangements that underlie human genetic disease. This concept, taken together with the intra- and intermolecular strand switch models that account for mutations that arise via quasipalindrome correction, suggest that intrachromosomal SRS in trans (SRStrans) mediated by short inverted repeats may also give rise to a diverse series of complex genomic rearrangements. If this were to be so, such rearrangements would invariably generate inversions. To test this idea, we collated all informative mutations involving inversions of >or=5 bp but <1 kb by screening the Human Gene Mutation Database (HGMD; www.hgmd.org) and conducting an extensive literature search. Of the 21 resulting mutations, only two (both of which coincidentally contain untemplated additions) were found to be incompatible with the SRStrans model. Eighteen (one simple inversion, six inversions involving sequence replacement by upstream or downstream sequence, five inversions involving the partial reinsertion of removed sequence, and six inversions that occurred in a more complicated context) of the remaining 19 mutations were found to be consistent with either two steps of intrachromosomal SRStrans or a combination of replication slippage in cis plus intrachromosomal SRStrans. The remaining lesion, a 31-kb segmental duplication associated with a small inversion in the SLC3A1 gene, is explicable in terms of a modified SRS model that integrates the concept of "break-induced replication." This study therefore lends broad support to our postulate that intrachromosomal SRStrans can account for a variety of complex gene rearrangements that involve inversions.

Amino Acid Transport Systems, Basic↗

Evolutionary conservation and selection of human disease gene orthologs in the rat and mouse genomes.

BACKGROUND: Model organisms have contributed substantially to our understanding of the etiology of human disease as well as having assisted with the development of new treatment modalities. The availability of the human, mouse and, most recently, the rat genome sequences now permit the comprehensive investigation of the rodent orthologs of genes associated with human disease. Here, we investigate whether human disease genes differ significantly from their rodent orthologs with respect to their overall levels of conservation and their rates of evolutionary change. RESULTS: Human disease genes are unevenly distributed among human chromosomes and are highly represented (99.5%) among human-rodent ortholog sets. Differences are revealed in evolutionary conservation and selection between different categories of human disease genes. Although selection appears not to have greatly discriminated between disease and non-disease genes, synonymous substitution rates are significantly higher for disease genes. In neurological and malformation syndrome disease systems, associated genes have evolved slowly whereas genes of the immune, hematological and pulmonary disease systems have changed more rapidly. Amino-acid substitutions associated with human inherited disease occur at sites that are more highly conserved than the average; nevertheless, 15 substituting amino acids associated with human disease were identified as wild-type amino acids in the rat. Rodent orthologs of human trinucleotide repeat-expansion disease genes were found to contain substantially fewer of such repeats. Six human genes that share the same characteristics as triplet repeat-expansion disease-associated genes were identified; although four of these genes are expressed in the brain, none is currently known to be associated with disease. CONCLUSIONS: Most human disease genes have been retained in rodent genomes. Synonymous nucleotide substitutions occur at a higher rate in disease genes, a finding that may reflect increased mutation rates in the chromosomal regions in which disease genes are found. Rodent orthologs associated with neurological function exhibit the greatest evolutionary conservation; this suggests that rodent models of human neurological disease are likely to most faithfully represent human disease processes. However, with regard to neurological triplet repeat expansion-associated human disease genes, the contraction, relative to human, of rodent trinucleotide repeats suggests that rodent loci may not achieve a 'critical repeat threshold' necessary to undergo spontaneous pathological repeat expansions. The identification of six genes in this study that have multiple characteristics associated with repeat expansion-disease genes raises the possibility that not all human loci capable of facilitating neurological disease by repeat expansion have as yet been identified.

Animals↗

Gross Rearrangement Breakpoint Database (GRaBD).

Translocations and gross gene deletions are an important cause of both cancer and inherited disease. Such DNA rearrangements are nonrandomly distributed in the human genome as a consequence of selection for growth advantage and/or the inherent potential of some DNA sequences to be particularly susceptible to breakage and recombination. The Gross Rearrangement Breakpoint Database (GRaBD; http://www.uwcm.ac.uk/uwcm/mg/grabd/) was established primarily for the analysis of the sequence context of translocation and deletion breakpoints in a search for characteristics that might have rendered these sequences prone to rearrangement. GRaBD, which contains 397 germline and somatic DNA breakpoint junction sequences derived from 219 different rearrangements underlying human inherited disease and cancer, is the only comprehensive collection of gross gene rearrangement breakpoint junctions currently available.

Chromosome Breakage↗

Human Gene Mutation Database (HGMD): 2003 update.

The Human Gene Mutation Database (HGMD) constitutes a comprehensive core collection of data on germ-line mutations in nuclear genes underlying or associated with human inherited disease (www.hgmd.org). Data catalogued includes: single base-pair substitutions in coding, regulatory and splicing-relevant regions; micro-deletions and micro-insertions; indels; triplet repeat expansions as well as gross deletions; insertions; duplications; and complex rearrangements. Each mutation is entered into HGMD only once in order to avoid confusion between recurrent and identical-by-descent lesions. By March 2003, the database contained in excess of 39,415 different lesions detected in 1,516 different nuclear genes, with new entries currently accumulating at a rate exceeding 5,000 per annum. Since its inception, HGMD has been expanded to include cDNA reference sequences for more than 87% of listed genes, splice junction sequences, disease-associated and functional polymorphisms, as well as links to data present in publicly available online locus-specific mutation databases. Although HGMD has recently entered into a licensing agreement with Celera Genomics (Rockville, MD), mutation data will continue to be made freely available via the Internet.

Databases, Genetic↗