PubMed HealthSearch

SEARCH · PubMed Health

Results for “Splicing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Improving spliced alignment by modeling splice sites with deep learning.

MOTIVATION: Spliced alignment refers to the alignment of messenger RNA (mRNA) or protein sequences to eukaryotic genomes. It plays a critical role in gene annotation and the study of gene functions. Accurate spliced alignment demands sophisticated modeling of splice sites, but current aligners use simple models, which may affect their accuracy given dissimilar sequences. RESULTS: We implemented minisplice to learn splice signals with a one-dimensional convolutional neural network (1D-CNN) and trained a model with 7,026 parameters for vertebrate and insect genomes. It captures conserved splice signals across phyla and reveals GC-rich introns specific to mammals and birds. We used this model to estimate the empirical splicing probability for every GT and AG in genomes, and modified minimap2 and miniprot to leverage pre-computed splicing probability during alignment. Evaluation on human long-read RNA-seq data and cross-species protein datasets showed our method greatly improves the junction accuracy especially for noisy long RNA-seq reads and proteins of distant homology. AVAILABILITY AND IMPLEMENTATION: https://github.com/lh3/minisplice.

Journal Article

Inefficient in vitro splicing of the regulatory intron of the L1 ribosomal protein gene of X.laevis depends on suboptimal splice site sequences.

The splicing of the third intron of the L1 r-protein gene of X.laevis was studied in the heterologous in vitro HeLa nuclear system. Despite the evolutionary distance, the cis-elements responsible for the default process play a similar role in the two organisms. Analysis of the splicing of various mutant substrates showed that the 5' splice site is primarily responsible for the low efficiency of splicing of the third intron. The suboptimal 5' splice site sequence leads to the utilization of an upstream alternative site which corresponds to the one utilized in vivo. The accumulation of splicing intermediates in the in vitro system allowed the identification of the branch site and of the branch consensus sequence. In contrast, the in vivo regulatory mechanism involving cleavage of the pre-mRNA is not mimicked in the HeLa extract.

Animals

A phosphorothioate at the 3' splice-site inhibits the second splicing step in a group I intron.

RNA polymerases can synthesize RNA containing phosphorothioate linkages in which a sulfur replaces one of the nonbridging oxygens. Only the Rp isomer is generated during transcription. A Rp phosphorothioate at the 5' splice-site of the Tetrahymena group I intron does not inhibit splicing (McSwiggen, J.A. and Cech, T.R. (1989) Science 244, 679). Transcription of mutants in which the first base of the 3' exon, U+1, was mutated to C or G, in the presence, respectively, of either cytosine or guanosine thiotriphosphate, introduced a phosphorothioate at the 3' splice-site. In both cases exon ligation was blocked. In the phosphorothioate substituted U+1G mutant, a new 3' splice-site was selected one base downstream of the correct site; despite the fact that the correct site was selected with very high fidelity in unsubstituted RNA. In contrast, the exon ligation reaction was successfully performed in reverse using unsubstituted intron RNA and ligated exons containing an Rp phosphorothioate at the exon junction site. Chirality was reversed during transesterification as in 5' splice-site cleavage (vide supra). This suggests that one non-bridging oxygen is particularly crucial for both splicing reactions.

Base Sequence

Expression of the tissue specific splicing protein SmN in neuronal cell lines and in regions of the brain with different splicing capacities.

The SmN protein is closely related to the ubiquitously expressed SmB and B' RNA splicing proteins but is expressed in only a limited range of tissues and cell types. The expression of SmN in a range of neuronal and non-neuronal cell lines correlates with their ability to splice the calcitonin/CGRP transcript to produce the mRNA encoding CGRP rather than that encoding calcitonin. Moreover, the SmN mRNA shows a widespread distribution within the brain and spinal ganglia being present in neuronal cells in all regions which naturally produce CGRP as well as in those areas which do not naturally express the calcitonin/CGRP gene but which can correctly splice the CGRP mRNA in transgenic mice expressing the calcitonin/CGRP gene in all cell types. Interestingly however the mRNA encoding SmN is also found in a few areas of the brain which can only carry out calcitonin-specific splicing in transgenic mice, such as the Purkinje layer of the cerebellum and the inferior colliculus. The possible role of SmN in the regulation of splicing in neuronal cells is discussed in the light of these results.

Animals

Site-specific cross-linking of mammalian U5 snRNP to the 5' splice site before the first step of pre-mRNA splicing.

We have used a site-specific cross-linking strategy to identify RNA and protein factors that interact with the 5' splice site region during mammalian pre-mRNA splicing. Two different pre-mRNA substrates were synthesized with a single 32P-labeled 4-thiouridine residue 2 nucleotides upstream of the 5' splice site. Selective photoactivation of the 4-thiouridine residue after incubation of either substrate under splicing conditions in HeLa nuclear extract resulted in cross-links to the U5 snRNA and the U5 snRNP protein p220. These ATP-dependent interactions occur before the first step of splicing. The U5 snRNA cross-links map to a phylogenetically invariant 9-nucleotide loop sequence and do not require Watson-Crick complementarity to the 5' exon. Cross-links of this position in the pre-mRNA to U1, but not to U2, U4, or U6 snRNAs, were also observed. The kinetics of U1 and U5 cross-link formation are similar, both peaking well before reaction intermediates appear.

Base Sequence

Trypanosoma brucei spliced-leader RNA methylations are required for trans splicing in vivo.

The Trypanosoma brucei spliced leader (SL) RNA donates its 5' leader sequence to all nuclear pre-mRNAs via trans RNA splicing. The SL RNA is a small-nuclear U RNA-like molecule which is present in the cell as part of a small ribonucleoprotein particle. However, unlike the trimethylguanosine-capped small nuclear U RNAs, the SL RNA has a highly modified 5' terminus containing an m7G cap and methylations on the first four transcribed nucleotides. Here, we show that incubation of procyclic-form T. brucei in the presence of the S-adenosylmethionine analog, sinefungin, leads to a rapid inhibition of SL RNA methylation. A concomitant inhibition of trans splicing and an accumulation of high-molecular-weight tubulin transcripts were also observed. The effects of sinefungin on SL RNA methylation and on trans splicing were correlated by labeling of cells incubated in the presence of the antibiotic. The results indicate that 5' modifications of the SL RNA are necessary for it to participate in trans splicing. SL RNA modification is not required for assembly of the core SL ribonucleoprotein, as these Cs2SO4-resistant particles can be formed with either methylated or undermethylated SL RNA.

Adenosine

A potential splicing factor is encoded by the opposite strand of the trans-spliced c-myb exon.

We previously established that the expression of a thymic c-myb mRNA species requires the intermolecular recombination of coding sequences expressed from transcriptional units localized on different chromosomes, in both chicken and human. We now report that a putative splicing factor (PR264), extremely well conserved in chicken and human, is encoded by the opposite strand of the c-myb trans-spliced exon. The PR264 polypeptide, which contains a typical ribonucleoprotein 80 and an arginine/serine-rich domain, is highly homologous to the Drosophila splicing regulators tra, tra-2, and su(wa) and to the human alternative splicing factor ASF/SF2. Furthermore, we show that PR264-specific mRNAs are expressed in normal hematopoietic cells of chicken and human origin and that the relative proportion of the PR264 transcripts is developmentally regulated in chicken.

Amino Acid Sequence

Characterization of the spectrum of alternative splicing of alpha 1 (XIII) collagen transcripts in HT-1080 cells and calvarial tissue resulted in identification of two previously unidentified alternatively spliced sequences, one previously unidentified exon, and nine new mRNA variants.

Amplification of a COL1-encoding region of alpha 1 (XIII) collagen transcripts of HT-1080 cell RNA suggested that exon 3 of the alpha 1 (XIII) collagen gene, which was previously deduced to be of 35 base pairs (bp) may consist of a constitutive 8-bp exon and an alternatively spliced 27-bp exon, termed here exons 3A and 3B, respectively. Furthermore, a previously unidentified alternatively spliced Gly-Xaa-Yaa-encoding exon designated as 4B was found between the sequences encoded by exons 4, redesignated here as 4A and 5. Six of the 16 potential combinations of the four consecutive alternatively spliced exons 3B, 4A, 4B, and 5 were found to exist in mRNAs, and as a result, the length of the COL1 domain may vary between 57 and 104 amino acid residues. Most of the NC2 domain is encoded by the alternatively spliced exons 12 and 13. Where previous analysis of cDNAs indicated that mRNA variants exist that contain either exon 12 or 13 sequences, amplification studies indicated here that there are also variants that lack both exons 12 and 13 but none that contain both exons simultaneously. Thus, the predicted length of this domain is either 12, 31, or 34 residues. Analyses covering both the COL1 and NC2 domains demonstrate that at least 12 mRNA species exist through the alternations of exons 3B-5, 12, and 13.

Abortion, Spontaneous

A base substitution at the splice acceptor site of intron 5 of the COL1A2 gene activates a cryptic splice site within exon 6 and generates abnormal type I procollagen in a patient with Ehlers-Danlos syndrome type VII.

The dermal type I collagen of a patient with Ehlers-Danlos type VIIB (EDS-VIIB) contained normal alpha 2(I) chains and mutant pN-alpha 2(I)' chains in which the amino-terminal propeptide (N-propeptide) remained attached to the alpha 2(I) chain. Similar alpha 2(I) chains were produced by cultured dermal fibroblasts. Amino acid sequencing of tryptic peptides, prepared from the mutant amino-terminal pN-alpha 2(I) CB1' peptide, indicated that five amino acids, including the N-proteinase (the specific proteinase that cleaves the procollagen N-propeptide) cleavage site, had been deleted from the junction of the N-propeptide and the N-telopeptide (the nonhelical domain at the amino-terminus of the alpha chains of fully processed type I polypeptide chains) of the mutant pro-alpha 2(I)' chain. The corresponding 15 nucleotides, which were deleted from approximately half of the alpha 2(I) cDNA polymerase chain reaction products, of the alpha 2(I) cDNA polymerase chain reaction products, were encoded by the +1 to +15 nucleotides of exon 6 of the normal alpha 2(I) gene (COL1A2). These 15 nucleotides were deleted in the splicing of alpha 2(I) pre-mRNA to mRNA as a result of inactivation of the 3' splice site of intron 5 by an AG to AC mutation and the activation of a cryptic AG splice acceptor site corresponding to positions +14 and +15 of exon 6. Loss of the N-proteinase cleavage site explained the persistence of the pN-alpha 2(I)' chains in the dermis and in fibroblast cultures. Collagen production by cultured dermal fibroblasts was doubled, possibly due to reduced feedback inhibition by the N-propeptides. In contrast to previously reported cases of EDS-VIIB, Lys5 of the N-telopeptide was not deleted and appeared to take part in the formation of intramolecular cross-linkages. However, increased collagen solubility and abnormal extraction profiles of the mutant type I collagen molecules indicated that collagen cross-linking was abnormal in the dermis. The proband and her son were heterozygous for the mutation. It is likely that the heterozygous loss of the N-proteinase cleavage site, with persistence of a shortened N-propeptide, was the major factor responsible for the EDS-VIIB phenotype.

Adult

Analysis of splice sites in the early region of bovine polyomavirus: evidence for a unique pattern of large T mRNA splicing.

The genetic organization of the early region of bovine polyomavirus (BPyV) was studied by analysis of the splice sites used in early mRNA maturation, using reverse transcription-polymerase chain reaction and DNA sequencing techniques. When compared to other polyomaviruses, the BPyV early region appears to have an uncommon organization. In the major early mRNA molecule two small intron sequences of 71 and 77 nucleotides, separated from one another by an 80 nucleotide exon sequence, were identified. Through splicing out both introns, a mRNA molecule is generated that contains an open reading frame with the capacity to encode 619 amino acids. Comparisons with the simian virus 40 large T antigen suggested that this mRNA molecule encodes the BPyV large T antigen. Remarkably, no mRNA product encoding a protein with a size comparable to that of the small t antigens of other polyomaviruses was detected. Another transcript was observed from which only the 77 nucleotide intron sequence had been removed, thereby creating a mRNA molecule with the capacity to encode only 45 amino acids. Whether this mRNA product represents a mature transcript which is translated in BPyV-infected cells or is an intermediate in the formation of the large T mRNA molecule is not known. Analysis of BPyV-specific early mRNA products isolated from BPyV-transformed murine cells revealed only the amplification product representing the putative large T antigen transcript.

Amino Acid Sequence

Mutations away from splice site recognition sequences might cis-modulate alternative splicing of goat alpha s1-casein transcripts. Structural organization of the relevant gene.

alpha s1-Casein variants F and D, synthesized in goat milk at lower levels than variant A, essentially differ from it by internal deletions of 37 and 11 amino acid residues, respectively. Northern blot analysis of mRNAs encoding alpha s1-casein F and A and sequencing of the relevant cloned cDNAs, as well as sequencing of in vitro amplified genomic fragments, revealed multiple alternatively processed transcripts, from the F allele. Although correctly spliced messengers were identified, most of the FmRNAs lacked three exons. These exons, further identified as exons 9, 10, and 11, together encode the 37 amino acid residues present in alpha s1-casein variant A but missing in variant F. Exon 9 codes for the sequence present in variant A but deleted in variant D. A single nucleotide deletion in exon 9 and two insertions, 11 and 3 base pairs in length, in the downstream intron, were identified as mutations potentially responsible for the alternative skipping of these 3 exons. From a computer-predicted secondary structure it appeared that the 11-base pair insertion might be involved in base-pairing interactions with the intron 5' splice site which might consequently be less accessible to U1 snRNA. We also report here the complete structural organization of the goat alpha s1-casein transcription unit, deduced from polymerase chain reaction experiments. It contains 19 exons scattered within a nucleotide stretch nearly 17-kilobase pairs long.

Amino Acid Sequence

A single-base change at a splice acceptor site in the ornithine aminotransferase gene causes abnormal RNA splicing in gyrate atrophy.

Gyrate atrophy (GA) is an autosomal recessive eye disease involving a progressive loss of vision due to chorioretinal degeneration in which the mitochondrial matrix enzyme ornithine aminotransferase (OAT) is defective. Two sisters with GA are described in this study in whom an A-to-G substitution at the 3' splice acceptor site of intron 4 in one allele of the OAT gene results in a truncated OAT mRNA devoid of exon 5 sequence. The mutation in the other allele was identified to be a mis-sense mutation at codon 318 by denaturing gradient gel electrophoresis and direct sequencing of the polymerase chain reaction (PCR)-amplified DNA. Thus, these GA patients are compound heterozygotes with respect to mutations in the OAT gene that result in inactivation of OAT.

Adenine

The RNA splicing factor PRPF8 is required for left-right organiser cilia differentiation and determination of cardiac left-right asymmetry via regulation of Arl13b splicing.

Cilia function in the left-right organizer (LRO) is critical for determining internal organ asymmetry in vertebrates. To further understand the genetics of left-right asymmetry, we isolated a mouse mutant with laterality defects, l11Jus27, from a random mutagenesis screen. l11Jus27 mutants carry a missense mutation in the pre-mRNA processing factor, Prpf8. cephalophŏnus (cph) mutant zebrafish, carrying a protein truncating mutation in prpf8, phenocopy the laterality defects of l11Jus27 mutants. Prpf8 mutant mouse and fish embryos have increased expression of an alternative transcript encoding the cilium-associated protein, ARL13B, that lacks exon 9. In zebrafish, over-expression of the arl13b transcript lacking exon 9 perturbed cilium formation and caused laterality defects. The shorter ARL13B protein isoform lacked interactions with intraflagellar transport proteins. Our data suggest that PRPF8 plays a prominent role in LRO cilia by through the regulation of alternative splicing of ARL13B, thus uncovering a new mechanism for cilia-linked developmental defects.

ARL13B

Alternatively-spliced p53 mRNA in the FAA-HTC1 rat hepatoma cell line without the splice site mutations.

A novel mutation of the p53 gene has been found in a rat hepatoma cell line, FAA-HTC1. This cell line carried two kinds of abnormal p53 transcripts; one lacked the exon 8 sequence, and the other had a single base substitution G to T which resulted in a new stop codon in exon 8. In the genomic DNA, this base substitution in exon 8 was present, indicating that both transcripts were transcribed from the mutated gene. No mutation was detected in its two flanking introns. In this cell line, the exon-deleted transcript seems to be generated by exon skipping due to an unknown mechanism other than splice site mutations.

Animals

Structure of the rat plasma membrane Ca(2+)-ATPase isoform 3 gene and characterization of alternative splicing and transcription products. Skeletal muscle-specific splicing results in a plasma membrane Ca(2+)-ATPase with a novel calmodulin-binding domain.

We have isolated the rat gene encoding isoform 3 of the plasma membrane Ca(2+)-ATPase (PMCA3) and have determined its exon/intron organization. The PMCA3 gene contains 24 exons and spans approximately 70 kilobases. In addition, we have analyzed the splicing and polyadenylation patterns leading to the production of an alternative 4.5-kilobase (PMCA3) skeletal muscle mRNA that differs from the previously characterized 7.5-kilobase brain mRNA (Greeb, J., and Shull, G. E. (1989) J. Biol. Chem. 264, 18569-18576). cDNA cloning, Northern blot hybridization, and polymerase chain reaction analyses of the 4.5-kilobase mRNA demonstrate (i) the inclusion of a novel 68-nucleotide exon (exon 22) that is specific for skeletal muscle and significantly alters the calmodulin-binding domain and (ii) the utilization of an alternative polyadenylation site following exon 23 which eliminates the last coding exon (exon 24) and 3'-untranslated sequence of the 7.5-kilobase mRNA. We have also identified a 42-nucleotide exon (exon 8) that is included in the skeletal muscle PMCA3 mRNAs, but may be either included or excluded in the brain mRNAs. Exon 8 is inserted immediately before the sequence encoding a putative phospholipid binding domain and thus may alter regulatory interactions of the enzyme with acidic phospholipids.

Amino Acid Sequence

Precursor RNA structural patterns at SF3B1 mutation sensitive cryptic 3' splice sites.

SF3B1 is a core component of the spliceosome involved in branch point recognition and 3' splice site selection. The SF3B1 K700E mutation (lysine to glutamic acid) is common in myelodysplastic syndrome and other blood disorders. SF3B1 K700E mutants utilize novel cryptic 3' splice sites; however, the properties distinguishing SF3B1-sensitive splice junctions from other alternatively spliced junctions are unknown. We identify a subset of 192 cryptic 3' splice junctions with significantly altered use in SF3B1 K700E cells, termed SF3B1-sensitive cryptic 3' splice sites, and 2800 cryptic 3' splice sites used in SF3B1 wild-type, termed SF3B1-resistant. We find that SF3B1-sensitive cryptic 3' splice sites are embedded in extended polypyrimidine tracts. Furthermore, canonical splice sites paired to SF3B1-sensitive cryptic 3' splice sites are significantly weaker than canonical 3' splice sites paired to SF3B1-resistant cryptic 3' splice sites. We test whether SF3B1-sensitive splice sites are structurally different from SF3B1-resistant 3' splice sites using chemical probing. We develop experimental RNA structure data for 83 SF3B1-sensitive junctions and 39 SF3B1-resistant junctions. We find that the pattern of structural accessibility at the NAG splicing motif in cryptic and canonical 3' splice sites is similar. However, the magnitude of accessibility differences is less in paired SF3B1-sensitive splice sites than in paired SF3B1-mutant splice sites. Additionally, SF3B1-sensitive splice junctions are more flexible than SF3B1-resistant junctions. Our results suggest that SF3B1-sensitive splice junctions have unique structure and sequence properties, containing poorly differentiated, weak splice sites that lead to altered 3' splice site recognition in the presence of SF3B1 mutation.

RNA Splicing Factors

The mutational spectrum of single base-pair substitutions in mRNA splice junctions of human genes: causes and consequences.

A total of 101 different examples of point mutations, which lie in the vicinity of mRNA splice junctions, and which have been held to be responsible for a human genetic disease by altering the accuracy of efficiency of mRNA splicing, have been collated. These data comprise 62 mutations at 5' splice sites, 26 at 3' splice sites and 13 that result in the creation of novel splice sites. It is estimated that up to 15% of all point mutations causing human genetic disease result in an mRNA splicing defect. Of the 5' splice site mutations, 60% involved the invariant GT dinucleotide; mutations were found to be non-randomly distributed with an excess over expectation at positions +1 and +2, and apparent deficiencies at positions -1 and -2. Of the 3' splice site mutations, 87% involved the invariant AG dinucleotide; an excess of mutations over expectation was noted at position -2. This non-randomness of mutation reflects the evolutionary conservation apparent in splice site consensus sequences drawn up previously from primate genes, and is most probably attributable to detection bias resulting from the differing phenotypic severity of specific lesions. The spectrum of point mutations was also drastically skewed: purines were significantly over-represented as substituting nucleotides, perhaps because of steric hindrance (e.g. in U1 snRNA binding at 5' splice sites). Furthermore, splice sites affected by point mutations resulting in human genetic disease were markedly different from the splice site consensus sequences. When similarity was quantified by a 'consensus value', both extremely low and extremely high values were notably absent from the wild-type sequences of the mutated splice sites. Splice sites of intermediate similarity to the consensus sequence may thus be more prone to the deleterious effects of mutation. Regarding the phenotypic effects of mutations on mRNA splicing, exon skipping occurred more frequently than cryptic splice site usage. Evidence is presented that indicates that, at least for 5' splice site mutations, cryptic splice site usage is favoured under conditions where (1) a number of such sites are present in the immediate vicinity and (2) these sites exhibit sufficient homology to the splice site consensus sequence for them to be able to compete successfully with the mutated splice site. The novel concept of a "potential for cryptic splice site usage" value was introduced in order to quantify these characteristics, and to predict the relative proportion of exon skipping vs cryptic splice site utilization consequent to the introduction of a mutation at a normal splice site.

Consensus Sequence