PubMed Health⌕ Search

Biomedical subjects

Mihaela Zavolan

Publications and source records attributed to Mihaela Zavolan.

At least 19 recordsLinked to original sources

HTSinfer: inferring metadata from bulk Illumina RNA-Seq libraries.

SUMMARY: The Sequencing Read Archive is one of the largest and fastest-growing repositories of sequencing data, containing tens of petabytes of sequenced reads. Its data is used by a wide scientific community, often beyond the primary study that generated them. Such analyses rely on accurate metadata concerning the type of experiment and library, as well as the organism from which the sequenced reads were derived. These metadata are typically entered manually by contributors in an error-prone process, and are frequently incomplete. In addition, easy-to-use computational tools that verify the consistency and completeness of metadata describing the libraries to facilitate data reuse, are largely unavailable. Here, we introduce HTSinfer, a Python-based tool to infer metadata directly and solely from bulk RNA-sequencing data generated on Illumina platforms. HTSinfer leverages genome sequence information and diagnostic genes to rapidly and accurately infer the library source and library type, as well as the relative read orientation, 3' adapter sequence and read length statistics. HTSinfer is written in a modular manner, published under a permissible free and open-source license and encourages contributions by the community, enabling easy addition of new functionalities, e.g. for the inference of additional metrics, or the support of different experiment types or sequencing platforms. AVAILABILITY AND IMPLEMENTATION: HTSinfer is released under the Apache License 2.0. Latest code is available via GitHub at https://github.com/zavolanlab/htsinfer, while releases are published on Bioconda. A snapshot of the HTSinfer version described in this article was deposited at Zenodo at 10.5281/zenodo.13985958.

Metadata↗

Effects of Dicer and Argonaute down-regulation on mRNA levels in human HEK293 cells.

RNA interference and the microRNA (miRNA) pathway can induce sequence-specific mRNA degradation and/or translational repression. The human genome encodes hundreds of miRNAs that can post-transcriptionally repress thousands of genes. Using reporter constructs, we observed that degradation of mRNAs bearing sites imperfectly complementary to the endogenous let-7 miRNA is considerably stronger in human HEK293 than HeLa cells. The degradation did not result from the Ago2-mediated endonucleolytic cleavage but it was Dicer- and Ago2-dependent. We used this feature of HEK293 to address the size of a pool of transcripts regulated by RNA silencing in a single cell type. We generated HEK293 cell lines depleted of Dicer or individual Ago proteins. The cell lines were used for microarray analyses to obtain a comprehensive picture of RNA silencing. The 3'-untranslated region sequences of a few hundred transcripts that were commonly up-regulated upon Ago2 and Dicer knock-downs showed a significant enrichment of putative miRNA-binding sites. The up-regulation upon Ago2 and Dicer knock-downs was moderate and we found no evidence, at the mRNA level, for activation of silenced genes. Taken together, our data suggest that, independent of the effect on translation, miRNAs affect levels of a few hundred mRNAs in HEK293 cells.

3' Untranslated Regions↗

A novel class of small RNAs bind to MILI protein in mouse testes.

Small RNAs bound to Argonaute proteins recognize partially or fully complementary nucleic acid targets in diverse gene-silencing processes. A subgroup of the Argonaute proteins--known as the 'Piwi family'--is required for germ- and stem-cell development in invertebrates, and two Piwi members--MILI and MIWI--are essential for spermatogenesis in mouse. Here we describe a new class of small RNAs that bind to MILI in mouse male germ cells, where they accumulate at the onset of meiosis. The sequences of the over 1,000 identified unique molecules share a strong preference for a 5' uridine, but otherwise cannot be readily classified into sequence families. Genomic mapping of these small RNAs reveals a limited number of clusters, suggesting that these RNAs are processed from long primary transcripts. The small RNAs are 26-31 nucleotides (nt) in length--clearly distinct from the 21-23 nt of microRNAs (miRNAs) or short interfering RNAs (siRNAs)--and we refer to them as 'Piwi-interacting RNAs' or piRNAs. Orthologous human chromosomal regions also give rise to small RNAs with the characteristics of piRNAs, but the cloned sequences are distinct. The identification of this new class of small RNAs provides an important starting point to determine the molecular function of Piwi proteins in mammalian spermatogenesis.

Animals↗

The types and prevalence of alternative splice forms.

The finding that eukaryotic gene structures are extremely complex prompted the development of new experimental techniques for the accurate measurement of transcription start site usage and of the expression of alternative splice forms. On the computational side, analyses of large databases of splice variants revealed differences in the length, motif composition and selection pressure between constitutive and alternatively spliced exons. Such features are being incorporated into novel computational tools for gene structure prediction. The result of these investigations is a continuously improving catalogue of alternative splice forms. How the expression of these alternative splice forms is regulated remains one of the major open questions.

Alternative Splicing↗

SPA: a probabilistic algorithm for spliced alignment.

Recent large-scale cDNA sequencing efforts show that elaborate patterns of splice variation are responsible for much of the proteome diversity in higher eukaryotes. To obtain an accurate account of the repertoire of splice variants, and to gain insight into the mechanisms of alternative splicing, it is essential that cDNAs are very accurately mapped to their respective genomes. Currently available algorithms for cDNA-to-genome alignment do not reach the necessary level of accuracy because they use ad hoc scoring models that cannot correctly trade off the likelihoods of various sequencing errors against the probabilities of different gene structures. Here we develop a Bayesian probabilistic approach to cDNA-to-genome alignment. Gene structures are assigned prior probabilities based on the lengths of their introns and exons, and based on the sequences at their splice boundaries. A likelihood model for sequencing errors takes into account the rates at which misincorporation, as well as insertions and deletions of different lengths, occurs during sequencing. The parameters of both the prior and likelihood model can be automatically estimated from a set of cDNAs, thus enabling our method to adapt itself to different organisms and experimental procedures. We implemented our method in a fast cDNA-to-genome alignment program, SPA, and applied it to the FANTOM3 dataset of over 100,000 full-length mouse cDNAs and a dataset of over 20,000 full-length human cDNAs. Comparison with the results of four other mapping programs shows that SPA produces alignments of significantly higher quality. In particular, the quality of the SPA alignments near splice boundaries and SPA's mapping of the 5' and 3' ends of the cDNAs are highly improved, allowing for more accurate identification of transcript starts and ends, and accurate identification of subtle splice variations. Finally, our splice boundary analysis on the human dataset suggests the existence of a novel non-canonical splice site that we also find in the mouse dataset. The SPA software package is available at http://www.biozentrum.unibas.ch/personal/nimwegen/cgi-bin/spa.cgi.

Algorithms↗

A simple physical model predicts small exon length variations.

One of the most common splice variations are small exon length variations caused by the use of alternative donor or acceptor splice sites that are in very close proximity on the pre-mRNA. Among these, three-nucleotide variations at so-called NAGNAG tandem acceptor sites have recently attracted considerable attention, and it has been suggested that these variations are regulated and serve to fine-tune protein forms by the addition or removal of a single amino acid. In this paper we first show that in-frame exon length variations are generally overrepresented and that this overrepresentation can be quantitatively explained by the effect of nonsense-mediated decay. Our analysis allows us to estimate that about 50% of frame-shifted coding transcripts are targeted by nonsense-mediated decay. Second, we show that a simple physical model that assumes that the splicing machinery stochastically binds to nearby splice sites in proportion to the affinities of the sites correctly predicts the relative abundances of different small length variations at both boundaries. Finally, using the same simple physical model, we show that for NAGNAG sites, the difference in affinities of the neighboring sites for the splicing machinery accurately predicts whether splicing will occur only at the first site, splicing will occur only at the second site, or three-nucleotide splice variants are likely to occur. Our analysis thus suggests that small exon length variations are the result of stochastic binding of the spliceosome at neighboring splice sites. Small exon length variations occur when there are nearby alternative splice sites that have similar affinity for the splicing machinery.

Animals↗

Virus-encoded microRNAs: novel regulators of gene expression.

MicroRNAs (miRNAs) are a class of small RNAs that have recently been recognized as major regulators of gene expression. They influence diverse cellular processes ranging from cellular differentiation, proliferation, apoptosis and metabolism to cancer. Bioinformatic approaches and direct cloning methods have identified >3500 miRNAs, including orthologues from various species. Experiments to identify the targets and potential functions of miRNAs in various species are continuing but the recent discovery of virus-encoded miRNAs indicates that viruses also use this fundamental mode of gene regulation. Virus-encoded miRNAs seem to evolve rapidly and regulate both the viral life cycle and the interaction between viruses and their hosts.

Animals↗

Cell-type-specific signatures of microRNAs on target mRNA expression.

Although it is known that the human genome contains hundreds of microRNA (miRNA) genes and that each miRNA can regulate a large number of mRNA targets, the overall effect of miRNAs on mRNA tissue profiles has not been systematically elucidated. Here, we show that predicted human mRNA targets of several highly tissue-specific miRNAs are typically expressed in the same tissue as the miRNA but at significantly lower levels than in tissues where the miRNA is not present. Conversely, highly expressed genes are often enriched in mRNAs that do not have the recognition motifs for the miRNAs expressed in these tissues. Together, our data support the hypothesis that miRNA expression broadly contributes to tissue specificity of mRNA expression in many human tissues. Based on these insights, we apply a computational tool to directly correlate 3' UTR motifs with changes in mRNA levels upon miRNA overexpression or knockdown. We show that this tool can identify functionally important 3' UTR motifs without cross-species comparison.

3' Untranslated Regions↗

Identification of clustered microRNAs using an ab initio prediction method.

BACKGROUND: MicroRNAs (miRNAs) are endogenous 21 to 23-nucleotide RNA molecules that regulate protein-coding gene expression in plants and animals via the RNA interference pathway. Hundreds of them have been identified in the last five years and very recent works indicate that their total number is still larger. Therefore miRNAs gene discovery remains an important aspect of understanding this new and still widely unknown regulation mechanism. Bioinformatics approaches have proved to be very useful toward this goal by guiding the experimental investigations. RESULTS: In this work we describe our computational method for miRNA prediction and the results of its application to the discovery of novel mammalian miRNAs. We focus on genomic regions around already known miRNAs, in order to exploit the property that miRNAs are occasionally found in clusters. Starting with the known human, mouse and rat miRNAs we analyze 20 kb of flanking genomic regions for the presence of putative precursor miRNAs (pre-miRNAs). Each genome is analyzed separately, allowing us to study the species-specific identity and genome organization of miRNA loci. We only use cross-species comparisons to make conservative estimates of the number of novel miRNAs. Our ab initio method predicts between fifty and hundred novel pre-miRNAs for each of the considered species. Around 30% of these already have experimental support in a large set of cloned mammalian small RNAs. The validation rate among predicted cases that are conserved in at least one other species is higher, about 60%, and many of them have not been detected by prediction methods that used cross-species comparisons. A large fraction of the experimentally confirmed predictions correspond to an imprinted locus residing on chromosome 14 in human, 12 in mouse and 6 in rat. Our computational tool can be accessed on the world-wide-web. CONCLUSION: Our results show that the assumption that many miRNAs occur in clusters is fruitful for the discovery of novel miRNAs. Additionally we show that although the overall miRNA content in the observed clusters is very similar across the three considered species, the internal organization of the clusters changes in evolution.

Algorithms↗

The developmental miRNA profiles of zebrafish as determined by small RNA cloning.

MicroRNAs (miRNAs) represent a family of small, regulatory, noncoding RNAs that are found in plants and animals. Here, we describe the miRNA profile of the zebrafish Danio rerio resolved in a developmental and cell-type-specific manner. The profiles were obtained from larger-scale sequencing of small RNA libraries prepared from developmentally staged zebrafish, and two adult fibroblast cell lines derived from the caudal fin (ZFL) and the liver epithelium (SJD). We identified a total of 154 distinct miRNAs expressed from 343 miRNA genes. Other experimental/computational sources support an additional 10 miRNAs encoded by 19 genes. The miRNAs can be classified into 87 distinct families. Cross-species comparison indicates that 81 families are conserved in mammals, 17 of which also have at least one member conserved in an invertebrate. Our analysis reveals that the zygotes are essentially devoid of miRNAs and that their expression begins during the blastula period with a zebrafish-specific family of miRNAs encoded by closely spaced multicopy genes. Computational predictions of zebrafish miRNA targets are provided that take into account the depth of evolutionary conservation. Besides miRNAs, we identified a prominent class of repeat-associated small interfering RNAs (rasiRNAs).

Animals↗

Identification of microRNAs of the herpesvirus family.

Epstein-Barr virus (EBV or HHV4), a member of the human herpesvirus (HHV) family, has recently been shown to encode microRNAs (miRNAs). In contrast to most eukaryotic miRNAs, these viral miRNAs do not have close homologs in other viral genomes or in the genome of the human host. To identify other miRNA genes in pathogenic viruses, we combined a new miRNA gene prediction method with small-RNA cloning from several virus-infected cell types. We cloned ten miRNAs in the Kaposi sarcoma-associated virus (KSHV or HHV8), nine miRNAs in the mouse gammaherpesvirus 68 (MHV68) and nine miRNAs in the human cytomegalovirus (HCMV or HHV5). These miRNA genes are expressed individually or in clusters from either polymerase (pol) II or pol III promoters, and share no substantial sequence homology with one another or with the known human miRNAs. Generally, we predicted miRNAs in several large DNA viruses, and we could neither predict nor experimentally identify miRNAs in the genomes of small RNA viruses or retroviruses.

Chromosome Mapping↗

Analysis of human immunodeficiency virus cytopathicity by using a new method for quantitating viral dynamics in cell culture.

Human immunodeficiency virus (HIV) causes complex metabolic changes in infected CD4(+) T cells that lead to cell cycle arrest and cell death by necrosis. To study the viral functions responsible for deleterious effects on the host cell, we quantitated the course of HIV type 1 infection in tissue cultures by using flow cytometry for a virally encoded marker protein, heat-stable antigen (HSA). We found that HSA appeared on the surface of the target cells in two phases: passive acquisition due to association and fusion of virions with target cells, followed by active protein expression from transcription of the integrated provirus. The latter event was necessary for decreased target cell viability. We developed a general mathematical model of viral dynamics in vitro in terms of three effective time-dependent rates: those of cell proliferation, infection, and death. Using this model we show that the predominant contribution to the depletion of viable target cells results from direct cell death rather than cell cycle blockade. This allows us to derive accurate bounds on the time-dependent death rates of infected cells. We infer that the death rate of HIV-infected cells is 80 times greater than that of uninfected cells and that the elimination of the vpr protein reduces the death rate by half. Our approach provides a general method for estimating time-dependent death rates that can be applied to study the dynamics of other viruses.

Antigens, CD↗

Identification of virus-encoded microRNAs.

RNA silencing processes are guided by small RNAs that are derived from double-stranded RNA. To probe for function of RNA silencing during infection of human cells by a DNA virus, we recorded the small RNA profile of cells infected by Epstein-Barr virus (EBV). We show that EBV expresses several microRNA (miRNA) genes. Given that miRNAs function in RNA silencing pathways either by targeting messenger RNAs for degradation or by repressing translation, we identified viral regulators of host and/or viral gene expression.

Animals↗

The small RNA profile during Drosophila melanogaster development.

Small RNAs ranging in size between 20 and 30 nucleotides are involved in different types of regulation of gene expression including mRNA degradation, translational repression, and chromatin modification. Here we describe the small RNA profile of Drosophila melanogaster as a function of development. We have cloned and sequenced over 4000 small RNAs, 560 of which have the characteristics of RNase III cleavage products. A nonredundant set of 62 miRNAs was identified. We also isolated 178 repeat-associated small interfering RNAs (rasiRNAs), which are cognate to transposable elements, satellite and microsatellite DNA, and Suppressor of Stellate repeats, suggesting that small RNAs participate in defining chromatin structure. rasiRNAs are most abundant in testes and early embryos, where regulation of transposon activity is critical and dramatic changes in heterochromatin structure occur.

Animals↗

Impact of alternative initiation, splicing, and termination on the diversity of the mRNA transcripts encoded by the mouse transcriptome.

We analyzed the FANTOM2 clone set of 60,770 RIKEN full-length mouse cDNA sequences and 44,122 public mRNA sequences. We developed a new computational procedure to identify and classify the forms of splice variation evident in this data set and organized the results into a publicly accessible database that can be used for future expression array construction, structural genomics, and analyses of the mechanism and regulation of alternative splicing. Statistical analysis shows that at least 41% and possibly as much as 60% of multiexon genes in mouse have multiple splice forms. Of the transcription units with multiple splice forms, 49% contain transcripts in which the apparent use of an alternative transcription start (stop) is accompanied by alternative splicing of the initial (terminal) exon. This implies that alternative transcription may frequently induce alternative splicing. The fact that 73% of all exons with splice variation fall within the annotated coding region indicates that most splice variation is likely to affect the protein form. Finally, we compared the set of constitutive (present in all transcripts) exons with the set of cryptic (present only in some transcripts) exons and found statistically significant differences in their length distributions, the nucleotide distributions around their splice junctions, and the frequencies of occurrence of several short sequence motifs.

Alternative Splicing↗

Decay rates of human mRNAs: correlation with functional characteristics and sequence attributes.

Although mRNA decay rates are a key determinant of the steady-state concentration for any given mRNA species, relatively little is known, on a population level, about what factors influence turnover rates and how these rates are integrated into cellular decisions. We decided to measure mRNA decay rates in two human cell lines with high-density oligonucleotide arrays that enable the measurement of decay rates simultaneously for thousands of mRNA species. Using existing annotation and the Gene Ontology hierarchy of biological processes, we assign mRNAs to functional classes at various levels of resolution and compare the decay rate statistics between these classes. The results show statistically significant organizational principles in the variation of decay rates among functional classes. In particular, transcription factor mRNAs have increased average decay rates compared with other transcripts and are enriched in "fast-decaying" mRNAs with half-lives <2 h. In contrast, we find that mRNAs for biosynthetic proteins have decreased average decay rates and are deficient in fast-decaying mRNAs. Our analysis of data from a previously published study of Saccharomyces cerevisiae mRNA decay shows the same functional organization of decay rates, implying that it is a general organizational scheme for eukaryotes. Additionally, we investigated the dependence of decay rates on sequence composition, that is, the presence or absence of short mRNA motifs in various regions of the mRNA transcript. Our analysis recovers the positive correlation of mRNA decay with known AU-rich mRNA motifs, but we also uncover further short mRNA motifs that show statistically significant correlation with decay. However, we also note that none of these motifs are strong predictors of mRNA decay rate, indicating that the regulation of mRNA decay is more complex and may involve the cooperative binding of several RNA-binding proteins at different sites.

Base Composition↗

Systematic characterization of the zinc-finger-containing proteins in the mouse transcriptome.

Zinc-finger-containing proteins can be classified into evolutionary and functionally divergent protein families that share one or more domains in which a zinc ion is tetrahedrally coordinated by cysteines and histidines. The zinc finger domain defines one of the largest protein superfamilies in mammalian genomes;46 different conserved zinc finger domains are listed in InterPro (http://www.ebi.ac.uk/InterPro). Zinc finger proteins can bind to DNA, RNA, other proteins, or lipids as a modular domain in combination with other conserved structures. Owing to this combinatorial diversity, different members of zinc finger superfamilies contribute to many distinct cellular processes, including transcriptional regulation, mRNA stability and processing, and protein turnover. Accordingly, mutations of zinc finger genes lead to aberrations in a broad spectrum of biological processes such as development, differentiation, apoptosis, and immunological responses. This study provides the first comprehensive classification of zinc finger proteins in a mammalian transcriptome. Specific detailed analysis of the SP/Krüppel-like factors and the E3 ubiquitin-ligase RING-H2 families illustrates the importance of such an analysis for a more comprehensive functional classification of large protein families. We describe the characterization of a new family of C2H2 zinc-finger-containing proteins and a new conserved domain characteristic of this family, the identification and characterization of Sp8, a new member of the Sp family of transcriptional regulators, and the identification of five new RING-H2 proteins.

Alternative Splicing↗

SMASHing regulatory sites in DNA by human-mouse sequence comparisons.

Regulatory sequence elements provide important clues to understanding and predicting gene expression. Although the binding sites for hundreds of transcription factors are known, there has been no systematic attempt to incorporate this information in the annotation of the human genome. Cross species sequence comparisons are critical to a meaningful annotation of regulatory elements since they generally reside in conserved non-coding regions. To take advantage of the recently completed drafts of the mouse and human genomes for annotating transcription factor binding sites, we developed SMASH, a computational pipeline that identifies thousands of orthologous human/ mouse proteins, maps them to genomic sequences, extracts and compares upstream regions and annotates putative regulatory elements in conserved, non-coding, upstream regions. Our current dataset consists of approximately 2,500 human/mouse gene pairs. Transcription start sites were estimated by mapping quasi-full length cDNA sequences. SMASH uses a novel probabilistic method to identify putative conserved binding sites that takes into account the competition between transcription factors for binding DNA. SMASH presents the results via a genome browser web interface which displays the predicted regulatory information together with the current annotations for the human genome. Our results are validated by comparison to previously published experimental data. SMASH results compare favorably to other existing computational approaches.

Algorithms↗