PubMed Health⌕ Search

Biomedical subjects

Magnus Larsson

Publications and source records attributed to Magnus Larsson.

3 recordsLinked to original sources

Transcript identification by analysis of short sequence tags--influence of tag length, restriction site and transcript database.

There exist a number of gene expression profiling techniques that utilize restriction enzymes for generation of short expressed sequence tags. We have studied how the choice of restriction enzyme influences various characteristics of tags generated in an experiment. We have also investigated various aspects of in silico transcript identification that these profiling methods rely on. First, analysis of 14 248 mRNA sequences derived from the RefSeq transcript database showed that 1-30% of the sequences lack a given restriction enzyme recognition site. Moreover, 1-5% of the transcripts have recognition sites located less than 10 bases from the poly(A) tail. The uniqueness of 10 bp tags lies in the range 90-95%, which increases only slightly with longer tags, due to the existence of closely related transcripts. Furthermore, 3-30% of upstream 10 bp tags are identical to 3' tags, introducing a risk of misclassification if upstream tags are present in a sample. Second, we found that a sequence length of 16-17 bp, including the recognition site, is sufficient for unique transcript identification by BLAST based sequence alignment to the UniGene Human non-redundant database. Third, we constructed a tag-to-gene mapping for UniGene and compared it to an existing mapping database. The mappings agreed to 79-83%, where the selection of representative sequences in the UniGene clusters is the main cause of the disagreement. The results of this study may serve to improve the interpretation of sequence-based expression studies and the design of hybridization arrays, by identifying short tags that have a high reliability and separating them from tags that carry an inherent ambiguity in their capacity to discriminate between genes. To this end, supplementary information in the form of a web companion to this paper is located at http:// biobase.biotech.kth.se/tagseq.

Algorithms↗

Single-vector three-frame expression systems for affinity-tagged proteins.

An effort is presented to create expression vectors which would allow expression of an inserted gene fragment in three reading frames in a single vector from a single promoter but with three separate ribosome binding sites (RBS). Each expression frame would generate an in-frame fusion with an affinity tag to allow efficient recovery of the produced fusion proteins. In the first generation vector, three identical polyhistidyl tags (His(6)) were used as affinity tags for the three expression frames. In the second generation vector, three different tags, an albumin binding domain derived from streptococcal protein G, an IgG binding Staphylococcus aureus protein A-derived domain (Z) and a His(6) tag, were employed to allow frame-specific affinity recovery. To evaluate the systems, model genes have been inserted in three different frames in both vectors. The first vector was demonstrated to produce fusion proteins in all three frames, whereas for the second, with a much wider spacing between the RBSs and affinity tags, expression could only be demonstrated from the first two translational start sites. For both systems, the first translation start was found to be significantly favored over the others. Nevertheless, we believe that the presented results represent the first successful attempt to create single-vector three-frame expression systems, a concept that could become valuable in future combined cloning-expression vectors.

Bacterial Proteins↗

Gene expression analysis by signature pyrosequencing.

We describe a novel method for transcript profiling based on high-throughput parallel sequencing of signature tags using a non-gel-based microtiter plate format. The method relies on the identification of cDNA clones by pyrosequencing of the region corresponding to the 3'-end of the mRNA preceding the poly(A) tail. Simultaneously, the method can be used for gene discovery, since tags corresponding to unknown genes can be further characterized by extended sequencing. The protocol was validated using a model system for human atherosclerosis. Two 3'-tagged cDNA libraries, representing macrophages and foam cells, which are key components in the development of atherosclerotic plaques, were constructed using a solid phase approach. The libraries were analyzed by pyrosequencing, giving on average 25 bases. As a control, conventional expressed sequence tag (EST) sequencing using slab gel electrophoresis was performed. Homology searches were used to identify the genes corresponding to each tag. Comparisons with EST sequencing showed identical, unique matches in the majority of cases when the pyrosignature was at least 18 bases. A visualization tool was developed to facilitate differential analysis using a virtual chip format. The analysis resulted in identification of genes with possible relevance for development of atherosclerosis. The use of the method for automated massive parallel signature sequencing is discussed.

Base Sequence↗