PubMed Health⌕ Search

Biomedical subjects

Duane E Smailus

Publications and source records attributed to Duane E Smailus.

9 recordsLinked to original sources

Sequencing and analysis of 10,967 full-length cDNA clones from Xenopus laevis and Xenopus tropicalis reveals post-tetraploidization transcriptome remodeling.

Sequencing of full-insert clones from full-length cDNA libraries from both Xenopus laevis and Xenopus tropicalis has been ongoing as part of the Xenopus Gene Collection Initiative. Here we present 10,967 full ORF verified cDNA clones (8049 from X. laevis and 2918 from X. tropicalis) as a community resource. Because the genome of X. laevis, but not X. tropicalis, has undergone allotetraploidization, comparison of coding sequences from these two clawed (pipid) frogs provides a unique angle for exploring the molecular evolution of duplicate genes. Within our clone set, we have identified 445 gene trios, each comprised of an allotetraploidization-derived X. laevis gene pair and their shared X. tropicalis ortholog. Pairwise dN/dS, comparisons within trios show strong evidence for purifying selection acting on all three members. However, dN/dS ratios between X. laevis gene pairs are elevated relative to their X. tropicalis ortholog. This difference is highly significant and indicates an overall relaxation of selective pressures on duplicated gene pairs. We have found that the paralogs that have been lost since the tetraploidization event are enriched for several molecular functions, but have found no such enrichment in the extant paralogs. Approximately 14% of the paralogous pairs analyzed here also show differential expression indicative of subfunctionalization.

Animals↗

A high-throughput screen identifying sequence and promiscuity characteristics of the loxP spacer region in Cre-mediated recombination.

BACKGROUND: Cre-loxP recombination refers to the process of site-specific recombination mediated by two loxP sequences and the Cre recombinase protein. Transgenic experiments exploit integrative recombination, where a donor plasmid carrying a loxP site and DNA of interest integrate into a recipient loxP site in a target genome. Unfortunately, integrative recombination is highly inefficient because the insert is flanked by two loxP sites, which themselves become targets for Cre and lead to subsequent excision of the insert. A small number of mutations have been discovered in parts of the loxP sequence, specifically the spacer and inverted repeat segments, that increase the efficiency of integrative recombination. In this study we introduce a high-throughput in vitro assay to rapidly detect novel loxP spacer mutants and describe the sequence characteristics of successful recombinants. RESULTS: We created synthetic loxP oligonucleotides that contained a combination of inverted repeat mutations (the lox66 and lox71 mutations) and mutant spacer sequences, degenerate at 6 of the 8 positions. After in vitro Cre recombination, 3,124 recombinant clones were identified by sequencing. Included in this set were 31 unique, novel, self-recombining sequences. Using network visualization tools, we recognized 12 spacer sets with restricted promiscuity. We observed that increased guanine content at all spacer positions save for position 8 resulted in increased recombination. Interestingly, recombination between identical spacers was not preferred over non-identical spacers. We also identified a set of 16 pairs of loxP spacers that reacted at least twice with another spacer, but not themselves. Further, neither the wild-type P1 phage loxP sequence nor any of the known loxP spacer mutants appeared to be kinetically favoured by Cre recombinase. CONCLUSION: This study approached loxP spacer mutant screening in an unbiased manner, assuming nothing about candidate loxP sites save for the conserved 4 and 5 spacer positions. Candidate sites were free to recombine with any other sequence in the pool of all possible sites. The subset of loxP sites identified here are candidates for in vivo serial recombination as they have already demonstrated limited promiscuity with other loxP spacer and stability in the presence of Cre.

Bacteriophage P1↗

Simple, robust methods for high-throughput nanoliter-scale DNA sequencing.

We have developed high-throughput DNA sequencing methods that generate high quality data from reactions as small as 400 nL, providing an approximate order of magnitude reduction in reagent use relative to standard protocols. Sequencing of clones from plasmid, fosmid, and BAC libraries yielded read lengths (PHRED20 bases) of 765 +/- 172 (n = 10,272), 621 +/- 201 (n = 1824), and 647 +/- 189 (n = 568), respectively. Implementation of these procedures at high-throughput genome centers could have a substantial impact on the amount of data that can be generated per unit cost.

Nanotechnology↗

Development and application of a salmonid EST database and cDNA microarray: data mining and interspecific hybridization characteristics.

We report 80,388 ESTs from 23 Atlantic salmon (Salmo salar) cDNA libraries (61,819 ESTs), 6 rainbow trout (Oncorhynchus mykiss) cDNA libraries (14,544 ESTs), 2 chinook salmon (Oncorhynchus tshawytscha) cDNA libraries (1317 ESTs), 2 sockeye salmon (Oncorhynchus nerka) cDNA libraries (1243 ESTs), and 2 lake whitefish (Coregonus clupeaformis) cDNA libraries (1465 ESTs). The majority of these are 3' sequences, allowing discrimination between paralogs arising from a recent genome duplication in the salmonid lineage. Sequence assembly reveals 28,710 different S. salar, 8981 O. mykiss, 1085 O. tshawytscha, 520 O. nerka, and 1176 C. clupeaformis putative transcripts. We annotate the submitted portion of our EST database by molecular function. Higher- and lower-molecular-weight fractions of libraries are shown to contain distinct gene sets, and higher rates of gene discovery are associated with higher-molecular weight libraries. Pyloric caecum library group annotations indicate this organ may function in redox control and as a barrier against systemic uptake of xenobiotics. A microarray is described, containing 7356 salmonid elements representing 3557 different cDNAs. Analyses of cross-species hybridizations to this cDNA microarray indicate that this resource may be used for studies involving all salmonids.

Animals↗

Systematic recovery and analysis of full-ORF human cDNA clones.

The Mammalian Gene Collection (MGC) consortium (http://mgc.nci.nih.gov) seeks to establish publicly available collections of full-ORF cDNAs for several organisms of significance to biomedical research, including human. To date over 15,200 human cDNA clones containing full-length open reading frames (ORFs) have been identified via systematic expressed sequence tag (EST) analysis of a diverse set of cDNA libraries; however, further systematic EST analysis is no longer an efficient method for identifying new cDNAs. As part of our involvement in the MGC program, we have developed a scalable method for targeted recovery of cDNA clones to facilitate recovery of genes absent from the MGC collection. First, cDNA is synthesized from various RNAs, followed by polymerase chain reaction (PCR) amplification of transcripts in 96-well plates using gene-specific primer pairs flanking the ORFs. Amplicons are cloned into a sequencing vector, and full-length sequences are obtained. Sequences are processed and assembled using Phred and Phrap, and analyzed using Consed and a number of bioinformatics methods we have developed. Sequences are compared with the Reference Sequence (RefSeq) database, and validation of sequence discrepancies is attempted using other sequence databases including dbEST and dbSNP. Clones with identical sequence to RefSeq or containing only validated changes will become part of the MGC human gene collection. Clones containing novel splice variants or polymorphisms have also been identified. Our approach to clone recovery, applied at large scale, has the potential to recover many and possibly most of the genes absent from the MGC collection.

Cloning, Molecular↗

Novel avian influenza H7N3 strain outbreak, British Columbia.

Genome sequences of chicken (low pathogenic avian influenza [LPAI] and highly pathogenic avian influenza [HPAI]) and human isolates from a 2004 outbreak of H7N3 avian influenza in Canada showed a novel insertion in the HA0 cleavage site of the human and HPAI isolate. This insertion likely occurred by recombination between the hemagglutination and matrix genes in the LPAI virus.

Amino Acid Sequence↗

The Genome sequence of the SARS-associated coronavirus.

We sequenced the 29,751-base genome of the severe acute respiratory syndrome (SARS)-associated coronavirus known as the Tor2 isolate. The genome sequence reveals that this coronavirus is only moderately related to other known coronaviruses, including two human coronaviruses, HCoV-OC43 and HCoV-229E. Phylogenetic analysis of the predicted viral proteins indicates that the virus does not closely resemble any of the three previously known groups of coronaviruses. The genome sequence will aid in the diagnosis of SARS virus infection in humans and potential animal hosts (using polymerase chain reaction and immunological tests), in the development of antivirals (including neutralizing antibodies), and in the identification of putative epitopes for vaccine development.

3' Untranslated Regions↗

Generation and initial analysis of more than 15,000 full-length human and mouse cDNA sequences.

The National Institutes of Health Mammalian Gene Collection (MGC) Program is a multiinstitutional effort to identify and sequence a cDNA clone containing a complete ORF for each human and mouse gene. ESTs were generated from libraries enriched for full-length cDNAs and analyzed to identify candidate full-ORF clones, which then were sequenced to high accuracy. The MGC has currently sequenced and verified the full ORF for a nonredundant set of >9,000 human and >6,000 mouse genes. Candidate full-ORF clones for an additional 7,800 human and 3,500 mouse genes also have been identified. All MGC sequences and clones are available without restriction through public databases and clone distribution networks (see http:mgc.nci.nih.gov).

Algorithms↗

An efficient strategy for large-scale high-throughput transposon-mediated sequencing of cDNA clones.

We describe an efficient high-throughput method for accurate DNA sequencing of entire cDNA clones. Developed as part of our involvement in the Mammalian Gene Collection full-length cDNA sequencing initiative, the method has been used and refined in our laboratory since September 2000. Amenable to large scale projects, we have used the method to generate >7 Mb of accurate sequence from 3695 candidate full-length cDNAs. Sequencing is accomplished through the insertion of Mu transposon into cDNAs, followed by sequencing reactions primed with Mu-specific sequencing primers. Transposon insertion reactions are not performed with individual cDNAs but rather on pools of up to 96 clones. This pooling strategy reduces the number of transposon insertion sequencing libraries that would otherwise be required, reducing the costs and enhancing the efficiency of the transposon library construction procedure. Sequences generated using transposon-specific sequencing primers are assembled to yield the full-length cDNA sequence, with sequence editing and other sequence finishing activities performed as required to resolve sequence ambiguities. Although analysis of the many thousands (22 785) of sequenced Mu transposon insertion events revealed a weak sequence preference for Mu insertion, we observed insertion of the Mu transposon into 1015 of the possible 1024 5mer candidate insertion sites.

Bacteriophage mu↗