PubMed Health⌕ Search

Biomedical subjects

Robert A Holt

Publications and source records attributed to Robert A Holt.

At least 19 recordsLinked to original sources

A de novo algorithm for allele reconstruction from Oxford nanopore amplicon reads, with application to CYP2D6.

MOTIVATION: The Oxford Nanopore Technologies' sequencing platform offers a path towards bedside genomics, producing long reads that can completely cover a gene of interest, and detect any known or novel variant the gene contains. However, the analysis of these long reads to identify actionable genotypes remains challenging and typically requires customization depending on the target gene. RESULTS: Here, we describe a generic algorithm to accurately reconstruct allele sequences derived from long-reads of amplicon-based data. Rather than calling variants directly from these long-reads, our method takes a "sequence-first" approach, performing an unbiased reconstruction of the underlying amplicon sequences to generate high-confidence reconstructed allele sequences. This is done without user input of the target gene, allowing for any source amplicon to be reconstructed. These high-confidence reconstructed allele sequences are then compared to the genomic reference sequence of the gene to infer the specific diplotype present in the sample. This approach is agnostic towards the number of genes and alleles present and readily detects novel variants. We demonstrate our approach using three independent data sets for CYP2D6, a diverse and complex gene with over 175 known alleles of clinical significance. We show how our approach can accurately recover validated CYP2D6 diplotypes from 20 Coriell samples covering 14 distinct alleles, using different amplicons, flow cell versions, and depths. This includes inferring occurrences of allele duplication events from relative abundances of each allele, a critical factor for ascribing functional effects to a diplotype. Further, we demonstrate our approach's utility for other genomic regions, including HLA. AVAILABILITY: Custom code is available at the following GitHub repository, along with instructions for use and test data: https://github.com/scottdbrown/allele-reconstruction-long-read-amplicon-data. A snapshot of the code at the time of publication is available on Zenodo.org; doi 10.5281/zenodo.19716004. Raw .fastq sequence data for our three sequencing runs is available at the SRA under Bioproject PRJNA1357883 (https://www.ncbi.nlm.nih.gov/bioproject/1357883).

Alleles↗

GPNMB-directed CAR T cell therapy against MiT/TFE-family fusion-driven solid tumors.

Chimeric antigen receptor (CAR) T cell therapy for solid tumors is constrained by the scarcity of safe, uniformly expressed cell-surface targets. Here we identify glycoprotein NMB (GPNMB)-an MiT/TFE-family fusion-driven protein-as being highly, homogeneously and stably expressed in primary and relapsed alveolar soft-part sarcoma (ASPS) and translocation renal cell carcinoma. We develop a GPNMB-directed CAR T cell product, GCAR1, which demonstrates potent activity against patient-matched cells, organoids and xenograft models. Post hoc interim analysis of a first-in-human open-label, individual-participant trial ( NCT07104682 ) for a participant with relapsed/refractory, metastatic ASPS showed that GCAR1 induces stable disease for up to 3 months, accompanied by resolution of many nontarget lesions (primary endpoint), and is well tolerated. GCAR1 T cells expand in peripheral blood as a polyclonal population and remain detectable for 1 month. Spatial transcriptomics identified immunosuppressive niches in a treatment-resistant lesion and immune checkpoint blockade synergized with GCAR1 in a xenograft model. Altogether, our data provide a proof of concept for treating GPNMB-expressing solid tumors with GCAR1 and more broadly targeting surface antigens driven by oncogenic gene fusions with CAR T cell therapies.

Animals↗

Assembling millions of short DNA sequences using SSAKE.

UNLABELLED: Novel DNA sequencing technologies with the potential for up to three orders magnitude more sequence throughput than conventional Sanger sequencing are emerging. The instrument now available from Solexa Ltd, produces millions of short DNA sequences of 25 nt each. Due to ubiquitous repeats in large genomes and the inability of short sequences to uniquely and unambiguously characterize them, the short read length limits applicability for de novo sequencing. However, given the sequencing depth and the throughput of this instrument, stringent assembly of highly identical sequences can be achieved. We describe SSAKE, a tool for aggressively assembling millions of short nucleotide sequences by progressively searching through a prefix tree for the longest possible overlap between any two sequences. SSAKE is designed to help leverage the information from short sequence reads by stringently assembling them into contiguous sequences that can be used to characterize novel sequencing targets. AVAILABILITY: http://www.bcgsc.ca/bioinfo/software/ssake.

Algorithms↗

The ELT-2 GATA-factor and the global regulation of transcription in the C. elegans intestine.

A SAGE library was prepared from hand-dissected intestines from adult Caenorhabditis elegans, allowing the identification of >4000 intestinally-expressed genes; this gene inventory provides fundamental information for understanding intestine function, structure and development. Intestinally-expressed genes fall into two broad classes: widely-expressed "housekeeping" genes and genes that are either intestine-specific or significantly intestine-enriched. Within this latter class of genes, we identified a subset of highly-expressed highly-validated genes that are expressed either exclusively or primarily in the intestine. Over half of the encoded proteins are candidates for secretion into the intestinal lumen to hydrolyze the bacterial food (e.g. lysozymes, amoebapores, lipases and especially proteases). The promoters of this subset of intestine-specific/intestine-enriched genes were analyzed computationally, using both a word-counting method (RSAT oligo-analysis) and a method based on Gibbs sampling (MotifSampler). Both methods returned the same over-represented site, namely an extended GATA-related sequence of the general form AHTGATAARR, which agrees with experimentally determined cis-acting control sequences found in intestine genes over the past 20 years. All promoters in the subset contain such a site, compared to <5% for control promoters; moreover, our analysis suggests that the majority (perhaps all) of genes expressed exclusively or primarily in the worm intestine are likely to contain such a site in their promoters. There are three zinc-finger GATA-type factors that are candidates to bind this extended GATA site in the differentiating C. elegans intestine: ELT-2, ELT-4 and ELT-7. All evidence points to ELT-2 being the most important of the three. We show that worms in which both the elt-4 and the elt-7 genes have been deleted from the genome are essentially wildtype, demonstrating that ELT-2 provides all essential GATA-factor functions in the intestine. The SAGE analysis also identifies more than a hundred other transcription factors in the adult intestine but few show an RNAi-induced loss-of-function phenotype and none (other than ELT-2) show a phenotype primarily in the intestine. We thus propose a simple model in which the ELT-2 GATA factor directly participates in the transcription of all intestine-specific/intestine-enriched genes, from the early embryo through to the dying adult. Other intestinal transcription factors would thus modulate the action of ELT-2, depending on the worm's nutritional and physiological needs.

Animals↗

Oligonucleotide microarray analysis of genomic imbalance in children with mental retardation.

The cause of mental retardation in one-third to one-half of all affected individuals is unknown. Microscopically detectable chromosomal abnormalities are the most frequently recognized cause, but gain or loss of chromosomal segments that are too small to be seen by conventional cytogenetic analysis has been found to be another important cause. Array-based methods offer a practical means of performing a high-resolution survey of the entire genome for submicroscopic copy-number variants. We studied 100 children with idiopathic mental retardation and normal results of standard chromosomal analysis, by use of whole-genome sampling analysis with Affymetrix GeneChip Human Mapping 100K arrays. We found de novo deletions as small as 178 kb in eight cases, de novo duplications as small as 1.1 Mb in two cases, and unsuspected mosaic trisomy 9 in another case. This technology can detect at least twice as many potentially pathogenic de novo copy-number variants as conventional cytogenetic analysis can in people with mental retardation.

Child↗

Duplication and divergence of 2 distinct pancreatic ribonuclease genes in leaf-eating African and Asian colobine monkeys.

Unique among primates, the colobine monkeys have adapted to a predominantly leaf-eating diet by evolving a foregut that utilizes bacterial fermentation to breakdown and absorb nutrients from such a food source. It has been hypothesized that pancreatic ribonuclease (pRNase) has been recruited to perform a role as a digestive enzyme in foregut fermenters, such as artiodactyl ruminants and the colobines. We present molecular analyses of 23 pRNase gene sequences generated from 8 primate taxa, including 2 African and 2 Asian colobine species. The pRNase gene is single copy in all noncolobine primate species assayed but has duplicated more than once in both the African and Asian colobine monkeys. Phylogenetic reconstructions show that the pRNase-coding and noncoding regions are under different evolutionary constraints, with high levels of concerted evolution among gene duplicates occurring predominantly in the noncoding regions. Our data suggest that 2 functionally distinct pRNases have been selected for in the colobine monkeys, with one group adapting to the role of a digestive enzyme by evolving at an increased rate with loss of positive charge, namely arginine residues. Conclusions relating our data to general hypotheses of evolution following gene duplication are discussed.

Africa↗

Sequencing and analysis of 10,967 full-length cDNA clones from Xenopus laevis and Xenopus tropicalis reveals post-tetraploidization transcriptome remodeling.

Sequencing of full-insert clones from full-length cDNA libraries from both Xenopus laevis and Xenopus tropicalis has been ongoing as part of the Xenopus Gene Collection Initiative. Here we present 10,967 full ORF verified cDNA clones (8049 from X. laevis and 2918 from X. tropicalis) as a community resource. Because the genome of X. laevis, but not X. tropicalis, has undergone allotetraploidization, comparison of coding sequences from these two clawed (pipid) frogs provides a unique angle for exploring the molecular evolution of duplicate genes. Within our clone set, we have identified 445 gene trios, each comprised of an allotetraploidization-derived X. laevis gene pair and their shared X. tropicalis ortholog. Pairwise dN/dS, comparisons within trios show strong evidence for purifying selection acting on all three members. However, dN/dS ratios between X. laevis gene pairs are elevated relative to their X. tropicalis ortholog. This difference is highly significant and indicates an overall relaxation of selective pressures on duplicated gene pairs. We have found that the paralogs that have been lost since the tetraploidization event are enriched for several molecular functions, but have found no such enrichment in the extant paralogs. Approximately 14% of the paralogous pairs analyzed here also show differential expression indicative of subfunctionalization.

Animals↗

A high-throughput screen identifying sequence and promiscuity characteristics of the loxP spacer region in Cre-mediated recombination.

BACKGROUND: Cre-loxP recombination refers to the process of site-specific recombination mediated by two loxP sequences and the Cre recombinase protein. Transgenic experiments exploit integrative recombination, where a donor plasmid carrying a loxP site and DNA of interest integrate into a recipient loxP site in a target genome. Unfortunately, integrative recombination is highly inefficient because the insert is flanked by two loxP sites, which themselves become targets for Cre and lead to subsequent excision of the insert. A small number of mutations have been discovered in parts of the loxP sequence, specifically the spacer and inverted repeat segments, that increase the efficiency of integrative recombination. In this study we introduce a high-throughput in vitro assay to rapidly detect novel loxP spacer mutants and describe the sequence characteristics of successful recombinants. RESULTS: We created synthetic loxP oligonucleotides that contained a combination of inverted repeat mutations (the lox66 and lox71 mutations) and mutant spacer sequences, degenerate at 6 of the 8 positions. After in vitro Cre recombination, 3,124 recombinant clones were identified by sequencing. Included in this set were 31 unique, novel, self-recombining sequences. Using network visualization tools, we recognized 12 spacer sets with restricted promiscuity. We observed that increased guanine content at all spacer positions save for position 8 resulted in increased recombination. Interestingly, recombination between identical spacers was not preferred over non-identical spacers. We also identified a set of 16 pairs of loxP spacers that reacted at least twice with another spacer, but not themselves. Further, neither the wild-type P1 phage loxP sequence nor any of the known loxP spacer mutants appeared to be kinetically favoured by Cre recombinase. CONCLUSION: This study approached loxP spacer mutant screening in an unbiased manner, assuming nothing about candidate loxP sites save for the conserved 4 and 5 spacer positions. Candidate sites were free to recombine with any other sequence in the pool of all possible sites. The subset of loxP sites identified here are candidates for in vivo serial recombination as they have already demonstrated limited promiscuity with other loxP spacer and stability in the presence of Cre.

Bacteriophage P1↗

DNA copy-number analysis in bipolar disorder and schizophrenia reveals aberrations in genes involved in glutamate signaling.

Using bacterial artificial chromosome (BAC) array comparative genome hybridization (aCGH) at approximately 1.4 Mbp resolution, we screened post-mortem brain DNA from bipolar disorder cases, schizophrenia cases and control individuals (n=35 each) for DNA copy-number aberrations. DNA copy number is a largely unexplored source of human genetic variation that may contribute risk for complex disease. We report aberrations at four loci which were seen in affected but not control individuals, and which were verified by quantitative real-time PCR. These aberrant loci contained the genes encoding EFNA5, GLUR7, CACNG2 and AKAP5; all brain-expressed proteins with known or postulated roles in neuronal function, and three of which (GLUR7, CACNG2 and AKAP5) are involved in glutamate signaling. A second cohort of psychiatric samples was also tested by quantitative PCR using the primer/probe sets for EFNA5, GLUR7, CACNG2 and AKAP5, and samples with aberrant copy number were found at three of the four loci (GLUR7, CACNG2 and AKAP5). Further scrutiny of these regions may reveal insights into the etiology and genetic risk factors for these complex psychiatric disorders.

A Kinase Anchor Proteins↗

Genomics of hybrid poplar (Populus trichocarpax deltoides) interacting with forest tent caterpillars (Malacosoma disstria): normalized and full-length cDNA libraries, expressed sequence tags, and a cDNA microarray for the study of insect-induced defences in poplar.

As part of a genomics strategy to characterize inducible defences against insect herbivory in poplar, we developed a comprehensive suite of functional genomics resources including cDNA libraries, expressed sequence tags (ESTs) and a cDNA microarray platform. These resources are designed to complement the existing poplar genome sequence and poplar (Populus spp.) ESTs by focusing on herbivore- and elicitor-treated tissues and incorporating normalization methods to capture rare transcripts. From a set of 15 standard, normalized or full-length cDNA libraries, we generated 139,007 3'- or 5'-end sequenced ESTs, representing more than one-third of the c. 385,000 publicly available Populus ESTs. Clustering and assembly of 107,519 3'-end ESTs resulted in 14,451 contigs and 20,560 singletons, altogether representing 35,011 putative unique transcripts, or potentially more than three-quarters of the predicted c. 45,000 genes in the poplar genome. Using this EST resource, we developed a cDNA microarray containing 15,496 unique genes, which was utilized to monitor gene expression in poplar leaves in response to herbivory by forest tent caterpillars (Malacosoma disstria). After 24 h of feeding, 1191 genes were classified as up-regulated, compared to only 537 down-regulated. Functional classification of this induced gene set revealed genes with roles in plant defence (e.g. endochitinases, Kunitz protease inhibitors), octadecanoid and ethylene signalling (e.g. lipoxygenase, allene oxide synthase, 1-aminocyclopropane-1-carboxylate oxidase), transport (e.g. ABC proteins, calreticulin), secondary metabolism [e.g. polyphenol oxidase, isoflavone reductase, (-)-germacrene D synthase] and transcriptional regulation [e.g. leucine-rich repeat transmembrane kinase, several transcription factor classes (zinc finger C3H type, AP2/EREBP, WRKY, bHLH)]. This study provides the first genome-scale approach to characterize insect-induced defences in a woody perennial providing a solid platform for functional investigation of plant-insect interactions in poplar.

Animals↗

Identification by full-coverage array CGH of human DNA copy number increases relative to chimpanzee and gorilla.

Duplication of chromosomal segments and associated genes is thought to be a primary mechanism for generating evolutionary novelty. By comparative genome hybridization using a full-coverage (tiling) human BAC array with 79-kb resolution, we have identified 63 chromosomal segments, ranging in size from 0.65 to 1.3 Mb, that have inferred copy number increases in human relative to chimpanzee. These segments span 192 Ensembl genes, including 82 gene duplicates (41 reciprocal best BLAST matches). Synonymous and nonsynonymous substitution rates across these pairs provide evidence for general conservation of the amino acid sequence, consistent with the maintenance of function of both copies, and one case of putative positive selection for an uncharacterized gene. Surprisingly, the core histone genes H2A, H2B, H3, and H4 have been duplicated in the human lineage since our split with chimpanzee. The observation of increased copy number of a human cluster of core histone genes suggests that altered dosage, even of highly constrained genes, may be an important evolutionary mechanism.

Animals↗

A mouse atlas of gene expression: large-scale digital gene-expression profiles from precisely defined developing C57BL/6J mouse tissues and cells.

We analyzed 8.55 million LongSAGE tags generated from 72 libraries. Each LongSAGE library was prepared from a different mouse tissue. Analysis of the data revealed extensive overlap with existing gene data sets and evidence for the existence of approximately 24,000 previously undescribed genomic loci. The visual cortex, pancreas, mammary gland, preimplantation embryo, and placenta contain the largest number of differentially expressed transcripts, 25% of which are previously undescribed loci.

Alternative Splicing↗

Simple, robust methods for high-throughput nanoliter-scale DNA sequencing.

We have developed high-throughput DNA sequencing methods that generate high quality data from reactions as small as 400 nL, providing an approximate order of magnitude reduction in reagent use relative to standard protocols. Sequencing of clones from plasmid, fosmid, and BAC libraries yielded read lengths (PHRED20 bases) of 765 +/- 172 (n = 10,272), 621 +/- 201 (n = 1824), and 647 +/- 189 (n = 568), respectively. Implementation of these procedures at high-throughput genome centers could have a substantial impact on the amount of data that can be generated per unit cost.

Nanotechnology↗

Satellog: a database for the identification and prioritization of satellite repeats in disease association studies.

BACKGROUND: To date, 35 human diseases, some of which also exhibit anticipation, have been associated with unstable repeats. Anticipation has been reported in a number of diseases in which repeat expansion may have a role in etiology. Despite the growing importance of unstable repeats in disease, currently no resource exists for the prioritization of repeats. Here we present Satellog, a database that catalogs all pure 1-16 repeat unit satellite repeats in the human genome along with supplementary data. Satellog analyzes each pure repeat in UniGene clusters for evidence of repeat polymorphism. RESULTS: A total of 5,546 such repeats were identified, providing the first indication of many novel polymorphic sites in the genome. Overall, polymorphic repeats were over-represented within 3'-UTR sequence relative to 5'-UTR and coding sequence. Interestingly, we observed that repeat polymorphism within coding sequence is restricted to trinucleotide repeats whereas UTR sequence tolerated a wider range of repeat period polymorphisms. For each pure repeat we also calculate its repeat length percentile rank, its location either within or adjacent to EnsEMBL genes, and its expression profile in normal tissues according to the GeneNote database. CONCLUSION: Satellog provides the ability to dynamically prioritize repeats based on any of their characteristics (i.e. repeat unit, class, period, length, repeat length percentile rank, genomic co-ordinates), polymorphism profile within UniGene, proximity to or presence within gene regions (i.e. cds, UTR, 15 kb upstream etc.), metadata of the genes they are detected within and gene expression profiles within normal human tissues. Unstable repeats associated with 31 diseases were analyzed in Satellog to evaluate their common repeat properties. The utility of Satellog was highlighted by prioritizing repeats for Huntington's disease and schizophrenia. Satellog is available online at http://satellog.bcgsc.ca.

3' Untranslated Regions↗

Analysis of long-lived C. elegans daf-2 mutants using serial analysis of gene expression.

We have identified longevity-associated genes in a long-lived Caenorhabditis elegans daf-2 (insulin/IGF receptor) mutant using serial analysis of gene expression (SAGE), a method that efficiently quantifies large numbers of mRNA transcripts by sequencing short tags. Reduction of daf-2 signaling in these mutant worms leads to a doubling in mean lifespan. We prepared C. elegans SAGE libraries from 1, 6, and 10-d-old adult daf-2 and from 1 and 6-d-old control adults. Differences in gene expression between daf-2 libraries representing different ages and between daf-2 versus control libraries identified not only single genes, but whole gene families that were differentially regulated. These gene families are part of major metabolic pathways including lipid, protein, and energy metabolism, stress response, and cell structure. Similar expression patterns of closely related family members emphasize the importance of these genes in aging-related processes. Global analysis of metabolism-associated genes showed hypometabolic features in mid-life daf-2 mutants that diminish with advanced age. Comparison of our results to recent microarray studies highlights sets of overlapping genes that are highly conserved throughout evolution and thus represent strong candidate genes that control aging and longevity.

Age Factors↗

High-throughput sequencing: a failure mode analysis.

BACKGROUND: Basic manufacturing principles are becoming increasingly important in high-throughput sequencing facilities where there is a constant drive to increase quality, increase efficiency, and decrease operating costs. While high-throughput centres report failure rates typically on the order of 10%, the causes of sporadic sequencing failures are seldom analyzed in detail and have not, in the past, been formally reported. RESULTS: Here we report the results of a failure mode analysis of our production sequencing facility based on detailed evaluation of 9,216 ESTs generated from two cDNA libraries. Two categories of failures are described; process-related failures (failures due to equipment or sample handling) and template-related failures (failures that are revealed by close inspection of electropherograms and are likely due to properties of the template DNA sequence itself). CONCLUSIONS: Preventative action based on a detailed understanding of failure modes is likely to improve the performance of other production sequencing pipelines.

Automation↗

Isolation and characterisation of bacterial strains containing enantioselective DMSO reductase activity: application to the kinetic resolution of racemic sulfoxides.

The kinetic resolution of racemic sulfoxides by dimethyl sulfoxide (DMSO) reductases was investigated with a range of microorganisms. Three bacterial isolates (provisionally identified as Citrobacter braakii, Klebsiella sp. and Serratia sp.) expressing DMSO reductase activity were isolated from environmental samples by anaerobic enrichment with DMSO as terminal electron acceptor. The organisms reduced a diverse range of racemic sulfoxides to yield either residual enantiomer depending upon the strain used. C. braakii DMSO-11 exhibited wide substrate specificity that included dialkyl, diaryl and alkylaryl sulfoxides, and was unique in its ability to reduce the thiosulfinate 1,4-dihydrobenzo-2, 3-dithian-2-oxide. DMSO reductase was purified from the periplasmic fraction of C. braakii DMSO-11 and was used to demonstrate unequivocally that the DMSO reductase was responsible for enantiospecific reductive resolution of racemic sulfoxides.

Citrobacter↗

Genome sequence of the Brown Norway rat yields insights into mammalian evolution.

The laboratory rat (Rattus norvegicus) is an indispensable tool in experimental medicine and drug development, having made inestimable contributions to human health. We report here the genome sequence of the Brown Norway (BN) rat strain. The sequence represents a high-quality 'draft' covering over 90% of the genome. The BN rat sequence is the third complete mammalian genome to be deciphered, and three-way comparisons with the human and mouse genomes resolve details of mammalian evolution. This first comprehensive analysis includes genes and proteins and their relation to human disease, repeated sequences, comparative genome-wide studies of mammalian orthologous chromosomal regions and rearrangement breakpoints, reconstruction of ancestral karyotypes and the events leading to existing species, rates of variation, and lineage-specific and lineage-independent evolutionary events such as expansion of gene families, orthology relations and protein evolution.

Animals↗