PubMed Health⌕ Search

Biomedical subjects

Steven J M Jones

Publications and source records attributed to Steven J M Jones.

At least 19 recordsLinked to original sources

Evaluating Language Models for Biomedical Fact-Checking: A Benchmark Dataset for Cancer Variant Interpretation Verification.

Accurate interpretation of genomic variants is critical for precision oncology but remains slow and dependent on specialized expertise. Public knowledgebases such as the Clinical Interpretation of Variants in Cancer (CIViC) help by curating literature-backed variant interpretations in a structured form, yet verification and review have become major bottlenecks. To address this, we developed CIViC-Fact, a benchmark dataset and pipeline for testing automated systems that verify the accuracy of cancer variant claims. CIViC-Fact links structured claims to sentence-level supporting or refuting evidence from full-text articles, and includes expert annotations and explanations. We evaluated multiple language models. Proprietary models performed well without training, but a smaller open-source model, fine-tuned on CIViC-Fact, achieved the highest accuracy (89%). Applying our fact-checking pipeline to real CIViC entries showed that reviewing less than 20% of content, focusing on flagged entries, would be sufficient to catch over half of all errors. This AI-assisted triage greatly accelerates the review process without replacing or reducing expert insight, ensuring that existing careful oversight remains in place while curators can work more efficiently. CIViC-Fact provides a realistic, high-consequence framework for biomedical fact-checking and a path toward more rigorous and efficient knowledgebase curation.

Journal Article↗

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans↗

Assembling millions of short DNA sequences using SSAKE.

UNLABELLED: Novel DNA sequencing technologies with the potential for up to three orders magnitude more sequence throughput than conventional Sanger sequencing are emerging. The instrument now available from Solexa Ltd, produces millions of short DNA sequences of 25 nt each. Due to ubiquitous repeats in large genomes and the inability of short sequences to uniquely and unambiguously characterize them, the short read length limits applicability for de novo sequencing. However, given the sequencing depth and the throughput of this instrument, stringent assembly of highly identical sequences can be achieved. We describe SSAKE, a tool for aggressively assembling millions of short nucleotide sequences by progressively searching through a prefix tree for the longest possible overlap between any two sequences. SSAKE is designed to help leverage the information from short sequence reads by stringently assembling them into contiguous sequences that can be used to characterize novel sequencing targets. AVAILABILITY: http://www.bcgsc.ca/bioinfo/software/ssake.

Algorithms↗

Meta-analysis and meta-review of thyroid cancer gene expression profiling studies identifies important diagnostic biomarkers.

PURPOSE: An estimated 4% to 7% of the population will develop a clinically significant thyroid nodule during their lifetime. In many cases, preoperative diagnoses by needle biopsy are inconclusive. Thus, there is a clear need for improved diagnostic tests to distinguish malignant from benign thyroid tumors. The recent development of high-throughput molecular analytic techniques should allow the rapid evaluation of new diagnostic markers. However, researchers are faced with an overwhelming number of potential markers from numerous thyroid cancer expression profiling studies. MATERIALS AND METHODS: To address this challenge, we have carried out a comprehensive meta-review of thyroid cancer biomarkers from 21 published studies. A gene ranking system that considers the number of comparisons in agreement, total number of samples, average fold-change and direction of change was devised. RESULTS: We have observed that genes are consistently reported by multiple studies at a highly significant rate (P < .05). Comparison with a meta-analysis of studies reprocessed from raw data showed strong concordance with our method. CONCLUSION: Our approach represents a useful method for identifying consistent gene expression markers when raw data are unavailable. A review of the top 12 candidates revealed well known thyroid cancer markers such as MET, TFF3, SERPINA1, TIMP1, FN1, and TPO as well as relatively novel or uncharacterized genes such as TGFA, QPCT, CRABP1, FCGBP, EPS8 and PROS1. These candidates should help to develop a panel of markers with sufficient sensitivity and specificity for the diagnosis of thyroid tumors in a clinical setting.

Biomarkers, Tumor↗

The ELT-2 GATA-factor and the global regulation of transcription in the C. elegans intestine.

A SAGE library was prepared from hand-dissected intestines from adult Caenorhabditis elegans, allowing the identification of >4000 intestinally-expressed genes; this gene inventory provides fundamental information for understanding intestine function, structure and development. Intestinally-expressed genes fall into two broad classes: widely-expressed "housekeeping" genes and genes that are either intestine-specific or significantly intestine-enriched. Within this latter class of genes, we identified a subset of highly-expressed highly-validated genes that are expressed either exclusively or primarily in the intestine. Over half of the encoded proteins are candidates for secretion into the intestinal lumen to hydrolyze the bacterial food (e.g. lysozymes, amoebapores, lipases and especially proteases). The promoters of this subset of intestine-specific/intestine-enriched genes were analyzed computationally, using both a word-counting method (RSAT oligo-analysis) and a method based on Gibbs sampling (MotifSampler). Both methods returned the same over-represented site, namely an extended GATA-related sequence of the general form AHTGATAARR, which agrees with experimentally determined cis-acting control sequences found in intestine genes over the past 20 years. All promoters in the subset contain such a site, compared to <5% for control promoters; moreover, our analysis suggests that the majority (perhaps all) of genes expressed exclusively or primarily in the worm intestine are likely to contain such a site in their promoters. There are three zinc-finger GATA-type factors that are candidates to bind this extended GATA site in the differentiating C. elegans intestine: ELT-2, ELT-4 and ELT-7. All evidence points to ELT-2 being the most important of the three. We show that worms in which both the elt-4 and the elt-7 genes have been deleted from the genome are essentially wildtype, demonstrating that ELT-2 provides all essential GATA-factor functions in the intestine. The SAGE analysis also identifies more than a hundred other transcription factors in the adult intestine but few show an RNAi-induced loss-of-function phenotype and none (other than ELT-2) show a phenotype primarily in the intestine. We thus propose a simple model in which the ELT-2 GATA factor directly participates in the transcription of all intestine-specific/intestine-enriched genes, from the early embryo through to the dying adult. Other intestinal transcription factors would thus modulate the action of ELT-2, depending on the worm's nutritional and physiological needs.

Animals↗

Locating mammalian transcription factor binding sites: a survey of computational and experimental techniques.

Fields such as genomics and systems biology are built on the synergism between computational and experimental techniques. This type of synergism is especially important in accomplishing goals like identifying all functional transcription factor binding sites in vertebrate genomes. Precise detection of these elements is a prerequisite to deciphering the complex regulatory networks that direct tissue specific and lineage specific patterns of gene expression. This review summarizes approaches for in silico, in vitro, and in vivo identification of transcription factor binding sites. A variety of techniques useful for localized- and high-throughput analyses are discussed here, with emphasis on aspects of data generation and verification.

Animals↗

THOR: targeted high-throughput ortholog reconstructor.

Low-coverage genomes (LCGs) are becoming an increasingly important source of data for phylogenetic studies. However, assembly of these genomes is time consuming, difficult and lags behind sequence generation. THOR is a fast, stringent application for targeted reconstruction of sequence orthologs in unassembled LCGs. Using a 4x coverage set of mouse whole-genome sequence reads, THOR could partially or completely reconstruct 416/1000 human promoter ortholog regions in approximately 7.3 min/promoter. THOR's reconstruction rate improves markedly with both higher-coverage, and less divergent target species.

Algorithms↗

The complete genome of Rhodococcus sp. RHA1 provides insights into a catabolic powerhouse.

Rhodococcus sp. RHA1 (RHA1) is a potent polychlorinated biphenyl-degrading soil actinomycete that catabolizes a wide range of compounds and represents a genus of considerable industrial interest. RHA1 has one of the largest bacterial genomes sequenced to date, comprising 9,702,737 bp (67% G+C) arranged in a linear chromosome and three linear plasmids. A targeted insertion methodology was developed to determine the telomeric sequences. RHA1's 9,145 predicted protein-encoding genes are exceptionally rich in oxygenases (203) and ligases (192). Many of the oxygenases occur in the numerous pathways predicted to degrade aromatic compounds (30) or steroids (4). RHA1 also contains 24 nonribosomal peptide synthase genes, six of which exceed 25 kbp, and seven polyketide synthase genes, providing evidence that rhodococci harbor an extensive secondary metabolism. Among sequenced genomes, RHA1 is most similar to those of nocardial and mycobacterial strains. The genome contains few recent gene duplications. Moreover, three different analyses indicate that RHA1 has acquired fewer genes by recent horizontal transfer than most bacteria characterized to date and far fewer than Burkholderia xenovorans LB400, whose genome size and catabolic versatility rival those of RHA1. RHA1 and LB400 thus appear to demonstrate that ecologically similar bacteria can evolve large genomes by different means. Overall, RHA1 appears to have evolved to simultaneously catabolize a diverse range of plant-derived compounds in an O(2)-rich environment. In addition to establishing RHA1 as an important model for studying actinomycete physiology, this study provides critical insights that facilitate the exploitation of these industrially important microorganisms.

Bacterial Proteins↗

Analysis of the prostate cancer cell line LNCaP transcriptome using a sequencing-by-synthesis approach.

BACKGROUND: High throughput sequencing-by-synthesis is an emerging technology that allows the rapid production of millions of bases of data. Although the sequence reads are short, they can readily be used for re-sequencing. By re-sequencing the mRNA products of a cell, one may rapidly discover polymorphisms and splice variants particular to that cell. RESULTS: We present the utility of massively parallel sequencing by synthesis for profiling the transcriptome of a human prostate cancer cell-line, LNCaP, that has been treated with the synthetic androgen, R1881. Through the generation of approximately 20 megabases (MB) of EST data, we detect transcription from over 10,000 gene loci, 25 previously undescribed alternative splicing events involving known exons, and over 1,500 high quality single nucleotide discrepancies with the reference human sequence. Further, we map nearly 10,000 ESTs to positions on the genome where no transcription is currently predicted to occur. We also characterize various obstacles with using sequencing by synthesis for transcriptome analysis and propose solutions to these problems. CONCLUSION: The use of high-throughput sequencing-by-synthesis methods for transcript profiling allows the specific and sensitive detection of many of a cell's transcripts, and also allows the discovery of high quality base discrepancies, and alternative splice variants. Thus, this technology may provide an effective means of understanding various disease states, discovering novel targets for disease treatment, and discovery of novel transcripts.

Adenocarcinoma↗

Oligonucleotide microarray analysis of genomic imbalance in children with mental retardation.

The cause of mental retardation in one-third to one-half of all affected individuals is unknown. Microscopically detectable chromosomal abnormalities are the most frequently recognized cause, but gain or loss of chromosomal segments that are too small to be seen by conventional cytogenetic analysis has been found to be another important cause. Array-based methods offer a practical means of performing a high-resolution survey of the entire genome for submicroscopic copy-number variants. We studied 100 children with idiopathic mental retardation and normal results of standard chromosomal analysis, by use of whole-genome sampling analysis with Affymetrix GeneChip Human Mapping 100K arrays. We found de novo deletions as small as 178 kb in eight cases, de novo duplications as small as 1.1 Mb in two cases, and unsuspected mosaic trisomy 9 in another case. This technology can detect at least twice as many potentially pathogenic de novo copy-number variants as conventional cytogenetic analysis can in people with mental retardation.

Child↗

Sequence biases in large scale gene expression profiling data.

We present the results of a simple, statistical assay that measures the G+C content sensitivity bias of gene expression experiments without the requirement of a duplicate experiment. We analyse five gene expression profiling methods: Affymetrix GeneChip, Long Serial Analysis of Gene Expression (LongSAGE), LongSAGELite, 'Classic' Massively Parallel Signature Sequencing (MPSS) and 'Signature' MPSS. We demonstrate the methods have systematic and random errors leading to a different G+C content sensitivity. The relationship between this experimental error and the G+C content of the probe set or tag that identifies each gene influences whether the gene is detected and, if detected, the level of gene expression measured. LongSAGE has the least bias, while Signature MPSS shows a strong bias to G+C rich tags and Affymetrix data show different bias depending on the data processing method (MAS 5.0, RMA or GC-RMA). The bias in the Affymetrix data primarily impacts genes expressed at lower levels. Despite the larger sampling of the MPSS library, SAGE identifies significantly more genes (60% more RefSeq genes in a single comparison).

Animals↗

Sequencing and analysis of 10,967 full-length cDNA clones from Xenopus laevis and Xenopus tropicalis reveals post-tetraploidization transcriptome remodeling.

Sequencing of full-insert clones from full-length cDNA libraries from both Xenopus laevis and Xenopus tropicalis has been ongoing as part of the Xenopus Gene Collection Initiative. Here we present 10,967 full ORF verified cDNA clones (8049 from X. laevis and 2918 from X. tropicalis) as a community resource. Because the genome of X. laevis, but not X. tropicalis, has undergone allotetraploidization, comparison of coding sequences from these two clawed (pipid) frogs provides a unique angle for exploring the molecular evolution of duplicate genes. Within our clone set, we have identified 445 gene trios, each comprised of an allotetraploidization-derived X. laevis gene pair and their shared X. tropicalis ortholog. Pairwise dN/dS, comparisons within trios show strong evidence for purifying selection acting on all three members. However, dN/dS ratios between X. laevis gene pairs are elevated relative to their X. tropicalis ortholog. This difference is highly significant and indicates an overall relaxation of selective pressures on duplicated gene pairs. We have found that the paralogs that have been lost since the tetraploidization event are enriched for several molecular functions, but have found no such enrichment in the extant paralogs. Approximately 14% of the paralogous pairs analyzed here also show differential expression indicative of subfunctionalization.

Animals↗

An interactive tool for visualization of relationships between gene expression profiles.

BACKGROUND: Application of phenetic methods to gene expression analysis proved to be a successful approach. Visualizing the results in a 3-dimentional space may further enhance these techniques. RESULTS: We designed and built TreeBuilder3D, an interactive viewer for visualizing the hierarchical relationships between expression profiles such as SAGE libraries or microarrays. The program allows loading expression data as plain text files and visualizing the relative differences of the analyzed datasets in 3-dimensional space using various distance metrics. CONCLUSION: TreeBuilder3D provides a simple interface and has a small size. Written in Java, TreeBuilder3D is a platform-independent, open source application, which may be useful in analysis of large-scale gene expression data.

Algorithms↗

Physical map-assisted whole-genome shotgun sequence assemblies.

We describe a targeted approach to improve the contiguity of whole-genome shotgun sequence (WGS) assemblies at run-time, using information from Bacterial Artificial Chromosome (BAC)-based physical maps. Clone sizes and overlaps derived from clone fingerprints are used for the calculation of length constraints between any two BAC neighbors sharing 40% of their size. These constraints are used to promote the linkage and guide the arrangement of sequence contigs within a sequence scaffold at the layout phase of WGS assemblies. This process is facilitated by FASSI, a stand-alone application that calculates BAC end and BAC overlap length constraints from clone fingerprint map contigs created by the FPC package. FASSI is designed to work with the assembly tool PCAP, but its output can be formatted to work with other WGS assembly algorithms able to use length constraints for individual clones. The FASSI method is simple to implement, potentially cost-effective, and has resulted in the increase of scaffold contiguity for both the Drosophila melanogaster and Cryptococcus gattii genomes when compared to a control assembly without map-derived constraints. A 6.5-fold coverage draft DNA sequence of the Pan troglodytes (chimpanzee) genome was assembled using map-derived constraints and resulted in a 26.1% increase in scaffold contiguity.

Animals↗

Genomics of hybrid poplar (Populus trichocarpax deltoides) interacting with forest tent caterpillars (Malacosoma disstria): normalized and full-length cDNA libraries, expressed sequence tags, and a cDNA microarray for the study of insect-induced defences in poplar.

As part of a genomics strategy to characterize inducible defences against insect herbivory in poplar, we developed a comprehensive suite of functional genomics resources including cDNA libraries, expressed sequence tags (ESTs) and a cDNA microarray platform. These resources are designed to complement the existing poplar genome sequence and poplar (Populus spp.) ESTs by focusing on herbivore- and elicitor-treated tissues and incorporating normalization methods to capture rare transcripts. From a set of 15 standard, normalized or full-length cDNA libraries, we generated 139,007 3'- or 5'-end sequenced ESTs, representing more than one-third of the c. 385,000 publicly available Populus ESTs. Clustering and assembly of 107,519 3'-end ESTs resulted in 14,451 contigs and 20,560 singletons, altogether representing 35,011 putative unique transcripts, or potentially more than three-quarters of the predicted c. 45,000 genes in the poplar genome. Using this EST resource, we developed a cDNA microarray containing 15,496 unique genes, which was utilized to monitor gene expression in poplar leaves in response to herbivory by forest tent caterpillars (Malacosoma disstria). After 24 h of feeding, 1191 genes were classified as up-regulated, compared to only 537 down-regulated. Functional classification of this induced gene set revealed genes with roles in plant defence (e.g. endochitinases, Kunitz protease inhibitors), octadecanoid and ethylene signalling (e.g. lipoxygenase, allene oxide synthase, 1-aminocyclopropane-1-carboxylate oxidase), transport (e.g. ABC proteins, calreticulin), secondary metabolism [e.g. polyphenol oxidase, isoflavone reductase, (-)-germacrene D synthase] and transcriptional regulation [e.g. leucine-rich repeat transmembrane kinase, several transcription factor classes (zinc finger C3H type, AP2/EREBP, WRKY, bHLH)]. This study provides the first genome-scale approach to characterize insect-induced defences in a woody perennial providing a solid platform for functional investigation of plant-insect interactions in poplar.

Animals↗

Conifer defence against insects: microarray gene expression profiling of Sitka spruce (Picea sitchensis) induced by mechanical wounding or feeding by spruce budworms (Choristoneura occidentalis) or white pine weevils (Pissodes strobi) reveals large-scale changes of the host transcriptome.

Conifers are resistant to attack from a large number of potential herbivores or pathogens. Previous molecular and biochemical characterization of selected conifer defence systems support a model of multigenic, constitutive and induced defences that act on invading insects via physical, chemical, biochemical or ecological (multitrophic) mechanisms. However, the genomic foundation of the complex defence and resistance mechanisms of conifers is largely unknown. As part of a genomics strategy to characterize inducible defences and possible resistance mechanisms of conifers against insect herbivory, we developed a cDNA microarray building upon a new spruce (Picea spp.) expressed sequence tag resource. This first-generation spruce cDNA microarray contains 9720 cDNA elements representing c. 5500 unique genes. We used this array to monitor gene expression in Sitka spruce (Picea sitchensis) bark in response to herbivory by white pine weevils (Pissodes strobi, Curculionidae) or wounding, and in young shoot tips in response to western spruce budworm (Choristoneura occidentalis, Lepidopterae) feeding. Weevils are stem-boring insects that feed on phloem, while budworms are foliage feeding larvae that consume needles and young shoot tips. Both insect species and wounding treatment caused substantial changes of the host plant transcriptome detected in each case by differential gene expression of several thousand array elements at 1 or 2 d after the onset of treatment. Overall, there was considerable overlap among differentially expressed gene sets from these three stress treatments. Functional classification of the induced transcripts revealed genes with roles in general plant defence, octadecanoid and ethylene signalling, transport, secondary metabolism, and transcriptional regulation. Several genes involved in primary metabolic processes such as photosynthesis were down-regulated upon insect feeding or wounding, fitting with the concept of dynamic resource allocation in plant defence. Refined expression analysis using gene-specific primers and real-time PCR for selected transcripts was in agreement with microarray results for most genes tested. This study provides the first large-scale survey of insect-induced defence transcripts in a gymnosperm and provides a platform for functional investigation of plant-insect interactions in spruce. Induction of spruce genes of octadecanoid and ethylene signalling, terpenoid biosynthesis, and phenolic secondary metabolism are discussed in more detail.

Animals↗

Selective targeting of indel-inferred differences in spatial structures of homologous proteins.

The eukaryotic pathogen Leishmania donovani possesses a housekeeping protein Elongation-Factor-1alpha (EF-1alpha) which has been found to be unexpectedly involved in the pathogen's virulence. Because it is associated with virulence and essential for cell survival, this protein is an attractive choice for drug targeting; however, its sequence is highly similar (> 80% sequence identity) to that of its human homolog, rendering it a risky choice for a drug target. The chief difference between these two proteins has been found to be a 12 amino acid sequence present in human EF-1alpha but absent from leishmania EF-1alpha. Furthermore, it has been shown that this 12 amino acid insert in the human sequence corresponds to a hairpin loop on the surface of the protein. In this study, we searched for those spatial features in leishmania EF-1alpha that are impacted or obscured by the extra hairpin loop in the human counterpart. We have also conducted a large-scale in silico screening for small molecules that could plausibly bind to these protein features. While experimental evidence is required to verify our results, our findings thus far appear to support this approach as a new strategy for the development of antagonists against pathogenic targets having close human homologs.

Amino Acid Sequence↗

Prediction of genomic functional elements.

As the number of sequenced genomes increases, the ability to deduce genome function becomes increasingly salient. For many genome sequences, the only annotation that will be available for the foreseeable future will be based on computational predictions and comparisons with functional elements in related species. Here we discuss computational approaches for automated genome-wide annotation of functional elements in mammalian genomes. These include methods for ab initio and comparative gene-structure predictions. Gene features such as intron splice sites, 3' untranslated regions, promoters, and cis-regulatory elements are discussed, as is a novel method for predicting DNaseI hypersensitive sites. Recent methodologies for predicting noncoding RNA genes, including microRNA genes and their targets, are also reviewed.

Animals↗