PubMed HealthSearch

Biomedical subjects

Gregory R Grant

Publications and source records attributed to Gregory R Grant.

2 recordsLinked to original sources

MAJIQ-CLIN: A novel tool to help identify Mendelian disease-causing variants from RNA-seq data.

PURPOSE: The current diagnostic rate for patients with suspected Mendelian genetic disorders is low, despite exome/genome sequencing being the standard of care. One reason for this low diagnostic rate is that traditional exome/genome sequencing analysis methods struggle to detect RNA splicing aberrations. Causative variants often involve splicing changes, with numerous splice-altering variants being responsible for known Mendelian disorders. Therefore, it is crucial to develop reliable tools to detect, quantify, prioritize, and visualize RNA splicing aberrations from patient RNA sequencing data. METHODS: We developed Modeling Alternative Junction Inclusion Quantification for Clinical Applications (MAJIQ-CLIN), a method to identify RNA splicing aberrations in patients' RNA sequencing data compared with a cohort of control samples. MAJIQ-CLIN can efficiently process large datasets, avoiding reprocessing when new data are added, while effectively detecting local splicing variations with deviations in a given patient, termed outlier local splicing variation, or unique to the patient, termed private local splicing variation. RESULTS: We performed a systematic evaluation of the accuracy of tools for detecting patients' RNA splicing aberrations from RNA sequence using synthetic data across several aberration types and transcript inclusion levels. Then, we used several real datasets to assess MAJIQ-CLINs ability to identify solved test cases and control for the effect of confounders such as batches. We showed that MAJIQ-CLIN compares favorably to existing tools in both accuracy and efficiency. We also used MAJIQ-CLIN to investigate several unsolved patient cases from the Undiagnosed Diseases Network. CONCLUSION: MAJIQ-CLIN offers an efficient, accurate, and user-friendly tool to aid in diagnosing Mendelian disease-causing variants from RNA sequence data.

Bioinformatics

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation