PubMed Health⌕ Search

Biomedical subjects

Yoseph Barash

Publications and source records attributed to Yoseph Barash.

2 recordsLinked to original sources

MAJIQ-CLIN: A novel tool to help identify Mendelian disease-causing variants from RNA-seq data.

PURPOSE: The current diagnostic rate for patients with suspected Mendelian genetic disorders is low, despite exome/genome sequencing being the standard of care. One reason for this low diagnostic rate is that traditional exome/genome sequencing analysis methods struggle to detect RNA splicing aberrations. Causative variants often involve splicing changes, with numerous splice-altering variants being responsible for known Mendelian disorders. Therefore, it is crucial to develop reliable tools to detect, quantify, prioritize, and visualize RNA splicing aberrations from patient RNA sequencing data. METHODS: We developed Modeling Alternative Junction Inclusion Quantification for Clinical Applications (MAJIQ-CLIN), a method to identify RNA splicing aberrations in patients' RNA sequencing data compared with a cohort of control samples. MAJIQ-CLIN can efficiently process large datasets, avoiding reprocessing when new data are added, while effectively detecting local splicing variations with deviations in a given patient, termed outlier local splicing variation, or unique to the patient, termed private local splicing variation. RESULTS: We performed a systematic evaluation of the accuracy of tools for detecting patients' RNA splicing aberrations from RNA sequence using synthetic data across several aberration types and transcript inclusion levels. Then, we used several real datasets to assess MAJIQ-CLINs ability to identify solved test cases and control for the effect of confounders such as batches. We showed that MAJIQ-CLIN compares favorably to existing tools in both accuracy and efficiency. We also used MAJIQ-CLIN to investigate several unsolved patient cases from the Undiagnosed Diseases Network. CONCLUSION: MAJIQ-CLIN offers an efficient, accurate, and user-friendly tool to aid in diagnosing Mendelian disease-causing variants from RNA sequence data.

Bioinformatics↗

Context-specific Bayesian clustering for gene expression data.

The recent growth in genomic data and measurements of genome-wide expression patterns allows us to apply computational tools to examine gene regulation by transcription factors. In this work, we present a class of mathematical models that help in understanding the connections between transcription factors and functional classes of genes based on genetic and genomic data. Such a model represents the joint distribution of transcription factor binding sites and of expression levels of a gene in a unified probabilistic model. Learning a combined probability model of binding sites and expression patterns enables us to improve the clustering of the genes based on the discovery of putative binding sites and to detect which binding sites and experiments best characterize a cluster. To learn such models from data, we introduce a new search method that rapidly learns a model according to a Bayesian score. We evaluate our method on synthetic data as well as on real life data and analyze the biological insights it provides. Finally, we demonstrate the applicability of the method to other data analysis problems in gene expression data.

Bayes Theorem↗