PubMed HealthSearch

SEARCH · PubMed Health

Results for “Read depth”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

LYCEUM: learning to call copy number variants on low-coverage ancient genomes.

MOTIVATION: Copy number variants (CNVs) are pivotal in driving phenotypic variation that facilitates species adaptation. They are significant contributors to various disorders, making ancient genomes crucial for uncovering the genetic origins of disease susceptibility across populations. However, detecting CNVs in ancient DNA (aDNA) samples poses substantial challenges due to several factors: (i) aDNA is often highly degraded; (ii) contamination from microbial DNA and DNA from closely related species introduces additional noise into sequencing data; and finally, (iii) the typically low-coverage of aDNA renders accurate CNV detection particularly difficult. Conventional CNV calling algorithms, which are optimized for high-coverage read-depth signals, underperform under such conditions. RESULTS: To address these limitations, we introduce LYCEUM, the first machine learning-based CNV caller for aDNA. To overcome challenges related to data quality and scarcity, we employ a two-step training strategy. First, the model is pre-trained on whole genome sequencing data from the 1000 Genomes Project, teaching it CNV-calling capabilities similar to conventional methods. Next, the model is fine-tuned using high-confidence CNV calls derived from only a few existing high-coverage aDNA samples. During this stage, the model adapts to making CNV calls based on the downsampled read depth signals of the same aDNA samples. LYCEUM achieves accurate detection of CNVs even in typically low-coverage ancient genomes. We also observe that the segmental deletion calls made by LYCEUM show correlation with the demographic history of the samples and exhibit patterns of negative selection inline with natural selection. AVAILABILITY AND IMPLEMENTATION: LYCEUM is available at https://github.com/ciceklab/LYCEUM.

DNA Copy Number Variations

FLASH-TB: an Application of Next-Generation CRISPR to Detect Drug Resistant Tuberculosis from Direct Sputum.

Offering patients with tuberculosis (TB) an optimal and timely treatment regimen depends on the rapid detection of Mycobacterium tuberculosis (Mtb) drug resistance from clinical samples. Finding Low Abundance Sequences by Hybridization (FLASH) is a technique that harnesses the efficiency, specificity, and flexibility of the Cas9 enzyme to enrich targeted sequences. Here, we used FLASH to amplify 52 candidate genes probably associated with resistance to first- and second-line drugs in the Mtb reference strain (H37Rv), then detect drug resistance mutations in cultured Mtb isolates, and in sputum samples. 92% of H37Rv reads mapped to Mtb targets, with 97.8% of target regions covered at a depth ≥ 10X. Among cultured isolates, FLASH-TB detected the same 17 drug resistance mutations as whole genome sequencing (WGS) did, but with much greater depth. Among the 16 sputum samples, FLASH-TB increased recovery of Mtb DNA compared with WGS (from 1.4% [IQR 0.5-7.5] to 33% [IQR 4.6-66.3]) and average depth reads of targets (from 6.3 [IQR 3.8-10.5] to 1991 [IQR 254.4-3623.7]). FLASH-TB identified Mtb complex in all 16 samples based on IS1081 and IS6110 copies. Drug resistance predictions for 15/16 (93.7%) clinical samples were highly concordant with phenotypic DST for isoniazid, rifampicin, amikacin, and kanamycin [15/15 (100%)], ethambutol [12/15 (80%)] and moxifloxacin [14/15 (93.3%)]. These results highlighted the potential of FLASH-TB for detecting Mtb drug resistance from sputum samples.

Humans

Optimizing GRIDSS for clinical use: A targeted NGS filtering strategy for germline structural variant detection.

Detecting intermediate-sized structural variants (SVs) remains challenging in diagnostics, as tools for single-nucleotide and copy-number variants, particularly read-depth-based methods, are often insufficient. GRIDSS addresses this gap by integrating paired-end mapping, split-read analysis, and assembly-based approaches. However, its use in targeted sequencing and diagnostic workflows remains complex. NGS panel data from 9726 patients with suspected hereditary cancer were analyzed using GRIDSS. A filtering strategy was developed to prioritize clinically relevant germline SVs. Multiple parameter settings were tested to optimize performance. The initial dataset of 1,307,592 variants was reduced to 89 candidates after applying the selected filtering strategy. Of these, 24 had been previously detected by routine callers and were not further analyzed. Among the remaining 65, 13 were considered likely true positives after visual inspection using IGV. Experimental validation was performed by Sanger/Nanopore long-read sequencing for these variants, all of which were confirmed. Eight were classified as (likely) pathogenic, including two frameshift duplications in MSH6, one splicing variant in BARD1, and five mobile element insertions in APC, BRCA2, and PALB2. Altogether, GRIDSS implementation increased diagnostic yield while maintaining feasibility for diagnostic workflows. Comprehensive workflow scheme for germline structural variant detection and results in our diagnostic setting.

Humans

COSIGT: population-scalable genotyping of complex loci from low-coverage sequencing data using pangenome graphs.

Pangenome graphs capture extensive structural diversity, but resolving complex loci from shallow sequencing remains challenging, particularly when samples are of low quality such as in ancient DNA. We introduce COSIGT (COsine SImilarity-based GenoTyper), which assigns diploid genotypes by matching read-depth distributions to haplotype paths via cosine similarity. Because this metric evaluates relative coverage profiles rather than absolute read counts, COSIGT substantially outperforms existing likelihood-based tools at low coverage (1-2X). We demonstrate scalability to thousands of modern and ancient genomes, enabling robust, population-scale analyses of complex variation directly from low-coverage datasets.

Humans

An enhanced multisegment RT-PCR method for influenza A virus sequencing: Improved performance and reduced preparation time over traditional methods.

Influenza A viruses (IAVs) remain a major global health threat, affecting both human and animal populations. Whole-genome sequencing is essential for monitoring viral evolution, zoonotic transmission, and emerging variants. However, conventional RT-PCR methods often result in incomplete gene coverage, amplification biases, and reduced sequencing accuracy, particularly in clinical samples. We developed a robust In-house method for IAV full-genome sequencing using the Oxford Nanopore Technologies (ONT) long-read sequencing platform. This method integrates an in-house multisegment Reverse Transcription PCR (RT-PCR) method with a streamlined 2-pool primer design targeting all eight IAV gene segments. RNA extracted from clinical and stock virus samples was reverse-transcribed and amplified using Superscript IV-based chemistry, followed by magnetic bead purification to ensure high-quality amplicons. Sequencing libraries were prepared with the Native Barcoding Kit 24 (SQK-NBD114.24) and sequenced on R10.4.1 flow cells on the MinION MK1C device. Data analysis using the Iterative Refinement Meta-Assembler (IRMA) confirmed improved read depth, uniform coverage, and complete genome recovery. Compared to conventional methods, our In-House Multisegment 2-Pool (IH-MS2P) RT-PCR method generated higher numbers of matched read counts, minimized chimeric artifacts, and delivered superior genome coverage across human, swine, and avian isolates. This optimized RT-PCR method provides a high-performance, time-efficient, and portable solution for influenza genomics, demonstrating robust applicability even with clinical samples of low RNA yield.

Influenza A virus

Anatomical and physiological localization of visual and infrared cell layers in tectum of pit vipers.

Visual and infrared cell layers were identified in the tectum of the pit vipers Crotalus viridis and Sistrurus melitus. Histologic reconstructions of 48 lesions utilizing the Prussian Blue technique were correlated with micrometer depth readings for 251 visual, infrared and bimodal single unit recordings. The visual cell layer extends caudally from approximately the level of the habenula to the rostral border of the posterior corpora quadrigemina. Neurons responding to visual stimulation are generally contained within zones 7b-13, i.e., the superficial 600--700 micrometer of the optic tectum (stratum fibrosum et griseum superficiale and the superficial sublayer of stratum griesum centrale). The infrared cell group is found in layer 7 (a and b; stratum griseum centrale) throughout the optic tectum. Eighty percent of the infrared neurons are found within 500--1,200 micrometer of the surface. In layer 7b the visual and infrared cell groups are mixed; bimodal neurons that respond to a combination of visual and infrared input are located predominantly in this sublamina. The lamination pattern for visual and nonvisual cell groups in the rattlesnake tectum appears to more closely resemble the colubrid tectum and mammalian superior colliculus than the tecta of other reptiles.

Animals

Precision ID mtDNA Whole Genome Panel and sequencing of telogen hairs - perspectives for validation and implementation in casework.

Shed hair is a commonly encountered type of forensic evidence. Shed telogen hairs generally contain insufficient or highly degraded nuclear DNA for STR profiling; however, mtDNA analysis of telogen hair and hair shafts remains possible. We validated whole mitochondrial genome (mtGenome) sequencing using the Precision ID mtDNA Whole Genome Panel (Thermo Fisher Scientific) and subsequently implemented the panel for the analysis of telogen hair, buccal, and casework samples. We analysed 90 diluted DNA samples containing 3-3,600 mtDNA copies, shed telogen hairs and their corresponding mtDNA from buccal swabs from 91 individuals, and 11 archived DNA extracts from hair samples in criminal cases. Complete mtGenome sequences were consistently recovered in 99% of samples across DNA dilution series at DNA input levels as low as 47 mtDNA copies, demonstrating the assay's robustness under low-template conditions. We obtained complete and reproducible mtGenome sequences with ≥ 327 mtDNA copies/µL from telogen hair samples. After applying ISFG recommendations and excluding low-confidence discrepancies associated with high-strand bias, heteroplasmic variants and sequencing artifacts, mtGenome sequence concordance increased from 93.4% to 100%. None of the 16 negative controls produced complete mtDNA sequences. Six negative controls showed low-level mtDNA signal (2-8 variants), consisting predominantly of common polymorphisms. These samples did not yield complete mtGenome sequences and showed no correspondence to any of the analysed samples. Finally, archived telogen hair samples from criminal cases presented complete mtGenome sequences with an average read depth of 1,037x.Our findings highlight the reliability of mtDNA analysis of telogen hairs using the Precision ID mtDNA Whole Genome Panel for implementation in forensic casework.

Forensic casework

Wastewater-based sequencing of respiratory syncytial virus to investigate lineage dynamics and antigenic site mutations: a retrospective genomic epidemiology study.

BACKGROUND: Respiratory syncytial virus (RSV) infections pose a substantial health burden, particularly for clinically vulnerable populations such as infants and older adults. Although novel immunoprophylactic interventions show promise in providing protection, many countries may not have robust surveillance systems to monitor circulating RSV lineages and detect mutations that might reduce the effectiveness of these new interventions. We aimed to assess the diversity and temporal dynamics of circulating RSV lineages in urban populations through amplicon-based sequencing and analysis of wastewater extracts. METHODS: In this prospective observational wastewater-based genomic surveillance study, 32 raw influent 24-h composite samples were collected during the 2022-23 and 2023-24 RSV seasons from both Zurich and Geneva, Switzerland. We applied an RSV subtype-specific amplicon-based sequencing approach to obtain RSV-A and RSV-B sequences from all 64 samples. Mutations relative to reference genomes were identified at positions with read depth above 30. Relative abundances of RSV lineages were estimated from frequencies of lineage-signature mutations, present in greater than 90% of publicly available sequences of that lineage. FINDINGS: Relative abundances of RSV-B (2022-23) and RSV-A (2023-24) lineages were estimated over the two RSV seasons. During the 2022-23 season, the RSV-B B.D.E.1 lineage prevailed in both cities. In the 2023-24 season, multiple RSV-A lineages cocirculated, including A.D.1, A.D.3, A.D.5, and their sub-lineages. Identification and frequency estimation of mutations showed low-frequency, non-synonymous mutations in antigenic sites on the fusion gene of both RSV-A and RSV-B, some of which have not been reported in clinical sequences. The primary outcome was identification and relative abundance of RSV lineages in wastewater samples. INTERPRETATION: These findings show the potential of wastewater-based genomic surveillance to identify and track circulating RSV lineages and clinically relevant mutations. As novel RSV immunoprophylaxis measures are introduced in upcoming RSV seasons, wastewater-derived genomic RSV data provide a valuable baseline for understanding RSV diversity and future viral evolution under increased immunological pressure. FUNDING: This study was funded by the Swiss National Science Foundation and in part by the National Institute Of Allergy And Infectious Diseases of the National Institutes of Health. Funding for sample collection and processing was provided by the Swiss Federal Office of Public Health.

Humans

dsRNAscan maps human dsRNAome, revealing conservation, intermolecular dsRNA, and correlates of ADAR dependency.

The human transcriptome contains millions of A-to-I editing sites arising from an unclear number of poorly characterized dsRNAs. Editing sites reveal the presence of dsRNA, but this method is limited by transcription levels, read depth, and ADAR expression and cannot identify unedited dsRNA. To address these limitations, we developed dsRNAscan. Applying dsRNAscan to the human genome predicted 5 million dsRNAs, mostly in repetitive and intergenic regions. Machine learning models trained on A-to-I editing and RNA structure-probing data identified ∼2.4 million high-confidence predictions, which were enriched at dsRNA-binding protein binding sites. Additionally, we predicted hundreds of dsRNAs conserved across vertebrates and observed thousands of editing-enriched regions suspected to arise from intermolecular dsRNAs formed with sense-antisense transcripts. Quantifying expression of intramolecular and intermolecular dsRNAs accessible to cytoplasmic immune sensors revealed that their ratio correlated with ADAR dependency across cancer cell lines. The human dsRNAome is available as a resource at https://dsrna.chpc.utah.edu/.

A-to-I RNA editing

Single-cell copy number calling and event history reconstruction.

MOTIVATION: Copy number alterations are driving forces of tumour development and the emergence of intra-tumour heterogeneity. A comprehensive picture of these genomic aberrations is therefore essential for the development of personalised and precise cancer diagnostics and therapies. Single-cell sequencing offers the highest resolution for copy number profiling down to the level of individual cells. Recent high-throughput protocols allow for the processing of hundreds of cells through shallow whole-genome DNA sequencing. The resulting low read-depth data poses substantial statistical and computational challenges to the identification of copy number alterations. RESULTS: We developed SCICoNE, a statistical model and MCMC algorithm tailored to single-cell copy number profiling from shallow whole-genome DNA sequencing data. SCICoNE reconstructs the history of copy number events in the tumour and uses these evolutionary relationships to identify the copy number profiles of the individual cells. We show the accuracy of this approach in evaluations on simulated data and demonstrate its practicability in applications to two breast cancer samples from different sequencing protocols. AVAILABILITY AND IMPLEMENTATION: SCICoNE is available at https://github.com/cbg-ethz/SCICoNE.

Single-Cell Analysis

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome‑wide coverage of only 0.1-5×, sWGS data display a pronounced zero‑inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several‑fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy‑number gains (false positives), and true deletions often become indistinguishable from pervasive zero‑coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Assessing the influence of different alignment tools on the accuracy of a forensic epigenetic clock.

MOTIVATION: DNA methylation (DNAm) has long been a commonly investigated biomarker in biomedical research. The current gold standard for DNAm detection is bisulfite sequencing which requires dedicated alignment tools that can handle reduced sequence complexity. One commonly used application of DNAm are epigenetic clock measurements. These clocks have been adapted by many fields for their specific needs, including forensic genetics. Here, epigenetic clocks were designed to help estimate the chronological age of a biological stain donor for investigative purposes. RESULTS: In this study, data generated with a well-established forensic epigenetic clock is aligned with four different bisulfite-specific alignment tools: "Bwa-meth," "Abismal," "Bismark," and "BS-Seeker2." For each tool, we tested up to six different settings, altering parameters such as the maximum number of mismatches or the score function setting. The goal was to investigate whether the final predicted ages differed considerably between the tested alignment tools and settings. Quality controls such as read depth, precision, recall, F1 score, and alignment run time were also assessed. To allow other researchers to easily perform such methylation comparison analyses on their own data, a Shiny app called "MethylAge Explorer" was developed within this study. None of the tested settings for the three alignment tools "Abismal," "Bismark," and "BS-Seeker2" outperformed the originally used alignment tool "Bwa-meth" in terms of age prediction accuracy. However, differences in final age predictions were observed between the different alignment tools. Therefore, it is necessary to be aware of which alignment tool to use for particular epigenetic clocks. AVAILABILITY AND IMPLEMENTATION: The data underlying this article and the code for the shiny app are available on GitHub (https://github.com/charlsut/methylage_explorer).

DNA Methylation

A Chromosome-Level Genome Assembly of the Potato Leafhopper Empoasca fabae (Hemiptera: Cicadellidae).

The potato leafhopper, Empoasca fabae (Harris, 1841), is a highly polyphagous, migratory insect pest of eastern North America that feeds on more than 200 herbaceous and woody plant species, causing substantial losses to forage and field crops. Despite its agricultural and ecological importance, no genome has been available for this species. Here, we present the first chromosome-level genome assembly of E. fabae, generated from Oxford Nanopore long reads, Illumina short reads, and Omni-C proximity-ligation data. The final assembly spans 908 Mb across 132 scaffolds, with 99.8% of the assembly captured in ten chromosome-length scaffolds (nine autosomes and an X chromosome) with a scaffold N50 of 96.2 Mb. The assembly is highly complete, recovering 92.9% of conserved hemipteran single-copy orthologs from protein annotations, and is composed of 47.6% repetitive sequence, dominated by long terminal repeat retrotransposons and unclassified elements. Read-depth comparison between male and female individuals supports assignment of a single sex-linked chromosome, consistent with an XO sex determination system. BRAKER3 gene annotation predicted 31,406 protein-coding genes after retaining the longest isoform per locus. Comparative genome analysis of the two closest related Typhlocybinae species with genomes available, Matsumurasca onukii and Hebata decipiens, revealed extensive chromosome-scale collinearity while defining a shared core gene repertoire. This reference genome provides a foundation for comparative and population genomic studies and for investigating genetic traits in this economically important crop pest species.

Animals

Proportionality-based association metrics in count compositional data.

Compositional data comprise vectors that describe the constituent parts of a whole. Data arising from various -omics platforms such as 16S and RNA sequencing are compositional in nature. In this kind of data, correlations between features on raw counts have no meaningful interpretation. Metrics of proportionality were formulated to address this problem. However, an inherent bias arises when these metrics are calculated empirically on count-based measures due to variability in read depths. We quantify the bias introduced by empirically calculating proportionality-based association metrics in count data. Additionally, we propose a means of estimating these metrics within a logit-normal multinomial model in pursuit of more accurate estimates. The model-based estimates are shown to outperform empirical estimates in simulated data and are applied to a mouse embryonic stem cell single-cell sequencing dataset, as well as a pediatric-onset multiple sclerosis metagenomic dataset.

Animals

Biochemical and Genomic Underpinnings of Carotenoid Colour Variation Across a Hybrid Zone Between South Asian Flameback Woodpeckers.

Colouration and patterning have been implicated in lineage diversification across various taxa, as colour traits are heavily influenced by sexual and natural selection. Investigating the biochemical and genomic foundations of these traits therefore provides deeper insights into the interplay between genetics, ecology and social interactions in shaping the diversity of life. In this study, we assessed the pigment chemistries and genomic underpinnings of carotenoid colour variation in naturally hybridising Dinopium flamebacks in tropical South Asia. We employed reflectance spectrometric analysis to quantify species-specific plumage colouration, High-Performance Liquid Chromatography (HPLC) to elucidate the feather carotenoids of flamebacks across the hybrid zone, and Genome-Wide Association Study (GWAS) using next-generation sequencing data to uncover the genetic factors underlying carotenoid colour variation in flamebacks. Our analysis revealed that the red mantle feathers of D. psarodes primarily contained astaxanthin, with small amounts of other 4-keto-carotenoids. In contrast, the yellow mantle feathers of D. benghalense predominantly contained lutein and 3'-dehydro-lutein, alongside minor amounts of zeaxanthin, β-cryptoxanthin and canary-xanthophylls A and B. Hybrids with an intermediate, orange colouration deposited all of these pigments in their mantle feathers, with notably higher concentrations of carotenoids with ε-end rings. The GWAS analysis identified the CYP2J2 gene, which plays a role in carotenoid ketolation, as associated with the expression of carotenoid colouration. Read depth data suggested variation in copy number of this gene in flamebacks. These findings contribute to the growing knowledge of avian carotenoid metabolism and highlight how genomic architecture can influence phenotypic diversity.

Animals

Clinical and functional characterization of a novel homozygous non-canonical splice mutation (c.1910-15_1910-11delinsTTACA) in CEP290 causing Joubert syndrome.

BACKGROUND: Joubert syndrome (JS) is a rare, predominantly autosomal recessive neurodevelopmental disorder characterized by hypotonia, motor delay, intellectual disability, oculomotor apraxia, and the hallmark "molar tooth sign" on axial view of MRI. JS is genetically heterogeneous, with pathogenic variants identified in more than 40 genes involved in primary cilia function. Among these, CEP290 is one of the most frequently mutated genes. RESULTS: In this study, we investigated two children-an 11-year-old boy (the proband) and his 5-year-old sister-both presenting with a similar phenotype consistent with JS. The parents, who self-identified as Chechen, reported distant consanguinity. The family also included a healthy 13-year-old daughter. The proband had previously been evaluated by a neurologist and underwent whole-genome sequencing (WGS); however, no causative variants were identified initially. After phenotype reassessment by a clinical geneticist, we performed a reanalysis of the raw WGS data and identified a novel homozygous intronic variant of uncertain significance (VUS), c.1910-15_1910-11delinsTTACA in CEP290 (NM_025114.4). Sanger sequencing confirmed that both the proband and his affected sister were homozygous for this variant, which they inherited from their heterozygous parents. Their healthy sister did not carry the variant. mRNA-sequencing and targeted cDNA sequencing (read depth ~ 100,000x) demonstrated that this intronic variant causes completely aberrant splicing of CEP290 pre-mRNA. Predominantly this variant causes the skipping of exon 20 in the main CEP290 transcript. Alternatively, the variant results in partial inclusion of intron 19 into the mRNA, elongation of exon 20 by 58 nucleotides, and a homozygous substitution chr12:88114573 (ACTGTGTA> TTACAGTA). No canonical mRNA isoform was detected when the variant was homozygous. Both the predicted severe truncation and the likely degradation of aberrant transcripts through nonsense-mediated decay (NMD) would correspond to complete loss of CEP290 function. Following the reclassification of this VUS to likely pathogenic, the family was able to pursue in vitro fertilization (IVF) with preimplantation genetic testing for monogenic disorders (PGT-M). CONCLUSION: Our study highlights the critical importance of proper phenotyping prior to referral for WES/WGS as well as of combining NGS with functional mRNA studies to achieve a molecular diagnosis for patients with predicted splice-site mutations in JS-associated genes. It also emphasizes the need for functional reassessment of VUS when genomic data are expected to guide reproductive decision-making within affected families.

Humans

APAV: An advanced pangenome analysis and visualization toolkit.

Traditional pangenome analysis focuses on gene presence/absence variations (gene PAVs). However, the current methods for gene PAV analysis are insensitive to detect small but valuable mutations within gene regions, and they overlook variations in intergenic regions. Additionally, the visual inspection of PAVs is an important but time-consuming step for pangenome analysis and result interpretation. To address these issues, we present APAV, an advanced toolkit designed for comprehensive PAV analysis and visualization. It integrates gene element-level PAV analysis and provides PAV analysis for arbitrary given regions in a genome. The resulted PAV profile can be visualized and investigated interactively with reports in HTML format, enabling researchers to conveniently verify sequencing read depth, target region coverage, and intervals of absence for each PAV. Furthermore, APAV offers various subsequent analysis and visualization functions based on the PAV profile table, including basic statistics, sample clustering, genome size estimation, and phenotype association analysis. We demonstrated the capability of APAV with pangenome analysis of tumor genomes and rice genomes. Performing PAV analysis at the element level not only provides more accurate information about the variations but also uncovers a larger number of variations for the phenotype-genotype association studies. In the rice genome analysis, we identified over twenty thousand distributed genes and more than fifty thousand distributed genetic elements. In the tumor genome analysis, element-level analysis revealed approximately three times as many phenotype-related genes as gene-level analysis. This indicates that altering the PAV unit from genes to smaller segments or elements can lead to more biological insights.

Software