PubMed HealthSearch

PubMed · 42667620

Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.

Abstract

Long-read sequencing technologies have revolutionized human genome exploration at an unparalleled resolution, particularly facilitating the analysis of structural variation (SV) at haplotype resolution. Here, we present a protocol for using cuteHap, a robust framework for haplotype-aware SV detection through phased alignment reads generated by diverse long-read sequencing platforms. We describe procedures for single-nucleotide variant (SNV) calling, read phasing, SV calling, and genotyping. We also establish a benchmarking pipeline to evaluate the detected SV callsets. For complete details on the use and execution of this protocol, please refer to Cao et al.1.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shuqi Cao, Chuanmin Wu, Yuejin He, Tao Jiang. 2026-08-28. Protocol for haplotype-resolved structural variant detection via long-read sequencing using cuteHap.. https://doi.org/10.1016/j.xpro.2026.104810

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Protocol for telomere-to-telomere assembly of Borrelia genomes using a hybrid method.

Borrelia has a linear chromosome and linear plasmids capped by hairpin telomeres that short-read sequencing cannot resolve. Here, we present a protocol for telomere-to-telomere assembly of Borrelia genomes. We describe steps for spanning B. burgdorferi culture, DNA extraction, and sequencing through hybrid genome assembly to generate complete Borrelia genomes. The pipeline integrates Oxford Nanopore long reads and Illumina short reads to assemble hairpin telomeres, resolve paralogous linear and circular plasmids, and annotate and validate the assembled complete Borrelia genome. For complete details on the use and execution of this protocol, please refer to Amin et al.1.

Bioinformatics

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics

Integrative analysis of the roles and prognostic value of RNA-binding proteins in papillary renal cell carcinoma.

RNA-binding proteins (RBPs) serve essential roles in various cancer types, but their functions in papillary renal cell carcinoma (pRCC) have not been elucidated to date. In our work, differentially expressed RBPs in pRCC were identified after acquisition of RNA-sequencing and clinical data related to pRCC from The Cancer Genome Atlas database(TCGA). Functional enrichment analysis and protein interaction network analysis, along with univariate and multivariate Cox regression analyses, were performed to uncover potential biological effects of the identified RBPs and screen the hub RBPs for pRCC prognosis. We identified 251 up-regulated and 129 down-regulated RBPs, and filtered out seven hub RBPs, namely, SRSF8, CD3EAP, HBS1L, ELAC2, MRPL34, NOP2 and IGF2BP2, for their prognostic relevance. A prognostic risk score model for overall survival of pRCC patients was constructed based on the seven hub RBPs. Further analysis showed that the low-risk group had higher survival rate than the high-risk group in both training and validation cohorts. The predictive accuracy was verified in the Human Protein Atlas database.In addition, we introduced the GSE15641 dataset from the Gene Expression Omnibus (GEO) database for independent external validation, and confirmed the expression levels of HBS1L, MRPL34 and IGF2BP2 through real-time quantitative PCR (RT-qPCR) and Western blotting (WB) using human renal tubular epithelial cell line HK-2 and human papillary renal cell carcinoma cell line Caki-2. In pRCC, CD3EAP was significantly elevated, while ELAC2, IGF2BP2, MRPL34, SRSF8 and HBS1L were significantly reduced. There was no significant difference between tumor and normal tissues in NOP2 expression. Risk score and tumor grade were independent prognostic factors associated with overall survival. In addition, we established a nomogram based on the seven prognostic RBPs to help predict overall survival at 1-3 years. In conclusion, seven differentially expressed hub RBPs were identified as potential prognostic biomarkers for pRCC. Our prognostic model might serve as a support for better treatment decision-making. Our work could provide potential new ideas for diagnosis and research on targeted drugs for pRCC.

Bioinformatics