PubMed HealthSearch

Biomedical subjects

David Porubsky

Publications and source records attributed to David Porubsky.

4 recordsLinked to original sources

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation.

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Journal Article

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

Complete chromosome 21 centromere sequencing of families with Down syndrome reveals centromere size asymmetry.

Down syndrome, the most common form of human intellectual disability, is caused by nondisjunction and chromosome 21 trisomy (T21). Small centromeres have been hypothesized to contribute to its aetiology and studies on mammals suggest that larger centromeres are more efficiently transmitted, yet complete sequencing of chromosome 21 (chr21) centromeres has been particularly challenging. Using long-read sequencing, we sequenced and assembled the centromeres from eight families that include a child with free T21 (1 trio, 6 child-mother duos, and 1 singleton) all resulting from maternal meiosis I errors. Two of these families carry the smallest chr21 centromeres (143 and 181 kbp) observed in female individuals to date, exhibiting a ~10.7- and ~19.4-fold centromeric α-satellite higher-order repeat array size difference between the maternally inherited homologs, respectively. In both cases, the longer centromere harbors a poorly defined centromere dip region, marked by DNA hypomethylation, in the proband but not in the mother. A comparison of all proband chr21 centromeres (n=24) to those of controls (n=261) shows that small centromeres are not enriched in families with T21 (p-value=0.73); contrarily, chr21 extreme centromere size asymmetry (>10-fold) is unique of T21 (p-value=0.003), suggesting that this feature may represent a genetic risk factor for a subset of families with free T21. Additionally, phylogenetic reconstruction reveals that human chr21 has been particularly prone to such variation with some of the biggest size differences occurring over the last ~17 thousand years of human evolution.

Down syndrome

SVbyEye: a visual tool to characterize structural variation among whole-genome assemblies.

MOTIVATION: We are now in the era of being able to routinely generate highly contiguous (near telomere-to-telomere) genome assemblies of human and nonhuman species. Complex structural variation and regions of rapid evolutionary turnover are being discovered for the first time. Thus, efficient and informative visualization tools are needed to evaluate and directly observe structural differences between two or more genomes. RESULTS: We developed SVbyEye, an open-source R package to visualize and annotate sequence-to-sequence alignments along with various functionalities to process these alignments. The tool facilitates the characterization of complex structural variants in the context of sequence homology helping resolve the mechanisms underlying their formation. AVAILABILITY AND IMPLEMENTATION: SVbyEye is available on GitHub (https://github.com/daewoooo/SVbyEye) and via Zenodo (https://doi.org/10.5281/zenodo.15303553).

Software