PubMed HealthSearch

Biomedical subjects

Arang Rhie

Publications and source records attributed to Arang Rhie.

3 recordsLinked to original sources

Long-read Sequences Mapped to a Complete Reference Genome Uncover Uncaptured Structural Variants across the Beta-globin Cluster in Africans with Sickle Cell Disease.

African genomes are marked by extensive complexity in the number and distribution of variants, yet remain under-represented in genetic databases and the human reference genome. This gap in representation limits the broad application of genomic medicine. Sickle cell disease (SCD) - one of the most common monogenic diseases - has its highest prevalence in Africa, and variation in disease severity has consistently been linked to the beta-globin locus, including levels of fetal hemoglobin (HbF). Modulation of HbF is central to current SCD gene therapies; however, the inherent complexity and variation at the locus in African genomes presents a challenge to translating these advances to Africa. Here, we align long-read single molecule sequences (LRS) targeted to the beta-globin region to the hg38 and T2T-CHM13v2 genome references in 40 individuals with SCD, predominantly recruited from three African countries. We demonstrate that the expanded T2T-CHM13v2 reference sequence at this locus reduces Structural Variant (SV) calls by 70% and uncovers uncaptured single nucleotide variants (SNVs). Across the cluster we report 343 SVs and 196 SNVs that have not been previously reported, including in LRS data from the All of Us project. By including African populations from ethnolinguistic groups that have not been previously surveyed we improve variant resolution and bolster evidence for observed variation. Finally, we identify a common ∼4kb insertion locus overlapping the HBB promoter among individuals with high HbF. These results demonstrate the utility of combining a comprehensive reference genome with LRS in African populations to uncover genomic variation at disease-associated loci.

SNV

Chromosome-specific epigenetic control and transmission of ribosomal DNA arrays in Hominidae genomes.

Ribosomal RNA (rRNA) genes are organized in tandem arrays known as ribosomal DNA (rDNA) on multiple chromosomes in Hominidae genomes. We measured copy number and transcriptional activity status of rRNA gene arrays across multiple individual genomes, revealing an identifiable fingerprint of rDNA copy number and activity. In some cases, entire arrays were transcriptionally silent, characterized by high DNA methylation across the rRNA gene, inaccessible chromatin, and the absence of transcription factors and transcripts. Silent arrays showed reduced association with the nucleolus and decreased interchromosomal interactions, consistent with the model that nucleolar organizer function depends on transcriptional activity. Removing rDNA methylation activated silent arrays. Array activity status remained stable through induced pluripotent stem cell reprogramming and differentiation into cerebral and intestinal organoids. Haplotype tracing in two unrelated family trios showed paternal transmission of silent arrays. We propose that the epigenetic state buffers rRNA gene dosage, specifies nucleolar organizer function, and can propagate transgenerationally.

Epigenesis, Genetic

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article