PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genome, Human”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

Nuclear-lamin-guided plastic positioning and folding of the human genome.

The human genome exhibits a highly ordered hierarchical architecture, yet the mechanisms governing its large-scale organization remain poorly understood. Here, we generate lamin single-, double-, and triple-knockout human embryonic and mesenchymal stem cells (hESCs and hMSCs) to investigate the role of lamins in the spatial organization of the human genome. Complete lamin depletion in hMSCs triggers extensive genome repositioning, disrupts chromosome territories, and dissolves long-range compartment clustering and mega-loops. Lamin loss affects both the nuclear periphery and interior, causing partial inversion and dispersion of nuclear speckles, accompanied by reduced global transcription and impaired stem cell homeostasis. Re-expression of wild-type lamin A, which interacts with the speckle scaffold protein SON, partially restores the organizational and transcriptional defects, while the disease-associated E161K mutant disrupts SON binding and shows limited recovery. Our results elucidate the multifaceted roles of lamins in nuclear organization and link their dysfunction to the pathogenesis of laminopathies.

Humans

Retrotransposon-based mechanisms for transgene addition to the human genome.

When human disease arises from a loss of function caused by diverse mutant alleles of the same gene, the patient population could be best served by a clinical therapy that achieves genome safe-harbor supplementation with a functional transgene. Until recently, transgene delivery strategies have shared the disadvantages of induced immune responses and/or genome mutagenesis from untargeted DNA insertion. As a different strategy, several groups recently described the use of retrotransposon proteins to accomplish transgene insertion by RNA-templated cDNA synthesis directly into the genome. In some strategies, gene insertion relies on the retrotransposon protein to bring a transgene-encoding template RNA to the target site. Retrotransposon protein positioning of template RNA for cDNA synthesis minimizes the requirement for RNA base-pairing to target-site DNA. This review presents an overview of RNA-templated DNA synthesis in cells as backdrop for describing recent uses of retrotransposon reverse transcriptases to supplement the human genome.

Journal Article

Archaic ancestry inference in imputed ancient human genomes.

When modern humans expanded from Africa into Eurasia, they interbred with archaic hominins such as Neanderthals and Denisovans. This introgression shaped human evolution, yet most insights have been gained from present-day genomes, leaving little known about how archaic variants evolved after interbreeding. Ancient genomes offer a direct view of this process, but low coverage and poor quality have limited their use. Recent advances in genotype imputation offer a way to overcome these challenges by reconstructing missing information from reference panels and recovering evolutionary signals from low-coverage data. Here, we show that imputation enables accurate detection and quantification of archaic introgression in ancient genomes, improves local archaic ancestry inference, and that regions of archaic ancestry are imputed with especially high accuracy. We further demonstrate that imputed genomes can reconstruct the trajectories of introgressed haplotypes, distinguish populations across time and geography, and identify both known and additional candidates for adaptive introgression.

Humans

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning

Genomic Language Model for Predicting Enhancers and Their Allele-Specific Activity in the Human Genome.

Predicting and deciphering the regulatory logic of enhancers is a challenging problem, due to the intricate sequence features and lack of consistent genetic or epigenetic signatures that can accurately discriminate enhancers from other genomic regions. Recent machine-learning based methods have spotlighted the importance of extracting nucleotide composition of enhancers but failed to learn the sequence context and perform suboptimally. Motivated by advances in genomic language models, we developed DNABERT-Enhancer, a novel enhancer prediction method, by applying DNABERT pre-trained language model on the human genome. We trained two different models, using large collection of enhancers curated from the ENCODE registry of candidate cis-Regulatory Elements. The best fine-tuned model achieved 88.05% accuracy with Matthews correlation coefficient of 76% on independent set aside data. Further, we present the analysis of the predicted enhancers for all chromosomes of the human genome by comparing with the enhancer regions reported in publicly available databases. Finally, we applied DNABERT-Enhancer along with other DNABERT based regulatory genomic region prediction models to predict candidate SNPs with allele-specific enhancer and transcription factor binding activity. The genome-wide enhancer annotations and candidate loss-of-function genetic variants predicted by DNABERT-Enhancer provide valuable resources for genome interpretation in functional and clinical genomics studies.

Journal Article

A comparison of software for analysis of rare and common short tandem repeat (STR) variation using human genome sequences from clinical and population-based samples.

Short tandem repeat (STR) variation is an often overlooked source of variation between genomes. STRs comprise about 3% of the human genome and are highly polymorphic. Some cause Mendelian disease, and others affect gene expression. Their contribution to common disease is not well-understood, but recent software tools designed to genotype STRs using short read sequencing data will help address this. Here, we compare software that genotypes common STRs and rarer STR expansions genome-wide, with the aim of applying them to population-scale genomes. By using the Genome-In-A-Bottle (GIAB) consortium and 1000 Genomes Project short-read sequencing data, we compare performance in terms of sequence length, depth, computing resources needed, genotyping accuracy and number of STRs genotyped. To ensure broad applicability of our findings, we also measure genotyping performance against a set of genomes from clinical samples with known STR expansions, and a set of STRs commonly used for forensic identification. We find that HipSTR, ExpansionHunter and GangSTR perform well in genotyping common STRs, including the CODIS 13 core STRs used for forensic analysis. GangSTR and ExpansionHunter outperform HipSTR for genotyping call rate and memory usage. ExpansionHunter denovo (EHdn), STRling and GangSTR outperformed STRetch for detecting expanded STRs, and EHdn and STRling used considerably less processor time compared to GangSTR. Analysis on shared genomic sequence data provided by the GIAB consortium allows future performance comparisons of new software approaches on a common set of data, facilitating comparisons and allowing researchers to choose the best software that fulfils their needs.

Humans

Public opinion survey on heritable human genome editing in South Africa: a study protocol.

Heritable human genome editing (HHGE) presents new possibilities for the prevention of genetic diseases but also raises ethical and societal questions. While international surveys have explored public attitudes, particularly in high-income countries, there is a lack of large-scale empirical data from the Global South. In South Africa, previous work used deliberative public engagement to examine public perspectives. The present study aims to complement this by capturing public opinion through a cross-sectional survey, enabling direct comparison with deliberative findings. This study will recruit 400 adult participants residing in South Africa using targeted Facebook advertisements. A two-phase sampling process will be employed: initial screening for demographic information, followed by stratified sampling to ensure a representative South African population. The opinion survey consists of 19 HHGE scenarios, each explored through private and public moral lenses. Additionally, participants will indicate their interpretation of 'safe and effective' genome editing. Quantitative data will be analysed using descriptive statistics, chi-square tests, and logistic regression. Qualitative responses will undergo thematic analysis using both manual coding and generative AI tools under human oversight. The study includes two stages of informed consent and ensures data confidentiality through strict data handling protocols. Results will be disseminated in peer-reviewed journals and policy forums. The study will also generate a secondary dataset for evaluating AI-assisted qualitative analysis, to be conducted under separate ethical clearance.

Humans

Environmentally responsible human genomic data governance: points for consideration.

We introduce five points for integrating environmental ethics into human genomic data governance: (i) recognizing the ethical imperative to consider environmental impacts of human genomic data; (ii) fostering collective responsibility for environmental harms; (iii) prospectively assessing benefits and harms; (iv) anticipating barriers to integration of environmental ethics into genomic data governance; and (v) meaningfully engaging all interest-holders. These points will be useful to all involved in the genomic data ecosystem.

Letter

Chromosome-specific centromeric patterns define the centeny map of the human genome.

Centromeres are epigenetically specified by distinct chromatin, whereas their DNA varies between species and individuals. This extensive sequence divergence makes comparative analyses between centromeres challenging. In this study, we identified a chromosome-specific architectural pattern across the human genome, defined by the conserved spacing of a functionally relevant centromeric DNA motif. The distribution of these sites along chromosome arms constitutes the human "centeny map." By using a custom Genomic Centromere Profiling (GCP) pipeline, we leveraged the motif's position, orientation, and organization to construct structural models that enable reclassification of human chromosomal clusters, detection of centromere expansion, and identification of structural variants and misassembled regions. The high-resolution maps derived from this pattern not only provide a framework for comparative analysis of centromeres across evolution and disease but also offer a new dimension for chromosome annotation, assembly, and characterization.

Humans

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans

A token-pruning framework enables efficient representation of the human genome for RNA modification analysis.

MOTIVATION: Modelling long genomic sequences remains challenging due to extreme sequence length, high redundancy, and the need for biological interpretability. Although Transformer-based architectures have achieved strong performance across genomic tasks, their high computational cost and reliance on fixed tokenization strategies limit their scalability and ability to focus on biologically informative regions. RESULTS: We propose ATSFormer, a token-pruning Transformer framework for efficient and biologically informed genomic sequence modelling. ATSFormer incorporates an attention-guided and parameter-free Adaptive Token Sampling (ATS) module into Transformer layers. Guided by attention-derived importance scores, ATS dynamically retains informative tokens while probabilistically discarding redundant ones, thereby reducing sequence length, FLOPs, and memory usage without introducing additional learnable parameters or extra training procedures. Importantly, the retained tokens correspond to key contributors to model predictions, enabling ATSFormer to highlight biologically meaningful sites and sequence motifs. We evaluated ATSFormer on four benchmark RNA modification datasets derived from RMVar 2.0, covering A-to-I, m1A, m5C, and m7G. Experimental results show that ATSFormer consistently outperforms existing state-of-the-art methods while achieving substantial computational savings. Furthermore, structural analysis using AlphaFold3 supports the biological relevance of the motifs identified by ATSFormer. AVAILABILITY AND IMPLEMENTATION: The source data and code are freely available at GitHub (https://github.com/1gao2/ATSFormer) and Zenodo (https://doi.org/10.5281/zenodo.21813541).

Humans

PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data.

MOTIVATION: Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. RESULTS: We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy-utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. AVAILABILITY AND IMPLEMENTATION: The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg.

Humans

An archaic reference-free method to jointly infer Neanderthal and Denisovan introgressed segments in modern human genomes.

Admixture between populations is a common feature of human history. Admixture events introduce new genetic variation that can fuel evolution. Characterizing the significance of admixture events on the evolution of populations across various species is of great interest to evolutionary geneticists. Local Ancestry Inference (LAI) methods infer genetic ancestry of an individual at a particular chromosomal location. Certain methods specialize in detecting archaic introgression, which consists of interbreeding between modern and archaic humans like Neanderthals and Denisovans. Most current LAI methods allow the detection of a single archaic ancestry, and post-processing may distinguish between multiple waves of introgression. These methods vary in how they choose archaic or modern reference genomes for the inference. Here, we present a new HMM-based method (DAIseg), which has the advantage of simultaneously distinguishing between multiple waves of ancient and recent admixture, using only modern human reference genomes. Simulations demonstrate that DAIseg achieves higher overall performance than state-of-the-art methods. We also apply DAIseg to Papuan populations to jointly detect Denisovan and Neanderthal introgressed segments, and identify a higher number of archaic segments than previous methods. Analysis of inferred introgressed segments, shows that we can identify evidence for two Denisovan introgression events in Papuans. Overall, on top of being able to deal with both Archaic and recent admixture, DAIseg provides a more principled approach for detecting and classifying Denisovan and Neanderthal segments which will improve downstream analysis of introgressed segments to infer the impact of archaic introgression in humans.

Denisovan

Integration of therapeutic cargo into the human genome with programmable type V-K CAST.

CRISPR-associated (Cas) transposases (CAST) are RNA-guided systems capable of programmable integration of large segments of DNA without creating double-strand breaks. Engineered Cascade CAST function in human cells but are challenging to deploy due to the complexity of the targeting components. Unlike Cascade, which require three Cas proteins, type V-K CAST require a single Cas12k effector for targeting. Here, we show that compact type V-K CAST from uncultivated microbes are repurposable for programmable DNA integration into the genome of human cells. Engineering for nuclear localization and function enables integration of a therapeutically relevant transgene at a safe-harbor site in multiple human cell types. Notably, off-targets are rare events reproducibly found in specific genomic regions. These CAST advancements are expected to accelerate applications of genome editing to therapeutic development, biotechnology, and synthetic biology.

Humans

Human Genome REWRITE for Off-the-Shelf Stem Cells Reveals an "Epigenetic Ghost".

Human leukocyte antigen (HLA) polymorphism hinders off-the-shelf cell therapies. We developed REWRITE, a modular platform for iterative, scar-minimized genome writing of synthetic constructs >100 kb in human pluripotent stem cells (hPSCs). Using REWRITE, we deleted 105-209 kb of the HLA locus and installed synthetic 24 kb or 100 kb HLA haplotypes, and a 62 kb antigen-processing locus. This uncovered a persistent, heritable "epigenetic ghost" - an active state lingering despite genetic removal - whose resolution to a silenced default state is driven by native intergenic DNA. These loci restored inducible expression in key lineages, sparing cells from NK-mediated killing and establishing HLA-matched T-cell tolerance, enabling off-the-shelf cell therapies. REWRITE facilitates extensible programming of multigenic functions in allogeneic human cells - from immune design to genome architecture discovery.

Journal Article

Learning a pairwise epigenomic and transcription factor binding association score across the human genome.

MOTIVATION: Identifying pairwise associations between genomic loci is an important challenge for which large and diverse collections of epigenomic and transcription factor (TF) binding data can potentially be informative. RESULTS: We developed Learning Evidence of Pairwise Association from Epigenomic and TF binding data (LEPAE). LEPAE uses neural networks to quantify evidence of association for pairs of genomic windows from large-scale epigenomic and TF binding data along with distance information. We applied LEPAE using thousands of human datasets. We show using additional data that LEPAE captures biologically meaningful pairwise relationships between genomic loci, and we expect LEPAE scores to be a resource. AVAILABILITY AND IMPLEMENTATION: The LEPAE scores and the software are available at https://github.com/ernstlab/LEPAE.

Humans