PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genetic Privacy”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

14 recordsLinked to original sources

Privacy-preserving framework for genomic computations via multi-key homomorphic encryption.

MOTIVATION: The affordability of genome sequencing and the widespread availability of genomic data have opened up new medical possibilities. Nevertheless, they also raise significant concerns regarding privacy due to the sensitive information they encompass. These privacy implications act as barriers to medical research and data availability. Researchers have proposed privacy-preserving techniques to address this, with cryptography-based methods showing the most promise. However, existing cryptography-based designs lack (i) interoperability, (ii) scalability, (iii) a high degree of privacy (i.e. compromise one to have the other), or (iv) multiparty analyses support (as most existing schemes process genomic information of each party individually). Overcoming these limitations is essential to unlocking the full potential of genomic data while ensuring privacy and data utility. Further research and development are needed to advance privacy-preserving techniques in genomics, focusing on achieving interoperability and scalability, preserving data utility, and enabling secure multiparty computation. RESULTS: This study aims to overcome the limitations of current cryptography-based techniques by employing a multi-key homomorphic encryption scheme. By utilizing this scheme, we have developed a comprehensive protocol capable of conducting diverse genomic analyses. Our protocol facilitates interoperability among individual genome processing and enables multiparty tests, analyses of genomic databases, and operations involving multiple databases. Consequently, our approach represents an innovative advancement in secure genomic data processing, offering enhanced protection and privacy measures. AVAILABILITY AND IMPLEMENTATION: All associated code and documentation are available at https://github.com/farahpoor/smkhe.

Computer Security

PRISM: privacy-preserving rare disease analysis using fully homomorphic encryption.

MOTIVATION: Rare diseases affect millions of people worldwide, yet their genomic foundations remain poorly understood due to limited patient data and strict privacy regulations, such as the General Data Protection Regulation (GDPR) (https://gdpr.eu/tag/gdpr/) in March 2025. These restrictions can hinder the collaborative analysis of genomic data necessary for uncovering disease-causing variants. RESULTS: We present PRISM, a novel privacy-preserving framework based on fully homomorphic encryption (FHE) that facilitates rare disease variant analysis across multiple institutions without exposing sensitive genomic information. To address the challenges of centralized trust, PRISM is built upon a Threshold FHE scheme. This approach decentralizes key management across participating institutions and ensures no single entity can unilaterally decrypt sensitive data. Our method filters disease-causing variants under recessive, dominant, and de novo inheritance models entirely on encrypted data. We propose two algorithmic variants: a multiplication-intensive (MUL-IN) approach and an addition-intensive (ADD-IN) approach. The ADD-IN algorithms minimize the number of costly multiplication operations, enabling up to a 17× improvement in runtime for recessive/dominant filtering and 22× for de novo filtering, compared to MUL-IN methods. While ADD-IN produces larger ciphertexts, efficient parallelization via SIMD and multithreading allows it to handle millions of variants in reasonable time. To the best of our knowledge, this is the first study that utilizes FHE for privacy-preserving rare disease analysis across multiple inheritance models, demonstrating its practicality and scalability in a single-cloud setting. AVAILABILITY AND IMPLEMENTATION: The source code and the data used in this work can be found in https://github.com/mdppml/PRISM.git.

Computer Security

PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data.

MOTIVATION: Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. RESULTS: We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy-utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. AVAILABILITY AND IMPLEMENTATION: The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg.

Humans

Private detection of relatives in forensic genomics using homomorphic encryption.

BACKGROUND: Forensic analysis heavily relies on DNA analysis techniques, notably autosomal Single Nucleotide Polymorphisms (SNPs), to expedite the identification of unknown suspects through genomic database searches. However, the uniqueness of an individual's genome sequence designates it as Personal Identifiable Information (PII), subjecting it to stringent privacy regulations that can impede data access and analysis, as well as restrict the parties allowed to handle the data. Homomorphic Encryption (HE) emerges as a promising solution, enabling the execution of complex functions on encrypted data without the need for decryption. HE not only permits the processing of PII as soon as it is collected and encrypted, such as at a crime scene, but also expands the potential for data processing by multiple entities and artificial intelligence services. METHODS: This study introduces HE-based privacy-preserving methods for SNP DNA analysis, offering a means to compute kinship scores for a set of genome queries while meticulously preserving data privacy. We present three distinct approaches, including one unsupervised and two supervised methods, all of which demonstrated exceptional performance in the iDASH 2023 Track 1 competition. RESULTS: Our HE-based methods can rapidly predict 400 kinship scores from an encrypted database containing 2000 entries within seconds, capitalizing on advanced technologies like Intel AVX vector extensions, Intel HEXL, and Microsoft SEAL HE libraries. Crucially, all three methods achieve remarkable accuracy levels (ranging from 96% to 100%), as evaluated by the auROC score metric, while maintaining robust 128-bit security. These findings underscore the transformative potential of HE in both safeguarding genomic data privacy and streamlining precise DNA analysis. CONCLUSIONS: Results demonstrate that HE-based solutions can be computationally practical to protect genomic privacy during screening of candidate matches for further genealogy analysis in Forensic Genetic Genealogy (FGG).

Humans

Examining gaps in institutional policies for clinical genomic data sharing: A cross-jurisdictional study.

The sharing of data generated by clinical genetic and genomic testing without explicit consent is important for timely diagnosis and treatment. While many jurisdictions permit the sharing of identifiable data for direct clinical care, institutional policies vary in how clearly they specify key elements, including when sharing is permitted, what data are covered, and what safeguards apply. Greater clarity around these elements may support responsible data sharing while balancing timely care with transparency and appropriate protections. We conducted a mixed-methods content analysis of data-sharing and privacy policies from 33 clinical genomic institutions across 17 countries and regions. Using a predefined analytical framework, we assessed how policies document key governance elements relevant to sharing without explicit consent. Two independent reviewers extracted information about clinical contexts, data types, justifications, and protections. Although 70% of institutions described circumstances permitting data sharing without explicit consent, most policies did not clearly define the scope or governance of such sharing. Policies also rarely distinguished clinical from research or secondary use and inconsistently specified privacy and security safeguards. While sharing was commonly justified for clinical care (78.3%) or testing services (43.5%), data recipient roles and onward-sharing expectations were often left undefined. This uneven documentation could make it difficult for clinical teams and institutional decision-makers to identify and justify decisions about what is permitted and under what conditions. A guidance framework specifying core governance elements and corresponding protections could help institutions communicate their governance choices more clearly and support comparable baseline practices for responsible data sharing.

Information Dissemination

Listening forward: emerging roles of bioacoustics in ecology, evolution, and conservation.

Bioacoustics is increasingly shifting from a mostly descriptive pursuit to one that can anticipate ecological change. Recent innovations-from autonomous recording units and edge-computing sensors to speech-inspired feature extraction and machine-learning techniques like transfer learning, unsupervised discovery, and explainable AI-are transforming the study of animal communication. These advances let us work at scales previously difficult to imagine. Automated species recognition, individual identification, and even tracking cultural evolution over decades are now within reach. Entire ecosystem soundscapes can be mapped with unprecedented resolution. Looking ahead, global listening networks, adaptive acoustic indices, and live biodiversity dashboards seem increasingly realistic. We may soon build digital models that simulate communication networks under future scenarios. Closer integration with genomics, physiology, and robotics could link vocal traits to their genetic, physiological, and ecological drivers. Challenges remain, including data governance, acoustic privacy, and equitable access to the planet's sonic heritage. Bioacoustics may be on the way to becoming a predictive, integrative science - one particularly well suited to monitoring, interpreting, and helping safeguard life's communication systems in a rapidly changing world.

Animals

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans

The Impact of Chatbot Type and Normative Messaging on Chatbot Usage Intention Based on the Health Technology Acceptance Model: Randomized Controlled Trial.

BACKGROUND: Digital health tools, such as health chatbots, may improve access to scalable health support, but adoption remains inconsistent. Existing models do not fully integrate technology acceptance factors with health motivation factors relevant to digital health use. OBJECTIVE: This study proposed and tested the health technology acceptance model and examined whether normative message framing and chatbot type were associated with health motivation, technology acceptance, and intention to use a health chatbot. METHODS: In October 2025, we conducted a 4 &#xd7; 2 between-participants online experiment with 1000 US adults recruited from a nationally representative YouGov panel. Participants were randomized to 1 of 8 conditions varying norm message type (self-oriented, peer-oriented, expert-oriented, or family-oriented) and chatbot type (AI-powered or rule-based) in a cancer prevention and genetic risk information scenario. Outcomes included descriptive norms, injunctive norms, perceived susceptibility, perceived severity, perceived benefits, self-efficacy, perceived ease of use, trust, privacy concerns, and usage intention. Data were analyzed using a multivariate ANOVA with Bonferroni-adjusted post hoc tests and multiple linear regression. RESULTS: Peer-oriented and family-oriented messages produced higher usage intention than expert-oriented messages, and peer-oriented messages also increased descriptive norms, injunctive norms, self-efficacy, and trust. AI-powered chatbots were associated with higher usage intention (P=.02) and greater trust (P=.008) than rule-based chatbots. In regression analyses, the model explained 50.8% of the variance in usage intention. Usage intention was positively associated with descriptive norms (&#x3b2;=0.087; P=.003), injunctive norms (&#x3b2;=0.078; P=.009), perceived susceptibility (&#x3b2;=0.051; P=.03), perceived benefits (&#x3b2;=0.253; P<.001), and trust (&#x3b2;=0.33; P<.001), and negatively associated with perceived severity (&#x3b2;=-0.047; P=.049) and privacy concerns (&#x3b2;=-0.11; P<.001). Perceived ease of use and self-efficacy were not significant predictors. CONCLUSIONS: The health technology acceptance model was a useful framework for explaining the intention to use a health chatbot by combining technology acceptance and health motivation constructs. Both social design features and chatbot design features shaped adoption-related beliefs, with peer-oriented and family-oriented framing and AI-powered chatbots showing particular promise. Trust and privacy concerns remained central determinants of intended use.

Humans

bioETH-PRS: confidential polygenic risk scoring with smart contracts on an FHE-enabled blockchain.

Polygenic risk scores (PRSs) aggregate genetic effect estimates to predict disease susceptibility, yet calculating one through an external service can require exposing raw genotype data. Homomorphic encryption hides those data during the calculation but, in prior work, still places a designated evaluator in a position of trust. We present bioETH-PRS, a protocol that replaces the evaluator with publicly auditable smart contracts on a blockchain supporting Fully Homomorphic Ethereum Virtual Machine (fhEVM). Using integer-exact encrypted arithmetic, bioETH-PRS computes the PRS dot product entirely in the encrypted domain, so genotype dosages and, at the model provider's discretion, the GWAS weights stay hidden from the parties performing the computation. A fixed-point encoding represents signed weights as nonnegative integers within a bound that rules out overflow, recovering the score to the precision of the published weights. A four-contract architecture separates data custody, model publication, computation, and output release, and supports both a classic path that stores encrypted inputs and an appreciably cheaper streaming path that discards them. A release oracle can return a randomized risk category instead of the raw score, limiting what a repeated querier learns. Prototype evaluation on real GWAS fixtures, including a run on a public testnet, shows cost growing linearly with variant count and suggests the approach may be practical where transaction fees are low. Trust is redistributed rather than removed: the system still depends on the contracts, the blockchain, and the fhEVM services. We evaluate additive models of moderate size, not genome-wide or clinical use.

Blockchain

How Have Massively Parallel Sequencing Technologies Furthered Our Understanding of Oncogenesis and Cancer Progression?

Massively parallel sequencing technologies have been a boon to many fields of biological science, including oncology. Cancer is an umbrella term for many diseases featuring abnormal cellular growth due to genetic and epigenetic aberrations. Advances in sequencing technology allow for interrogation of the DNA and RNA of cancer cells and other cells in the tumor microenvironment down to a single-base resolution. However, these strides come after a rich history of ground-breaking biological assays, like the discovery of the Philadelphia chromosome in the context of leukemia. Many specific genetic and epigenetic modifications have been implicated in oncogenesis, cancer progression, and response to treatment. Sequencing technologies have also helped to associate populations of bacteria in the microbiome to cancer development and prognosis. However, all this new information, especially when procured via high-throughput methods, comes at the cost of being more computationally and staff-resource intensive. There is also more risk to the privacy of the individuals with sequenced genomes. Notwithstanding, the overall benefit of sequencing technologies can greatly outweigh the risks with careful advancements and continued focus on the goal: helping those affected by cancer via precision medicine. Cancer biology has been and will continue to be elucidated by sequencing innovations in ways unimaginable without it.

Humans

The European Health Data Space and the Secondary Use of Sensitive Health Data.

INTRODUCTION: The European Health Data Space (EHDS) is one of the European Union's most ambitious data-governance projects. It aims to create a common framework through which electronic health data can be accessed and reused across Member States for care, research, innovation, policy, and public-interest purposes. Its practical viability depends not only on digital infrastructure, but also on legal, ethical, and organisational harmonisation, particularly for genetic and genomic data. METHODS: This paper examines the EHDS with emphasis on the secondary use of health data. It reviews the EHDS institutional architecture, discusses Finland's Findata as a national model for structured access, and analyses challenges for data holders and data donors, including interoperability, governance burdens, privacy protection, residual re-identification risk, and genomic-data sensitivity. RESULTS: A cross-border cancer-genomics case study shows that the EHDS can streamline data discovery and the routing of access requests, but does not by itself eliminate legal fragmentation, heterogeneous ethics review, and consent-related barriers. DISCUSSION: Effective implementation will require harmonisation beyond infrastructure, including clearer consent standards, more consistent ethics procedures, interoperable metadata, and proportionate safeguards for genomic data.

Electronic Health Records

Beacon Reconstruction Attack: Reconstruction of genomes in genomic data-sharing beacons using summary statistics.

MOTIVATION: Genomic data-sharing beacon protocol, developed by the Global Alliance for Genomics and Health, offers a privacy-preserving mechanism for querying genomic datasets while restricting direct data access. Despite their design, beacons remain vulnerable to privacy attacks. This study introduces a novel privacy vulnerability of the protocol: one can reconstruct large portions of the genomes of all beacon participants by only using the summary statistics reported by the protocol. RESULTS: We introduce a novel optimization-based algorithm that leverages beacon responses and SNP correlations for reconstruction. By optimizing for the SNP correlations and allele frequencies, the proposed approach achieves genome reconstruction with a substantially higher F1-score (70%) compared to baseline methods (45%) on beacons generated using individuals from the HapMap and OpenSNP datasets. We show that reconstructed genomes can be used by downstream applications such as in membership inference attacks against other beacons. Our findings reveal that beacons releasing allele frequencies substantially increase the reconstruction risk, underscoring the need for enhanced privacy-preserving mechanisms to protect genomic data. AVAILABILITY AND IMPLEMENTATION: Our implementation is available at https://github.com/ASAP-Bilkent/Beacon-Reconstruction-Attack.

Genomics

Gene-environment interactions within a precision environmental health framework.

Understanding the complex interplay of genetic and environmental factors in disease etiology and the role of gene-environment interactions (GEIs) across human development stages is important. We review the state of GEI research, including challenges in measuring environmental factors and advantages of GEI analysis in understanding disease mechanisms. We discuss the evolution of GEI studies from candidate gene-environment studies to genome-wide interaction studies (GWISs) and the role of multi-omics in mediating GEI effects. We review advancements in GEI analysis methods and the importance of large-scale datasets. We also address the translation of GEI findings into precision environmental health (PEH), showcasing real-world applications in healthcare and disease prevention. Additionally, we highlight societal considerations in GEI research, including environmental justice, the return of results to participants, and data privacy. Overall, we underscore the significance of GEI for disease prediction and prevention and advocate for integrating the exposome into PEH omics studies.

Humans

Biological Foundation Models for Complex Disease Research and Clinical Translation.

Complex diseases, including cancer, rare genetic disorders, neurodevelopmental and psychiatric conditions, and neurodegenerative diseases, arise from interactions among genetic variation, gene regulation, and cellular states that are difficult to capture using a single data type or biological scale. Biological foundation models address this challenge by treating nucleotides and genes as tokens and learning representations that can be transferred to downstream biomedical and clinical tasks. In this review, we examine two major model classes, genomic sequence foundation models and cell foundation models, and compare their tokenization strategies, model architectures, pretraining objectives, and adaptation methods. We summarize their emerging applications in regulatory variant interpretation, disease-associated cell-state analysis, drug-response prediction, and therapeutic target discovery across complex diseases. We distinguish applications supported by experimental or retrospective validation from those that remain primarily computational or conceptual. We further discuss key challenges to clinical translation, including multimodal data integration, model interpretability, benchmarking, patient-specific prediction, and privacy protection. We highlight future opportunities to integrate biological foundation models with emerging frameworks of medical digital twins, agentic AI, and federated learning. By linking model design to translational goals, this review provides a practical framework for evaluating biological foundation models and their readiness for complex disease research and clinical use.

biological foundation model