PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genotype Data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Integrating plant phenotypic and genotypic data in the AGENT project: a BrAPI service implementation.

MOTIVATION: The AGENT project established a network of actively cooperating European genebanks, integrating genomic and phenotypic data from accessions of wheat and barley. Due to specific storage demands for phenotypic and genotypic data, the project used separate database instances and backend technologies to manage integrated phenotypic and genotypic data. RESULTS: We discuss the challenges encountered when integrating dispersed data to serve through a single interface such as the Plant Breeding Application Programming Interface, BrAPI. We examine how the consistent mappability of genebank data to the BrAPI model can enable the implementation of effective services. The advantages of BrAPI in transparently linking distributed data entities through embedded, unique identifiers are highlighted. We present a technical solution involving a BrAPI proxy, which combines and merges separate BrAPI endpoints. Finally, we demonstrate the AGENT BrAPI implementation with an illustrative example that validates a suggested SNP for a trait from the literature by linking phenotypic, genotypic and passport data. AVAILABILITY AND IMPLEMENTATION: The BrAPI proxy implementation and documentation is available at the Python Package Index (https://pypi.org/project/brapi-proxy) and archived in Zenodo (doi: 10.5281/zenodo.19436445). SUPPLEMENTARY INFORMATION: A Jupyter Notebook file for the validation example using a marker-trait relationship found in the literature.

Phenotype

Optimizing genetic ancestry adjustment in DNA methylation studies: a comparative analysis of approaches.

BACKGROUND: Genetic ancestry is an important factor to account for in DNA methylation studies because genetic variation influences DNA methylation patterns. One approach uses principal components (PCs) calculated from CpG sites that overlap with common SNPs to adjust for ancestry when genotyping data is not available. However, this method does not remove technical and biological variations, such as sex and age, prior to calculating the PCs. The first PC is therefore often associated with factors other than ancestry. METHODS: We developed and adapted the adapted EpiAnceR+ approach, which includes (1) residualizing the CpG data overlapping with common SNPs for control probe PCs, sex, age, and cell type proportions to remove the effects of technical and biological factors, and (2) integrating the residualized data with genotype calls from the SNP probes (commonly referred to as rs probes) present on the arrays, before calculating PCs and evaluated the clustering ability and relationship to genetic ancestry. RESULTS: The PCs generated by EpiAnceR+ led to improved clustering for repeated samples from the same individual and stronger associations with genetic ancestry groups predicted from genotype information compared to the original approach. EpiAnceR+ also outperformed the use of DNA methylation PCs or surrogate variables for ancestry adjustment. CONCLUSIONS: We show that the EpiAnceR+ approach improves the adjustment for genetic ancestry in DNA methylation studies. EpiAnceR+ can be integrated into existing R pipelines for commercial methylation arrays, such as 450 K, EPIC v1, and EPIC v2. The code is available on GitHub ( https://github.com/KiraHoeffler/EpiAnceR ).

DNA Methylation

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Clinically Relevant Pharmacogenomic Variant Frequencies in Kazakh, Russian, and Uzbek Population Groups Residing in Kazakhstan.

Central Asian populations remain underrepresented in pharmacogenomic research, limiting the availability of population-specific data for genotype-informed prescribing and precision medicine. This study analyzed clinically relevant pharmacogenomic variant frequencies in Kazakh, Russian, and Uzbek population groups residing in Kazakhstan using genome-wide genotype data from 1301 individuals: Kazakh (n = 1111), Russian (n = 156), and Uzbek (n = 34). ClinPGx, a PharmGKB-based clinical annotation framework that prioritizes variant-drug associations according to levels of evidence, was used to select variants with evidence levels 1A, 1B, and 2A. In total, 112 directly genotyped variants were retained for population-specific allele and genotype frequency analysis. All 112 variants were queried against the gnomAD v4.1 genome and exome reference datasets. Of these, matching allele-frequency data for the predefined reported allele were available in at least one of the two gnomAD datasets for 103 variants, whereas for 9 variants the VEP-based query did not return a matching gnomAD frequency for that allele. Frequencies were reported for the same predefined reported allele across all groups, and differences between the study groups were assessed using 95% confidence intervals, Fisher's exact tests, and false discovery rate correction. Genotype counts and the proportions of individuals carrying at least one copy of the reported allele were also summarized for all selected variants. Several pharmacogenomic variants showed population-specific frequency patterns, including NUDT15 rs116855232, SLCO1B1 rs4149056, VKORC1 rs9934438, and UGT1A1 rs10929302. Comparison with gnomAD showed that the observed frequencies were variant-specific and could not be consistently approximated by a single broad genetic ancestry group. Reference-based population structure analysis provided additional ancestry context and supported separate reporting by population group. The study did not evaluate clinical outcomes or make individual prescribing recommendations, and the small Uzbek sample size limits the precision of frequency estimates for this group, particularly for rare variants. Overall, this study provides a clinically prioritized pharmacogenomic frequency resource for underrepresented population groups in Kazakhstan and supports broader Central Asian representation in pharmacogenomic implementation research.

Central Asia

CRISPR-Cas9-induced genetic mosaicism in three species of the microcrustacean Daphnia.

Genetic mosaicism can arise from in vivo CRISPR-Cas9 gene editing, especially in the embryos. This study evaluates the extent of genetic mosaicism resulted from CRISPR-Cas9-mediated knockout for 11 genes in the freshwater microcrustacean Daphnia magna, Daphnia pulex, and Daphnia sinensis. Based on extensive genotyping data of the asexually produced progenies of successfully edited females, we find strong evidence of mosaicism in 9 of these genes. The genotyping data also suggest that the gene editing activity can take place as early as the one-cell embryo stage and extends into the 32-cell and later stages. This study establishes genetic mosaicism as an important feature of Cas9-mediated gene editing in Daphnia.

Animals

Genome-wide cis-expression Quantitative Trait Loci (eQTL) and transcriptomic signals reveal distinct molecular regulation across correlated feed efficiency traits.

INTRODUCTION: Feed efficiency (FE) is a complex trait which determines livestock production profitability, yet the molecular mechanisms behind it remain unclear. This study investigated the blood transcriptomic profile of lambs, alongside genotype data with the aim to uncover the genetic basis of FE traits such as absolute dry matter intake (DMIabsolute), DMI adjusted for body size (DMIadjusted), average daily live weight gain (ADG), and residual feed intake (RFI). MATERIALS AND METHODS: Bulk RNA-Seq and genotype data were analysed using three complementary approaches: differential gene expression (DGE) analysis, weighted gene co-expression network analysis (WGCNA), and cis-expression Quantitative Trait Loci (cis-eQTL) mapping. These methods were used independently to identify genes and regulatory networks associated with FE traits and to investigate evidence supporting multi-trait candidate gene selection. RESULTS: DGE analysis revealed 2, 24, 85 and 4 differentially expressed genes for DMIabsolute, DMIadjusted, ADG, and RFI (Padjusted < 0.05), functionally enriched in sensory perception, ATP-dependent chromatin remodeling, Notch signaling and immune response pathways. 9 gene modules significantly associated with the FE traits (P &#x2264; 0.05) with correlations ranging from r = -0.56 to 0.49, were identified using WGCNA. Single nucleotide polymorphism (SNP)-level cis-eQTL analysis identified 93 eSNPs associated with 74 genes (false discovery rate (FDR) < 0.05), while permutation-derived gene level analysis identified 280 eGenes (FDR < 0.2, empirical P < 0.03). Across the three analyses, applying thresholds of DGE (Padjusted < 0.05), WGCNA (correlation, P &#x2264; 0.05), and cis-eQTL gene-level significance (empirical P < 0.05), multiple overlapping genes were identified including DNMT3A, KANSL1, NCOR1 for DMIadjusted, ACOX2, FANCF, CIMIP2B, LOC101115106, ARMH2, LOC132657496 for ADG, and LOC114114576 for RFI representing regulators of variations in FE. DISCUSSION: The integration of DGE, WGCNA, and cis-eQTL analyses identified key genes and regulatory mechanisms associated with variation in FE traits. These results highlight that integrated multi-trait candidate gene identification approaches can reveal key genes that lower feed intake while maintaining animal growth, supporting breeding strategies aimed at improving efficiency and long-term economic sustainability in sheep.

average daily gain (ADG)

Understanding recurrence in Mycobacterium avium complex pulmonary disease: genotypic strategies to support clinical decision-making.

Pulmonary disease caused by Mycobacterium avium complex (MAC-PD) is a chronic, recurrent disease, and its high recurrence rate after treatment makes clinical management difficult. Distinguishing whether recurrence is due to persistence of existing strains or reinfection with new strains is essential for establishing treatment strategies, preventing overuse of antimicrobials, and establishing infection control measures. According to reports, 54%-74% of MAC-PD recurrence is due to reinfection, which may be mainly related to environmental reservoirs such as household water supply. In this review, we present various clinical scenarios in which MAC-PD recurrence may occur and examine genotyping techniques as a strategy to distinguish and respond to them. From traditional methods such as IS1245-based restriction fragment length polymorphism, pulsed-field gel electrophoresis, and hsp65 and rpoB gene sequencing to high-resolution analysis techniques such as multilocus sequence testing and whole-genome sequencing, the latest molecular typing methods are comprehensively summarized. Integrating these genotype data into clinical settings, standardizing single-nucleotide polymorphism-based interpretation thresholds, and promoting the establishment of a global MAC strain database will make a substantial contribution to more accurately distinguishing the recurrence mechanisms of MAC-PD and establishing personalized treatment strategies.IMPORTANCEThe global burden of nontuberculous mycobacterial pulmonary disease (PD) is increasing, with Mycobacterium avium (MAC)-PD being the most prevalent and clinically challenging form. Its low treatment success rates, high frequency of recurrence, and persistent environmental exposure complicate both diagnosis and management. A critical clinical issue is determining whether recurrence represents true relapse, due to persistence of the original strain, or reinfection with a new strain, as this guides treatment and prevents overtreatment. Genotypic strategies capable of resolving strain-level differences can improve diagnostic accuracy, prevent misclassification, and ultimately support more informed treatment decisions. Therefore, integrating genotyping data into clinical workflows, standardizing single-nucleotide polymorphism thresholds, and establishing a global MAC strain database will not only support personalized treatment but also enhance the broader public health response to this disease.

Humans

vcfgl: a flexible genotype likelihood simulator for VCF/BCF files.

MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.

Software

Estimating the heritability of longitudinal rate-of-change: genetic insights into PSA velocity in prostate cancer-free individuals.

Serum prostate-specific antigen (PSA) is widely used for prostate cancer screening. While the genetics of PSA levels have been studied to enhance screening accuracy, the genetic basis of PSA velocity, the rate of PSA change over time, remains unclear. The Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trial, a large, randomized study with longitudinal PSA data (15,260 cancer-free males, averaging 5.34 samples per subject) and genome-wide genotype data, provides a unique opportunity to estimate PSA velocity heritability. We developed a mixed model to jointly estimate the heritability of PSA levels at age 54 and PSA velocity. To accommodate the large dataset, we implemented 2 efficient computational approaches: a partitioning and meta-analysis strategy using average information restricted maximum likelihood (AI-REML) and a fast restricted Haseman-Elston (REHE) regression method. Simulations showed that both methods yield unbiased estimates of both heritability metrics, with AI-REML providing smaller variability in the estimation of velocity heritability than REHE. Applying AI-REML to PLCO data, we estimated heritability at 0.32 (s.e. = 0.07) for baseline PSA and 0.45 (s.e. = 0.18) for PSA velocity. These findings reveal a substantial genetic contribution to PSA velocity, supporting future genome-wide studies to identify variants affecting PSA dynamics and improve PSA-based screening.

Humans

Recall-by-genotype of neurodevelopmental disorder copy number variants in a multi-ancestry, healthcare-system biobank.

Clinical biobanks linking electronic health records (EHRs) with genotype data enable the study of genomic risk factors in real-world populations. However, recall-by-genotype (RbG) of psychiatric risk variants in diverse healthcare-system biobanks remains scarce. Leveraging BioMe, a multi-ancestry biobank within the Mount Sinai Health System, we recalled carriers of rare copy number variants (CNVs) that confer increased risk for neurodevelopmental disorders (NDDs) to establish empirical benchmarks for RbG implementation. We recontacted 892 participants: 335 NDD CNV carriers, 217 individuals with schizophrenia without NDD CNVs, and 340 neurotypical controls without NDD CNVs. Participants completed clinical and cognitive assessments. Overall, 18% of recontacted participants responded to recruitment, and 8% completed the study: 30 NDD CNV carriers, 20 individuals with schizophrenia, and 23 controls. The mean age was 48.8 years, 66% were female, and self-reported ancestry was 37% African, 34% Hispanic, and 26% European. Seventy percent of NDD CNV carriers had at least one neuropsychiatric or developmental condition, including mood or anxiety disorders (40%). Among 22 NDD CNV carriers at loci implicated in impaired cognition, performance was lower than controls on Digit Span Backward (&#x3b2;&#x2009;=&#x2009;-1.76, FDR&#x2009;=&#x2009;0.04) and Digit Span Sequencing (&#x3b2;&#x2009;=&#x2009;-2.01, FDR&#x2009;=&#x2009;0.04). NDD CNV carriers also outperformed the schizophrenia group on verbal learning (&#x3b2;&#x2009;=&#x2009;4.5, FDR&#x2009;=&#x2009;0.05). Recall of individuals-including those with psychiatric illness-yielded phenotypes not captured in EHRs and provides empirical benchmarks relevant to RbG implementation and precision psychiatry in diverse healthcare systems.

Journal Article

CLINICAL AND COGNITIVE PHENOTYPING OF COPY NUMBER VARIANTS ASSOCIATED WITH NEURODEVELOPMENTAL DISORDERS FROM A MULTI-ANCESTRY BIOBANK.

Clinical biobanks with electronic health records (EHRs) linked to genotype data continue to expand yielding an opportunity to further characterize disease-relevant genomic risk factors, yet few recall-by-genotype studies from biobanks have been published to date. For example, copy number variants (CNVs) that significantly increase risk for multiple neurodevelopmental disorders (NDDs) and negatively affect neurocognition, may present in up to 2% of population cohorts, with public health implications for ascertaining NDD CNV carriers. From BioMe, a multi-ancestry biobank derived from the Mount Sinai healthcare system (New York, NY), 892 adult participants were recontacted for deep phenotyping, including 335 NDD CNV carriers as well as comparators, 217 individuals with schizophrenia and 340 controls. Clinical and cognitive assessments were administered to each participant. There was no disclosure of genetic information. Eight percent of recontacted biobank participants completed the study (30 NDD CNV carriers across 15 unique loci, 20 schizophrenia and 23 controls). The study sample had a mean age of 48.8 (10.2) years, was 66% female and of diverse ancestry, 36% African, 34% Hispanic, and 26% European. Overall, 70% of 30 NDD-CNV carriers harbored at least one neuropsychiatric or developmental phenotype, including 40% with mood or anxiety disorders. Further, 22 NDD CNV carriers were significantly impaired compared to controls on digit span backwards (Beta=-1.76, FDR=0.04) and digit span sequencing (Beta=-2.01, FDR=0.04), but higher performing than schizophrenia on verbal learning (Beta=4.5, FDR=0.05). Thirty NDD CNV carriers were successfully recruited from a multi-ancestry biobank, as well as healthy controls and low-functioning individuals with schizophrenia. Deep phenotyping corroborated past reports, while also identifying discordance with EHRs. Future recall-by-genotype studies may further benchmark the study design and elucidate feasibility.

Biobank

Profile of the BREATHE cohort for risk-based breast screening in Singapore.

The BREAst screening Tailored for HEr (BREATHE) study was established to pilot a personalised, risk-based breast cancer screening programme for a multi-ethnic Asian population. Between October 2021 and December 2023, 4592 women aged 35-59 were enrolled (73% response rate). Follow-up from February 2022 to June 2024 included 4112 participants (8.6% loss to follow-up). Data collected encompassed demographics, lifestyle factors, reproductive risks, breast cancer awareness, screening behaviours, programme satisfaction, mammography outcomes (density and recall status), and genotype data. Above-average breast cancer risk increased with age: 2% in women aged 35-39, 31% in 40-49, and 42% in 50-59. Findings suggests that women valued the knowledge of their breast cancer risk and were motivated to attend mammography. The BREATHE study confirms the feasibility and acceptability of a personalised risk-based screening programme in a multi-ethnic Asian population, highlighting the value of tailored protocols to improve early detection and outcomes. Further details are available in the study protocol and online at https://blog.nus.edu.sg/breathe/ .

Humans

Association of the PROC rs146922325 variant with venous thrombosis in a Taiwanese population.

BACKGROUND: Hereditary protein C deficiency, caused by pathogenic variants in the PROC gene, is a known risk factor for venous thrombosis. However, data on PROC variants in Asian populations are limited. This study evaluated the clinical relevance of rs146922325 and its association with thrombotic outcomes in Taiwanese patients. METHODS: Using genotyping data from a single-nucleotide polymorphism array as part of the Taiwan Precision Medicine Initiative, we conducted a retrospective case-control study that included 805 carriers of the PROC rs146922325 variant and 8,050 age- and sex-matched non-carriers. The baseline characteristics, coagulation profiles, and thrombotic outcomes were systematically compared. Univariable and multivariable logistic regression analyses were performed to assess the association between rs146922325 and venous thrombosis. Sensitivity analyses were conducted by restricting the cohort to warfarin-na&#xef;ve participants and incident venous thrombosis events occurring after genotyping. RESULTS: Carriers of the rs146922325 T allele exhibited significantly lower protein C levels than non-carriers (71.28% vs. 114.58%, p&#x2009;<&#x2009;0.001) and a higher prevalence of venous thrombosis (4.10% vs. 2.48%; p&#x2009;=&#x2009;0.009). After multivariable adjustment, rs146922325 carrier status remained independently associated with an increased risk of venous thrombosis (adjusted odds ratio [aOR], 1.61; p&#x2009;=&#x2009;0.015). Allelic analysis further indicated that the T allele was associated with elevated thrombotic risk (aOR, 1.74; p&#x2009;=&#x2009;0.004). No clear dose-response pattern was observed because of the limited number of homozygous TT individuals. Sex-stratified analyses suggested a similar association across sexes; however, the sex&#x2009;&#xd7;&#x2009;genotype interaction was not statistically significant. CONCLUSION: The PROC rs146922325 variant was associated with an increased risk of venous thrombosis in the Taiwanese population. These findings expand the current knowledge of PROC-related thrombophilia in East Asians and support the potential value of genetic risk stratification in thrombosis research.

PROC rs146922325

Host Genetic Factors and Clinical Comorbidities Associated With Tuberculosis Risk.

HLA influence the immune response, shaping genetic susceptibility or resistance to tuberculosis (TB). This study aimed to investigate the associations of host genetics and comorbidities with TB infection in Taiwanese populations. This retrospective case-control study utilised data from the Taiwan Precision Medicine Initiative. TB cases and non-TB controls were compared using genome-wide association studies (GWAS), HLA allele typing, and genotype data. Multivariate logistic regression identified independent predictors of TB and interactions between risk factors. A total of 390 TB cases and 3,909 controls were analysed. Risk factors for TB included bronchiectasis (OR&#x2009;=&#x2009;2.76; 95% CI 1.54-4.44; p&#x2009;<&#x2009;0.001), diabetes mellitus (OR&#x2009;=&#x2009;1.30; 95% CI 1.00-1.68; p&#x2009;=&#x2009;0.050), malignancy (OR&#x2009;=&#x2009;1.46; 95% CI 1.15-1.85; p&#x2009;=&#x2009;0.002), smoking (OR&#x2009;=&#x2009;1.42; 95% CI 1.08-1.88; p&#x2009;=&#x2009;0.012), and steroid use (OR&#x2009;=&#x2009;1.66; 95% CI 1.29-2.13; p&#x2009;<&#x2009;0.001). HLA-DRB1*16:02 was associated with a higher frequency in the TB group (OR&#x2009;=&#x2009;1.47; 95% CI 1.04-2.09; p&#x2009;=&#x2009;0.030). Interaction analysis showed HLA-DRB1*16:02 increased TB risk in non-smokers (OR&#x2009;=&#x2009;1.58; 95% CI 1.02-2.46; p&#x2009;=&#x2009;0.042), but not in smokers. HLA-DRB1*16:02 was associated with a higher risk for TB. While carriers of HLA-DRB1*16:02 did not exhibit an increased risk of TB among smokers, we demonstrated a heightened risk among non-smokers.

Humans

Sparse polygenic risk score inference with the spike-and-slab LASSO.

MOTIVATION: Large-scale biobanks, with rich phenotypic and genomic data across hundreds of thousands of samples, provide ample opportunities to elucidate the genetics of complex traits and diseases. Consequently, there is growing demand for robust and scalable methods for disease risk prediction from genotype data. Inference in this setting is challenging due to the high-dimensionality of genomic data, especially when coupled with smaller sample sizes. Popular Polygenic Risk Score (PRS) inference methods address this challenge by adopting sparse Bayesian priors or penalized regression techniques, such as the Least Absolute Shrinkage and Selection Operator (LASSO). However, the former class of methods are not as scalable and do not produce exact sparsity, while the latter tends to over-shrink large coefficients. RESULTS: In this study, we present SSLPRS, a novel PRS method based on the Spike-and-Slab LASSO (SSL) prior, which offers a theoretical bridge between the two frameworks. We extend previous work to derive a coordinate-ascent inference algorithm that operates on GWAS summary statistics, which is orders-of-magnitude more efficient than corresponding individual-level-based implementations. To illustrate the statistical properties of the proposed model, we conducted experiments involving nine simulation configurations and nine quantitative phenotypes from the UK Biobank. Our results demonstrate that SSLPRS is competitive with state-of-the-art methods in terms of prediction accuracy and exhibits superior variable selection performance, especially in sparse genetic architectures. In simulations, this translates to upwards of 50% improvement in positive predictive value. In analysis of real phenotypes, we show that selected variants are highly enriched for meaningful genomic annotations and have better replication rates in larger meta-analyses. AVAILABILITY AND IMPLEMENTATION: SSLPRS is available in the open-source package https://github.com/li-lab-mcgill/penprs.

Multifactorial Inheritance

Large-Scale Genomic Analysis of Stripe Rust Resistance in Chinese Wheat Germplasm Using Multi-Environment Trial Data.

Wheat stripe rust, caused by Puccinia striiformis f. sp. tritici (Pst), is a significant disease affecting global wheat crops and causing substantial economic losses. This study aimed to identify effective resistance genes by evaluating 120 common wheat accessions from diverse regions in China. These samples were tested with three Pst races at the seedling stage and with natural Pst inoculum at four field locations in three crop seasons. Genotypic data were collected through a Wheat55K iSelect single-nucleotide polymorphism array. The genome-wide association study identified 17 distinct loci linked to stripe rust response, accounting for 1.07 to 30.58% of the phenotypic variation across trials. These loci were distributed among three wheat genome groups: 2 in Group A, 10 in Group B, and 5 in Group D. Among these, eight loci overlapped with the reported stripe rust resistance genes or quantitative trait loci, while nine loci were novel and mainly distributed on chromosomes 2A, 6B, and 7D. This research enhances the understanding of genetic mechanisms underlying wheat stripe rust resistance and provides valuable germplasm resources for breeding new cultivars with enhanced disease resilience.

Puccinia striiformis f. sp. tritici

bioETH-PRS: confidential polygenic risk scoring with smart contracts on an FHE-enabled blockchain.

Polygenic risk scores (PRSs) aggregate genetic effect estimates to predict disease susceptibility, yet calculating one through an external service can require exposing raw genotype data. Homomorphic encryption hides those data during the calculation but, in prior work, still places a designated evaluator in a position of trust. We present bioETH-PRS, a protocol that replaces the evaluator with publicly auditable smart contracts on a blockchain supporting Fully Homomorphic Ethereum Virtual Machine (fhEVM). Using integer-exact encrypted arithmetic, bioETH-PRS computes the PRS dot product entirely in the encrypted domain, so genotype dosages and, at the model provider's discretion, the GWAS weights stay hidden from the parties performing the computation. A fixed-point encoding represents signed weights as nonnegative integers within a bound that rules out overflow, recovering the score to the precision of the published weights. A four-contract architecture separates data custody, model publication, computation, and output release, and supports both a classic path that stores encrypted inputs and an appreciably cheaper streaming path that discards them. A release oracle can return a randomized risk category instead of the raw score, limiting what a repeated querier learns. Prototype evaluation on real GWAS fixtures, including a run on a public testnet, shows cost growing linearly with variant count and suggests the approach may be practical where transaction fees are low. Trust is redistributed rather than removed: the system still depends on the contracts, the blockchain, and the fhEVM services. We evaluate additive models of moderate size, not genome-wide or clinical use.

Blockchain

Genetic evidence for predisposition to acute leukemias due to a missense mutation (p.Ser518Arg) in ZAP70 kinase: a case-control study.

BACKGROUND: The apparent lack of additional missense mutations data on mixed-phenotype leukemia is noteworthy. Single amino acid substitution by these non-synonymous single nucleotide variations can be related to many pathological conditions and may influence susceptibility to disease. This case-control study aimed to unravel whether the ZAP70 missense variant (rs104893674 (C&#x2009;>&#x2009;A)) underpinning mixed-phenotype leukemia. METHODS: The rs104893674 was genotyped in clients who were mixed-phenotype acute leukemia-, acute lymphoblastic leukemia- and acute myeloid leukemia-positive and matched healthy controls, which have been referred to all major urban hospitals from multiple provinces of country- wide, IRAN, from February 11' 2019 to June 10' 2023, by amplification refractory mutation system-polymerase chain reaction method. Direct sequencing for rs104893674 of the ZAP70 gene was performed in a 3130 Genetic Analyzer. RESULTS: We found that the AC genotype of individuals with A allele at this polymorphic site (heterozygous variant-type) contribute to the genetic susceptibility to acute leukemia of both forms, acute myeloid leukemia and acute lymphoblastic leukemia as well as with a mixed phenotype. In other words, the ZAP70 missense variant (rs104893674 (C&#x2009;>&#x2009;A)) increases susceptibility of distinct cell populations of different (myeloid and lymphoid) lineages to exhibiting cancer phenotype. The results were all consistent with genotype data obtained using a direct DNA sequencing technique. CONCLUSION: Of special interest are pathogenic missense mutations, since they generate variants that cause specific molecular phenotypes through protein destabilization. Overall, we discovered that the rs104893674 (C&#x2009;>&#x2009;A) variant chance in causing mixed-phenotype leukemia is relatively high.

Humans