PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “trait imputation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

An autosomal dominant syndrome of hemiplegic migraine, nystagmus, and tremor.

A mother and son suffer from hemiplegic migraine with onset in childhood. Both have nystagmus which has not changed for many years, but the date of onset is uncertain. They have an asymmetrical tremor, clinically indistinguishable from essential tremor. Neuroophthalmological examination revealed inability to produce smooth pursuit, gaze-paretic nystagmus, rebound nystagmus, failure of fixation suppression of the vestibuloocular reflex both horizontally and vertically, and low gain of the optokinetic system. These abnormalities, confirmed by electrooculography, are commonly seen in disease of the cerebellum and brainstem. Treatment with propranolol and pizotyline lessened the number of episodes of hemiplegia and improved the tremor. Hemiplegic migraine has been reported in association with nystagmus, retinal degeneration, deafness, and ataxia in varying combinations in three other families with autosomal dominant inheritance. These associated neurological manifestations likely represent system degenerations rather than the effect of repeated ischemia imputable to the migraine itself. The syndrome of hemiplegic migraine, tremor, and ocular smooth pursuit system disorder seen in this family appears to be inherited as a single autosomal dominant trait, although more than one autosomal dominant gene may be involved.

Adolescent↗

Selphi, a tool for improving genotype imputation accuracy.

Genotype imputation is a powerful tool for inferring missing genotype data in large-scale genetic studies. Over the last two decades, multiple imputation algorithms have been developed, steadily improving in speed and overall accuracy. However, accurate imputation of rare and infrequent variants remains a challenge, largely because existing methods rely on local haplotype matching within genomic windows and do not fully exploit the extended patterns of haplotype sharing that span entire chromosomes. Here we present Selphi, a new genotype imputation algorithm that combines the Positional Burrows-Wheeler Transform (PBWT) with a multi-stage haplotype selection heuristic operating across entire chromosomes. When compared to state-of-the-art methods Beagle 5.4, IMPUTE5, and Minimac4, Selphi showed higher accuracy on the 1000 Genomes Project and TOPMed datasets, across all super-populations and allele frequencies. Similarly, Selphi achieved higher accuracy than Beagle 5.4 on the UK Biobank dataset, which translated into improved concordance with hc-WGS GWAS summary statistics at known trait-associated loci and more accurate polygenic risk scores (PRS). Selphi outputs standard VCF files with genotype dosages (DS), haplotype-specific allele probabilities (AP1, AP2), and a per-variant dosage R-squared quality score (DR2), enabling direct integration with downstream analytical pipelines including standard post-imputation quality filtering.

Genome-Wide Association Study↗

Digital Mindfulness Intervention for Pregnant Women With Affective Disorders and Acute Stress Reactions: Prespecified Secondary Analysis of a Randomized Controlled Trial.

BACKGROUND: Pregnant women with ICD-10 (International Statistical Classification of Diseases, Tenth Revision) affective or stress-related disorders face an elevated risk of perinatal depression and anxiety, yet evidence on digital nonpharmacologic interventions for this population remains limited. OBJECTIVE: This study evaluated the effectiveness of an 8-week digital mindfulness-based intervention (eMBI) compared with treatment as usual (TAU) among pregnant women with ICD-10 affective or stress-related disorders participating in a randomized controlled trial (RCT). METHODS: This prespecified secondary analysis was conducted within a multicenter RCT in Baden-Württemberg, Germany. Pregnant women aged 18 years and older with elevated depressive symptoms (Edinburgh Postnatal Depression Scale [EPDS]>9) and ICD-10-diagnosed affective or stress-related disorders were randomized 1:1 to eMBI or TAU. The intervention consisted of 8 weekly app-based mindfulness sessions (45 min each) delivered during gestational weeks 29-36, with no direct therapist contact. The primary outcome was continuous depressive symptom severity measured with the EPDS at 4-6 weeks post partum. Secondary outcomes included the EPDS at 6 months post partum, generalized anxiety (State-Trait Anxiety Inventory-State [STAI-S], State-Trait Anxiety Inventory-Trait [STAI-T]), and Pregnancy-Related Anxiety Questionnaire-Revised (PRAQ-R). Analyses followed the intention-to-treat (ITT) principle, using mixed models for repeated measures and multiple imputation. RESULTS: Of the 5299 screened women, 147 met the inclusion criteria for this subgroup analysis (intervention group [IG] had n=73 women and control group had n=74 women). Groups were comparable at baseline. The IG showed significantly greater reductions in EPDS scores at gestational week 34 (Δ=-2.21, P=.01), week 36 (Δ=-3.25, P=.01), and 4-6 weeks post partum (Δ=-4.81, P=.007). Treatment effects remained robust under conservative missing-data assumptions. At 4-6 weeks post partum, a higher proportion of participants in the IG achieved clinically meaningful improvement (31/73, 42.5% vs 21/74, 28.4%; adjusted odds ratio 1.56, 95% CI 1.19-2.05; P=.001). Anxiety outcomes followed a similar pattern, whereas pregnancy-related anxiety did not differ between groups. CONCLUSIONS: In this prespecified subgroup of pregnant women with ICD-10 affective or stress-related disorders, the eMBI was associated with clinically meaningful reductions in depressive symptoms from late pregnancy to 4-6 weeks post partum. Effects at 6 months post partum were attenuated and less stable across missing-data assumptions. These findings support eMBIs as a scalable, nonpharmacological adjunct to perinatal mental health care for women with affective or stress-related disorders, while confirmation in adequately powered trials with strategies to reduce postpartum attrition is warranted.

Humans↗

Genetic Analysis of Asymptomatic Antinuclear Antibody Production.

OBJECTIVE: Antinuclear antibodies (ANA) are detected in up to 14% of the population, and many individuals with ANA are asymptomatic. The literature on the genetic contribution to asymptomatic ANA positivity is limited. In this study, we aimed to perform a genome-wide association study of asymptomatic ANA positivity in multiple populations. METHODS: Asymptomatic individuals who were either ANA positive or ANA negative from the All of Us Research Program were included in this study, selecting those with an ANA test performed by immunofluorescence and no evidence of autoimmune disease. Imputation was performed, and a multipopulation meta-analysis including approximately 6 million single-nucleotide polymorphisms (SNPs) was conducted. Genome-wide SNP-based heritability was estimated using the Genome-wide Complex Trait Analysis&#xa0;software. A cumulative genetic risk score for lupus was constructed using previously reported genome-wide significant loci. RESULTS: A total of 1,955 asymptomatic ANA positive and 3,634 asymptomatic ANA negative individuals across three populations were included. The multipopulation meta-analysis revealed SNPs with a suggestive association (P <1 &#xd7; 10-5) across 8 different loci, but no genome-wide significant loci were identified. A gene variant upstream of HLA-DQB1, (rs17211748, P = 1.4 &#xd7; 10-6, odds ratio 0.82, 95% confidence interval 0.76-0.89), showed the most significant association. The heritability of asymptomatic ANA positivity was estimated to be 24.9%. Individuals who were asymptomatic and ANA positive did not exhibit increased cumulative genetic risk for lupus compared with individuals who were ANA negative. CONCLUSION: ANA production is not associated with significant genetic risk and is primarily determined by environmental factors.

Humans↗

Integrative cross-tissue transcriptome-wide association and metabolomic analysis reveals novel genetic risk loci for aortic aneurysm.

BACKGROUND: Aortic aneurysm (AA) is a life-threatening cardiovascular condition with a strong genetic component, however, its molecular mechanisms remain poorly understood. Although genome-wide association studies (GWAS) have identified numerous risk loci, most prior studies have investigated genetic and metabolic factors separately, leaving the causal pathways from genetic variants to disease largely unexplored. METHODS: We established an integrative framework combining cross-tissue transcriptome-wide association studies (TWAS) with metabolomic mediation analysis. First, we integrated GWAS data from FinnGen R12 with multi-tissue expression quantitative trait loci (eQTL) data from Genotype-Tissue Expression Project (GTEx) V8, then performed cross-tissue TWAS using the Unified Test for MOlecular SignaTures (UTMOST) and single-tissue validation with the Functional Summary-based Imputation (FUSION) to prioritize susceptibility genes. Second, we applied Mendelian randomization (MR), colocalization, and Fine-mapping Of CaUsal gene Sets (FOCUS) to assess causality and identify high-confidence genes. Third, we performed metabolite mediation analysis to uncover metabolic pathways linking genetic variants to disease risk. Finally, we validated key findings in mouse models of thoracic aortic aneurysm (TAA) and abdominal aortic aneurysm (AAA) using Quantitative Real-Time Reverse Transcription Polymerase Chain Reaction (RT-qPCR) and Western blotting. RESULTS: We identified multiple novel susceptibility genes for AA and its subtypes. Key genes included ADH family members (ADH1A, ADH1B, ADH4, ADH6) and ZNF827, which showed cross-subtype associations with strong colocalization evidence in vascular tissues. Metabolite mediation analysis revealed significant pathways involving N-acetylphenylalanine and methionine sulfoxide. Functional enrichment revealed distinct biological mechanisms: AA and AAA were primarily associated with metabolic pathways, whereas TAA-related genes were enriched in developmental and contractile processes. PheWAS indicated no significant off-target associations. Critically, experimental validation in mouse models confirmed significant upregulation of ZNF827 in TAA and ADH6 in AAA at both mRNA and protein levels, corroborating the genetic predictions. CONCLUSION: This integrated cross-omics analysis identifies novel genetic loci and, crucially, uncovers specific nutrient-related metabolic pathways that mediate genetic risk. These findings provide a mechanistic basis for future nutritional and metabolic intervention studies in AA and its subtypes.

MAGMA↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗

A method for identifying genes related to a quantitative trait, incorporating multiple siblings and missing parents.

When studying either qualitative or quantitative traits, tests of association in the presence of linkage are necessary for fine-mapping. In a previous report, we suggested a polytomous logistic approach to testing linkage and association between a di-allelic marker and a quantitative trait locus, using genotyped triads, consisting of an individual whose quantitative trait has been measured and his or her two parents. Here we extend that approach to incorporate marker information from entire nuclear families. By computing a weighted score function instead of a maximum likelihood test, we allow for both an unspecified correlation structure between siblings and "informative" family size. Both this approach and our original approach allow for population admixture by conditioning on parental genotypes. The proposed method allows for missing parental genotype data through a multiple imputation procedure. We use simulations based on a population with admixture to compare our method to a popular non-parametric family-based association test (FBAT), testing the null of no association in the presence of linkage.

Alleles↗

Genome-wide linkage analysis of systolic blood pressure slope using the Genetic Analysis Workshop 13 data sets.

Systolic blood pressure (SBP) is an age-dependent complex trait for which both environmental and genetic factors may play a role in explaining variability among individuals. We performed a genome-wide scan of the rate of change in SBP over time on the Framingham Heart Study data and one randomly selected replicate of the simulated data from the Genetic Analysis Workshop 13. We used a variance-component model to carry out linkage analysis and a Markov chain Monte Carlo-based multiple imputation approach to recover missing information. Furthermore, we adopted two selection strategies along with the multiple imputation to deal with subjects taking antihypertensive treatment. The simulated data were used to compare these two strategies, to explore the effectiveness of the multiple imputation in recovering varying degrees of missing information, and its impact on linkage analysis results. For the Framingham data, the marker with the highest LOD score for SBP slope was found on chromosome 7. Interestingly, we found that SBP slopes were not heritable in males but were for females; the marker with the highest LOD score was found on chromosome 18. Using the simulated data, we found that handling treated subjects using the multiple imputation improved the linkage results. We conclude that multiple imputation is a promising approach in recovering missing information in longitudinal genetic studies and hence in improving subsequent linkage analyses.

Age Factors↗

Untangling genetic influences on smoking, body mass index and longevity: a multivariate study of 2464 Danish twins followed for 28 years.

A multivariate twin study was conducted in order to evaluate to what extent smoking, BMI and longevity are influenced by common genetic factors. The study was based on a 28-year follow-up of a sample of 2464 Danish twins who were born in the period 1890-1920 and who answered a questionnaire, including requests for information on smoking status, height and weight, in 1966. By 1994, approximately 2/3 of the sample had died. To compensate for the right-censoring, age at death was imputed for twins who were still alive by using survival analysis; all living subjects were more than 73 years old (mean 80 years, SD 5) in 1994. Proportions of covariance resulting from genetic and environmental factors in common and unique to the three traits were estimated from covariance matrices using the structural equation model approach. The study found no evidence for a substantial impact of common genetic factors on smoking, BMI and longevity. This suggests that only a small fraction of the genetic influences on longevity is mediated via a genetic influence on smoking and BMI and, furthermore, that it is unlikely that the associations between smoking and mortality and between BMI and mortality are confounded by common genetic factors.

Aged↗

Bayesian modelling of multivariate quantitative traits using seemingly unrelated regressions.

We investigate a Bayesian approach to modelling the statistical association between markers at multiple loci and multivariate quantitative traits. In particular, we describe the use of Bayesian Seemingly Unrelated Regressions (SUR) whereby genotypes at the different loci are allowed to have non-simultaneous effects on the phenotypes considered with residuals from each regression assumed correlated. We present results from simulations showing that, under rather general conditions that are likely to hold in real situations, the Bayesian SUR approach has increased probability of selecting the true model compared to univariate analyses. Finally, we apply our methods to data from subjects genotyped for 12 SNPs in the apolipoprotein E (APOE) gene. Phenotypes relate to response to treatment with atorvastatin and include changes in total cholesterol, low-density lipoprotein cholesterol, and triglycerides. Missing genotype data are naturally accommodated in our Bayesian framework by imputing them using a nested haplotype phasing algorithm.

Algorithms↗

Bayesian methods for quantitative trait loci mapping based on model selection: approximate analysis using the Bayesian information criterion.

We describe an approximate method for the analysis of quantitative trait loci (QTL) based on model selection from multiple regression models with trait values regressed on marker genotypes, using a modification of the easily calculated Bayesian information criterion to estimate the posterior probability of models with various subsets of markers as variables. The BIC-delta criterion, with the parameter delta increasing the penalty for additional variables in a model, is further modified to incorporate prior information, and missing values are handled by multiple imputation. Marginal probabilities for model sizes are calculated, and the posterior probability of nonzero model size is interpreted as the posterior probability of existence of a QTL linked to one or more markers. The method is demonstrated on analysis of associations between wood density and markers on two linkage groups in Pinus radiata. Selection bias, which is the bias that results from using the same data to both select the variables in a model and estimate the coefficients, is shown to be a problem for commonly used non-Bayesian methods for QTL mapping, which do not average over alternative possible models that are consistent with the data.

Alleles↗

On locating multiple interacting quantitative trait loci in intercross designs.

A modified version (mBIC) of the Bayesian Information Criterion (BIC) has been previously proposed for backcross designs to locate multiple interacting quantitative trait loci. In this article, we extend the method to intercross designs. We also propose two modifications of the mBIC. First we investigate a two-stage procedure in the spirit of empirical Bayes methods involving an adaptive (i.e., data-based) choice of the penalty. The purpose of the second modification is to increase the power of detecting epistasis effects at loci where main effects have already been detected. We investigate the proposed methods by computer simulations under a wide range of realistic genetic models, with nonequidistant marker spacings and missing data. In the case of large intermarker distances we use imputations according to Haley and Knott regression to reduce the distance between searched positions to not more than 10 cM. Haley and Knott regression is also used to handle missing data. The simulation study as well as real data analyses demonstrates good properties of the proposed method of QTL detection.

Algorithms↗

Adjusting for treatment effects in studies of quantitative traits: antihypertensive therapy and systolic blood pressure.

A population-based study of a quantitative trait may be seriously compromised when the trait is subject to the effects of a treatment. For example, in a typical study of quantitative blood pressure (BP) 15 per cent or more of middle-aged subjects may take antihypertensive treatment. Without appropriate correction, this can lead to substantial shrinkage in the estimated effect of aetiological determinants of scientific interest and a marked reduction in statistical power. Correction relies upon imputation, in treated subjects, of the underlying BP from the observed BP having invoked one or more assumptions about the bioclinical setting. There is a range of different assumptions that may be made, and a number of different analytical models that may be used. In this paper, we motivate an approach based on a censored normal regression model and compare it with a range of other methods that are currently used or advocated. We compare these methods in simulated data sets and assess the estimation bias and the loss of power that ensue when treatment effects are not appropriately addressed. We also apply the same methods to real data and demonstrate a pattern of behaviour that is consistent with that in the simulation studies. Although all approaches to analysis are necessarily approximations, we conclude that two of the adjustment methods appear to perform well across a range of realistic settings. These are: (1) the addition of a sensible constant to the observed BP in treated subjects; and (2) the censored normal regression model. A third, non-parametric, method based on averaging ordered residuals may also be advocated in some settings. On the other hand, three approaches that are used relatively commonly are fundamentally flawed and should not be used at all. These are: (i) ignoring the problem altogether and analysing observed BP in treated subjects as if it was underlying BP; (ii) fitting a conventional regression model with treatment as a binary covariate; and (iii) excluding treated subjects from the analysis. Given that the more effective methods are straightforward to implement, there is no argument for undertaking a flawed analysis that wastes power and results in excessive bias.

Aged↗

Revealing the Shared Genetic Architecture of Metabolic Dysfunction-Associated Steatotic Liver Disease-Related Traits Through Genomic Structural Equation Modeling.

Although individual traits related to metabolic dysfunction-associated steatotic liver disease (MASLD) have been investigated through large-scale genome-wide association studies (GWASs), the shared genetic susceptibility across these traits remains unclear. We therefore conducted a multivariate GWAS of key MASLD-related traits to elucidate their common genetic architecture. We applied genomic structural equation modeling to model a latent genetic factor (MASLD-F) underlying genetically correlated MASLD-related traits, leveraging their GWAS-derived genetic correlations. We then performed functional annotations, including fine-mapping, transcriptome-wide association study, and cell- and tissue-type-specific enrichment analyses, and conducted Mendelian randomization analyses to identify modifiable risk factors. Our multivariate MASLD-F GWAS identified 50 independent variants across 48 genomic loci. Transcriptomic imputation identified several MASLD-F-associated genes, including ARNTL, NPC1, BTBD10, VDAC2, TSKU, SFMBT1, and ABHD17C. We observed significant enrichment of MASLD-F-related genetic signals predominantly in brain tissues, pancreatic islets, and the adrenal gland. Additionally, six modifiable risk factors and four modifiable protective factors for MASLD-F were identified. These findings reveal a complex shared genetic architecture underlying MASLD components, thereby expanding our understanding of disease pathogenesis and providing novel insights for precision medicine and public health interventions.

Humans↗

Summary report: Missing data and pedigree and genotyping errors.

Genetic epidemiology is faced with mapping complex traits to genes with relatively small effects whose phenotypes may be modulated by temporal factors. To do this, detailed and accurate data must be available on families, perhaps collected over time. The Framingham Heart Study data supplied to Genetic Analysis Workshop 13 (GAW13), along with its simulated counterpart, contain longitudinal measurements and genomic scan data on 2,885 individuals in 330 families, and offer an opportunity to examine data quality and completeness issues as they affect analytical conclusions. Six GAW13 contributions applied methods to deal with missing data, both phenotypic and genotypic, at a single time point and longitudinally, and with possible errors in pedigree structure and genotypes. The methods included missing phenotypic data imputation by Markov chain Monte Carlo sampling, propensity scoring, regression, and adjusted mean values, as well as the assessment of transmission-disequilibrium tests when missing marker data may be allele-specific. Pedigree structural errors were found by genome-wide allele-sharing probabilities, while Mendelian consistent genotype errors were evaluated through likelihoods of double-recombination events. Each of the methods reviewed here offered insights into how to better take advantage of large, time-dependent, familial data sets. However, no one of them dealt with the longitudinal and familial aspects simultaneously. Overall, more consideration needs to be given to the effects that missing data and data errors have on our ability to map complex traits efficiently and accurately.

Cardiovascular Diseases↗

New Genetic Loci Implicated in Cardiac Morphology and Function Using Three-Dimensional Population Phenotyping.

BACKGROUND: Cardiac remodeling occurs in the mature heart and is a cascade of adaptations in response to stress, which are primed in early life. A key question remains as to the processes that regulate the geometry and motion of the heart and how it adapts to stress. METHODS: We performed spatially resolved phenotyping using machine learning-based analysis of cardiac magnetic resonance imaging in 47&#x2009;549 UK Biobank participants. We analyzed 16 left ventricular spatial phenotypes, including regional myocardial wall thickness and systolic strain in both circumferential and radial directions. In up to 40&#x2009;058 participants, genetic associations across the allele frequency spectrum were assessed using genome-wide association studies with imputed genotype participants, and exome-wide association studies and gene-based burden tests using whole-exome sequencing data. We integrated transcriptomic data from the GTEx project and used pathway enrichment analyses to further interpret the biological relevance of identified loci. To investigate causal relationships, we conducted Mendelian randomization analyses to evaluate the effects of blood pressure on regional cardiac traits and the effects of these traits on cardiomyopathy risk. RESULTS: We found 42 loci associated with cardiac structure and contractility, many of which reveal patterns of spatial organization in the heart. Whole-exome sequencing revealed 3 additional variants not captured by the genome-wide association study, including a missense variant in CSRP3 (minor allele frequency 0.5%). The majority of newly discovered loci are found in cardiomyopathy-associated genes, suggesting that they regulate spatially distinct patterns of remodeling in the left ventricle in an adult population. Our causal analysis also found regional modulation of blood pressure on cardiac wall thickness and strain. CONCLUSIONS: These findings provide a comprehensive description of the pathways that orchestrate heart development and cardiac remodeling. These data highlight the role that cardiomyopathy-associated genes have on the regulation of spatial adaptations in those without known disease.

Humans↗

Variant harmonization critically determines polygenic score transferability for lipid traits in Samoan populations.

Dyslipidemia is a significant risk factor for cardiovascular disease (CVD), the leading cause of death in Samoa. Polygenic scores (PGSs) for lipid traits offer promise for improved CVD risk prediction; however, their performance in Pacific Islander populations-comprising only 0.002% of genome-wide association study (GWAS) participants as of 2024-remains unknown. We evaluated the transferability of multi-ancestry PGS for LDL cholesterol (LDL-C), HDL cholesterol (HDL-C), triglycerides (TGs), and total cholesterol (TC) in 4,342 Samoan adults across five cohorts spanning 1990-2010. PGSs from Graham et al. and Kanoni et al. multi-ancestry meta-analyses were harmonized with genome-wide imputed genotypes using a Samoan-specific reference panel, and performance was assessed via incremental R2 from linear mixed models with bootstrapped confidence intervals. HDL-C showed the highest performance (incremental R2 5.0%-15.0%), followed by TC (5.0%-10.7%), LDL-C (5.7%-8.6%), and TG (3.5%-7.0%). Critically, meaningful LDL-C performance was achieved only with the genome-wide PRS-CS score (99.6%-99.7% variant matching), while a curated pruning-and-thresholding score achieved &#x223c;9% matching and near-zero performance. These findings establish systematic lipid PGS benchmarks in Samoans, demonstrating meaningful transferability when genome-wide variant coverage is ensured, and highlight variant harmonization as a critical precondition for PGS deployment in underrepresented populations.

Pacific Islanders↗

Multi-omics analysis identifies key genes and functional loci affecting teat number in American Large White and Landrace pigs and their application in optimizing genomic selection models.

BACKGROUND: Teat number is a crucial economic trait in pigs. It directly affects the ability of sows to lactate, which in turn influences the survival and health of piglets. The teat number of French Large White pigs is close to 16, while the teat number of American Large White and Landrace pigs is about 14. In order to improve the teat number of American Landrace and Large White pigs through molecular approaches and precise breeding techniques, we genotyped 2,131 American Landrace and 4,564 American Large White with teat number phenotype using a 50&#xa0;K SNP chip. Then, the SNP-chip data was imputed to the level of whole-genome sequencing (iWGS). Based on iWGS data, we conducted GWAS to identify novel, significant SNPs associated with teat number and to incorporate them into genomic selection. RESULTS: In Landrace pigs, significant SNPs for TTN mapped to SSC2, SSC7, SSC8, and SSC14; the SSC8 and SSC14 effects are novel. LTN mapped to SSC7, RTN to SSC7 and SSC8. The lead SSC7 SNP explained 2.60% of TTN phenotypic variance. In Large White pigs, significant SNPs were detected on SSC7 and SSC10 for TTN; SSC7, SSC10, and SSC12 for LTN; and SSC7 and SSC10 for RTN. The most significant locus on SSC7 accounted for 2.99% of the phenotypic variance in TTN. Additionally, a multi-population meta-analysis detected significant novel SNPs for LTN on SSC1 and SSC8. By utilizing Bayesian fine mapping, the most precise QTL confidence interval on SSC7 for both TTN and RTN in Large White pigs was reduced to 40&#xa0;kb. By integrating functional gene annotation with RNA-seq and ATAC-seq data from Erhualian and Bamaxiang pigs mammary placodes at embryonic day 26, we prioritized PTPN13, TRPV3, ZDHHC13, and BRD2 as novel candidate genes for teat number. We then incorporated the significant SNPs to GBLUP and benchmarked genomic-selection accuracy. In both breeds, fitting the top SNP as fixed maximized prediction for TTN and RTN, whereas treating all significant loci as an additional random effect optimized LTN. CONCLUSIONS: Our findings provide a theoretical basis for dissecting new key genes affecting teat number and for advancing molecular breeding of teat number in pigs.

Animals↗