PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “trait imputation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross‐taxon transferability↗

Gene expression in mental illness: a navigation chart to future progress.

An initial course in disentangling complex causal interactions in psychiatric illnesses, we suggest, is finding co-familial traits with classical Mendelian segregation. Starting with non-Mendelian traits, three methods can be used to find underlying Mendelian phenotypes. (1) Statistically-inferred latent traits, with more nearly Mendelian transmission than the measures from which they are derived, can serve as pointers to concrete Mendelian phenotypes. (2) Linkage of non-Mendelian traits to genetic markers, if it can be established, can be followed by searching for phenotypes that discriminate carriers from non-carriers of the imputed trait gene. (3) In the long run, the most successful method is likely to be direct refinement of non-Mendelian behavioral and physiological traits into more fundamental components.

Bipolar Disorder↗

CYClones: a highly powered, fully genotyped, eight-parent yeast mapping population.

The budding yeast Saccharomyces cerevisiae is a remarkably adaptable organism that thrives in diverse environments. Global sequencing of natural isolates has revealed extensive genetic diversity within the species. Here, we describe the construction and characterization of CYClones (Collaborative Yeast Cross clones), a library of 11,392 segregants generated from a multiparent funnel cross of eight genetically diverse parental strains. To enable the genetic dissection of complex traits, we imputed whole-genome sequences for all segregants and show that CYClones captures a substantial fraction of the global genetic diversity of S. cerevisiae. Haplotype representation is well maintained, with each parental haplotype present at >5% frequency across >95% of the genome. Simulations demonstrate that CYClones has ≥95% power to detect variants with heritability as low as 0.36%, with mapping resolution often finer than the length of a single gene. In summary, CYClones is a powerful community resource for dissecting the genetic architecture of complex and quantitative traits, uncovering context-dependent mutational effects, and identifying causal variants underlying phenotypic diversity.

Saccharomyces cerevisiae↗

Association testing with Mendel.

This report presents an overview of association testing strategies from a user's perspective, with particular attention to the capabilities of the computer program Mendel. Association testing is driven by the nature of the study sample, the nature of the disease trait, and the kind of markers employed. The practicing statistician must also choose whether to conduct parametric or nonparametric tests. Because of the complexities involved, Mendel offers users several analysis options. The different options are tied together by shared input and output conventions and a shared language for defining models. Mendel also features new statistics and theory found in no other genetics software. The most important innovations include: association testing by penetrance estimation, expansion of matched-pair designs to permutation unit designs, and a rigorous implementation of the measured genotype approach for quantitative trait loci. This report explains how Mendel imputes allele counts and conducts both asymptotic and permutation tests in the measured genotype framework.

Analysis of Variance↗

Penalised regression improves imputation of cell-type specific expression using RNA-seq data from mixed cell populations compared to domain-specific methods.

Gene expression studies often use bulk RNA sequencing of mixed cell populations because single cell or sorted cell sequencing may be prohibitively expensive. However, mixed cell studies may miss expression patterns that are restricted to specific cell populations. Computational deconvolution can be used to estimate cell fractions from bulk expression data and infer average cell-type expression in a set of samples (e.g., cases or controls), but imputing sample-level cell-type expression is required for more detailed analyses, such as relating expression to quantitative traits, and is less commonly addressed. Here, we assessed the accuracy of imputing sample-level cell-type expression using a real dataset where mixed peripheral blood mononuclear cells (PBMC) and sorted (CD4, CD8, CD14, CD19) RNA sequencing data were generated from the same subjects (N=158), and pseudobulk datasets synthesised from eQTLgen single cell RNA-seq data. We compared three domain-specific methods, CIBERSORTx, bMIND and debCAM/swCAM, and two cross-domain machine learning methods, multiple response LASSO and ridge, that had not been used for this task before. We also assessed the methods according to their ability to recover differential gene expression (DGE) results. LASSO/ridge showed higher sensitivity but lower specificity for recovering DGE signals seen in observed data compared to deconvolution methods, although LASSO/ridge had higher area under curves than deconvolution methods. Machine learning methods have the potential to outperform domain-specific methods when suitable training data are available.

Humans↗

Anthropometric and cardio-metabolic trait variation and genetic associations in sub-Saharan Africa.

The genetics of complex traits in Africa has been historically understudied, which can contribute to healthcare inequalities. Here, we present observations of 27 anthropometric, cardiovascular, and blood biomarker measurements across 2,124 individuals from sub-Saharan Africa for whom we also have dense genotype data. First, we identified trait values that differ significantly across populations and subsistence lifestyles (e.g., hemoglobin levels and height). We then identified traits with high degrees of sexual dimorphism (e.g., weight and grip strength). ADMIXTURE analyses revealed substantial population structure in our dataset, and many of the phenotypes studied here are correlated with genetic ancestry components, particularly skin color and body size traits. A variance partitioning approach further revealed traits in which much of the SNP heritability is due to polymorphisms that also contribute to differences between ancestry components. Following genomic imputation, we performed genome-wide association studies (GWASs) for all 27 traits and identified >100 independent autosomal SNPs with genome-wide significant associations for at least one trait (p < 5 &#xd7; 10-8). Many of these trait-associated variants are rare outside of Africa (minor-allele frequency [MAF] < 1%). We found that 100 kb windows surrounding the top GWAS hits from our African-ancestry cohort were enriched for trait associations in an identically sized European cohort and vice versa. We performed a more detailed analysis of height prediction from genetic data, finding that genome-wide admixture proportions predict height in Africans better than polygenic predictors based on large-scale European height GWASs.

Female↗

OmicsPred as a centralised resource for genetic prediction of multi-omic traits.

Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.

Journal Article↗

A novel reusable transcriptome-wide association study workflow used to map key genes linked to important cattle traits.

Transcriptome-wide association studies (TWAS) are a powerful approach for studying the genes underlying complex traits by directly integrating GWAS and gene expression datasets. In cattle, they have been previously applied to identify genes driving fertility, milk production, and health. However, these studies have also highlighted several challenges, from difficulties in reproducing these complex analyses to limitations from poor genotype calls, especially when called directly from RNA sequencing data. To address these and other challenges, for the H2020 BovReg Project, we have developed a streamlined, species-agnostic, and reusable Nextflow TWAS workflow to integrate transcriptomic and GWAS summary statistic datasets. Our workflow first generates accurate genotype calls and gene expression prediction models from transcriptomic datasets and then applies these tools to impute gene expression levels into GWAS cohorts, enabling the association of genes with traits of interest. We explore optimal strategies for calling genetic variants directly from transcriptomic data and illustrate that using imputation approaches specifically designed for low-pass sequencing data can improve variant calling over previously adopted methods. We demonstrate the utility of our TWAS workflow by applying it to both novel and publicly available GWAS cohorts for cattle, detecting novel gene-trait associations for complex traits. Using a new transcriptome annotation of the cattle genome generated for the BovReg project we also illustrate how previously un-assayable associations can be detected. The results and the workflow we present, provide a new resource for the community and contribute to a better understanding of the molecular drivers of complex traits in cattle with the goal of eventually leveraging this information in future breeding decisions.

Animals↗

Comparison of missing data approaches in linkage analysis.

BACKGROUND: Observational cohort studies have been little used in linkage analyses due to their general lack of large, disease-specific pedigrees. Nevertheless, the longitudinal nature of such studies makes them potentially valuable for assessing the linkage between genotypes and temporal trends in phenotypes. The repeated phenotype measures in cohort studies (i.e., across time), however, can have extensive missing information. Existing methods for handling missing data in observational studies may decrease efficiency, introduce biases, and give spurious results. The impact of such methods when undertaking linkage analysis of cohort studies is unclear. Therefore, we compare here six methods of imputing missing repeated phenotypes on results from genome-wide linkage analyses of four quantitative traits from the Framingham Heart Study cohort. RESULTS: We found that simply deleting observations with missing values gave many more nominally statistically significant linkages than the other five approaches. Among the latter, those with similar underlying methodology (i.e., imputation- versus model-based) gave the most consistent results, although some discrepancies remained. CONCLUSION: Different methods for addressing missing values in linkage analyses of cohort studies can give substantially diverse results, and must be carefully considered to protect against biases and spurious findings.

Algorithms↗

Genome-wide association studies for feed efficiency, production and feeding behavior traits in Canadian purebred Duroc pigs.

This study aimed to identify potential genetic variants and candidate genes associated with feed efficiency (FE), production, and feeding behavior traits in Canadian purebred Duroc pigs. Genome-wide association studies (GWAS) were conducted using 8,861 individuals and an imputed Affymetrix PigGen Canada 50K panel v2.0 using a linear mixed model (LMM) and a Bayesian B model. This analysis used an adjusted P-value threshold (ranging from 6.6&#x202f;&#xd7;&#x202f;10-5 to 1.3&#x202f;&#xd7;&#x202f;10-4) using a false-discovery rate to determine significance. The number of significant SNPs identified for each trait was as follows: average daily gain (ADG, 48), daily feed intake (DFI, 85), feed conversion ratio (FCR, 101), residual feed intake (RFI, 37), residual gain (RG, 64), residual intake and gain (RIG, 55), backfat thickness (BF, 100), loin depth (LD, 6), Kleiber's ratio (KR, 0), total time spent eating per day (TPD, 7), and number of visits to the feeder per day (NVD, 6). Several traits (BF, DFI, FCR, RFI, RG, and RIG) showed strong overlapping signals on chromosomes 7 and 10 with 24 shared significant SNPs, indicating potential shared genetic mechanisms. These traits also had 71 overlapping candidate genes, such as PACSIN1, PTCH1, ADIPOR1, and ITPR3, associated with glucose, lipid, and cholesterol metabolism. Well-known candidate genes in literature associated with growth and fatness such as MC4R and CDH20 were also identified to be associated with ADG, BF, FCR, and DFI in this study. Gene ontology enrichment analysis revealed that a set of the candidate genes were involved in the gonadotropin-releasing hormone (GnRH) and the platelet-derived growth factor (PDGF) signaling pathways. Overall, this study contributed to understanding the genetic architecture and provided a biological foundation for improving FE, production, and feeding behavior traits in Canadian Duroc pigs, facilitating the selection of more efficient pigs.

Sus scrofa↗

Genome sharing in large pedigrees: multiple imputation of ibd for linkage detection.

Our objective is the development of robust methods for assessment of evidence for linkage of loci affecting a complex trait to a marker linkage group, using data on extended pedigrees. Using Markov chain Monte Carlo (MCMC) methods, it is possible to sample realizations from the distribution of gene identity by descent (IBD) patterns on a pedigree, conditional on observed data YM at multiple marker loci. Measures of gene IBDW which capture joint genome sharing in extended pedigrees often have unknown and highly skewed distributions, particularly when conditioned on marker data. MCMC provides a direct estimate of the distribution of such measures. Let W be the IBD measure from data YM, and W* the IBD measure from pseudo-data Y*M simulated with the same data availability and genetic marker model as the true data YM, but in the absence of linkage. Then measures of the difference in distributions of W and W* provide evidence for linkage. This approach extracts more information from the data YM than either comparison to the pedigree prior distribution of W or use of statistics that are expectations of W given the data YM. A small example is presented.

Chromosome Mapping↗

Genetic Architecture of Idiopathic Inflammatory Myopathies From Meta-Analyses.

OBJECTIVE: Idiopathic inflammatory myopathies (IIMs, myositis) are rare systemic autoimmune disorders that lead to muscle inflammation, weakness, and extramuscular manifestations, with a strong genetic component influencing disease development and progression. Previous genome-wide association studies identified loci associated with IIMs. In this study, we imputed data from two prior genome-wide myositis studies and analyzed the largest myositis data set to date to identify novel risk loci and susceptibility genes associated with IIMs and its clinical subtypes. METHODS: We performed association analyses on 14,903 individuals (3,206 patients and 11,697 controls) with genotypes and imputed data from the Trans-Omics for Precision Medicine reference panel. Fine-mapping and expression quantitative trait locus colocalization analyses in myositis-relevant tissues indicated potential causal variants. Functional annotation and network analyses using the random walk with restart (RWR) algorithm explored underlying genetic networks and drug repurposing opportunities. RESULTS: Our analyses identified novel risk loci and susceptibility genes, such as FCRLA, NFKB1, IRF4, DCAKD, and ATXN2 in overall IIMs; NEMP2 in polymyositis; ACBC11 in dermatomyositis; and PSD3 in myositis with anti-histidyl-transfer RNA synthetase autoantibodies (anti-Jo-1). We also characterized effects of HLA region variants and the role of C4. Colocalization analyses suggested putative causal variants in DCAKD in skin and muscle, HCP5 in lung, and IRF4 in Epstein-Barr virus (EBV)-transformed lymphocytes, lung, and whole blood. RWR further prioritized additional candidate genes, including APP, CD74, CIITA, NR1H4, and TXNIP, for future investigation. CONCLUSION: Our study uncovers novel genetic regions contributing to IIMs, advancing our understanding of myositis pathogenesis and offering new insights for future research.

Humans↗

Conditional multipoint linkage analysis using affected sib pairs: an alternative approach.

Recently, Liang et al. ([2001b] Genet. Epidemiol. 21:105-122) proposed a conditional approach to assess linkage evidence on the target region by incorporating linkage information from an unlinked (reference) region using allele shared IBD (identity-by-decent) from affected sib pairs. This is carried out by conditioning on the IBD sharing value at the estimated trait locus of the reference region. Since markers considered are typically non-fully informative, the IBD sharing at each marker needs to be estimated (or imputed). In this report, we propose an alternative approach to deal with the IBD sharing in the reference region. This new approach makes full use of the observed data without having to categorize the imputed IBD sharing as needed in Liang et al. ([2001b] Genet. Epidemiol. 21:105-122). We compare these two approaches by simulating data from a variety of two-locus models including heterogeneity, additive and multiplicative with either fully informative markers or non-fully informative markers. The performance of both approaches is quite comparable showing consistent estimates of the trait locus and key genetic parameters.

Alleles↗

LungGENIE: the lung gene-expression and network imputation engine.

BACKGROUND: Few cohorts have study populations large enough to conduct molecular analysis of ex vivo lung tissue for genomic analyses. Transcriptome imputation is a non-invasive alternative with many potential applications. We present a novel transcriptome-imputation method called the Lung Gene Expression and Network Imputation Engine (LungGENIE) that uses principal components from blood gene-expression levels in a linear regression model to predict lung tissue-specific gene-expression. METHODS: We use paired blood and lung RNA sequencing data from the Genotype-Tissue Expression (GTEx) project to train LungGENIE models. We replicate model performance in a unique dataset, where we generated RNA sequencing data from paired lung and blood samples available through the SUNY Upstate Biorepository (SUBR). We further demonstrate proof-of-concept application of LungGENIE models in an independent blood RNA sequencing data from the Genetic Epidemiology of COPD (COPDGene) study. RESULTS: We show that LungGENIE prediction accuracies have higher correlation to measured lung tissue expression compared to existing cis-expression quantitative trait loci-based methods (median Pearson's r&#x2009;=&#x2009;0.25, IQR 0.19-0.32), with close to half of the reliably predicted transcripts being replicated in the testing dataset. Finally, we demonstrate significant correlation of differential expression results in chronic obstructive pulmonary disease (COPD) from imputed lung tissue gene-expression and differential expression results experimentally determined from lung tissue. CONCLUSION: Our results demonstrate that LungGENIE provides complementary results to existing expression quantitative trait loci-based methods and outperforms direct blood to lung results across internal cross-validation, external replication, and proof-of-concept in an independent dataset. Taken together, we establish LungGENIE as a tool with many potential applications in the study of lung diseases.

Humans↗

Imputation methods for missing data for polygenic models.

Methods to handle missing data have been an area of statistical research for many years. Little has been done within the context of pedigree analysis. In this paper we present two methods for imputing missing data for polygenic models using family data. The imputation schemes take into account familial relationships and use the observed familial information for the imputation. A traditional multiple imputation approach and multiple imputation or data augmentation approach within a Gibbs sampler for the handling of missing data for a polygenic model are presented.We used both the Genetic Analysis Workshop 13 simulated missing phenotype and the complete phenotype data sets as the means to illustrate the two methods. We looked at the phenotypic trait systolic blood pressure and the covariate gender at time point 11 (1970) for Cohort 1 and time point 1 (1971) for Cohort 2. Comparing the results for three replicates of complete and missing data incorporating multiple imputation, we find that multiple imputation via a Gibbs sampler produces more accurate results. Thus, we recommend the Gibbs sampler for imputation purposes because of the ease with which it can be extended to more complicated models, the consistency of the results, and the accountability of the variation due to imputation.

Adult Children↗

Meta-analysis of over 8,000 individuals from Hawai'i and Samoa for genetic associations to cardiometabolic phenotypes.

Although genome-wide association studies (GWAS) now routinely reveal genetic associations and biological insights in millions of individuals, underrepresentation of global populations, such as those from Polynesia, continue to persist. These exclusions, often driven by logistical challenges and lack of data, prevent systematic identification of population-enriched associations, such as the association of the missense variant at the CREBRF locus to BMI and type 2 diabetes discovered commonly occurring in Polynesian populations due to its rarity in global populations. Armed with the recently updated TOPMed imputation panel that could benefit studies in diverse populations that previously had poorer imputation performance, we performed the first GWAS of Native Hawaiians and largest to date of Polynesian-ancestry populations (combined N up to 8,461) to identify population-enriched associations for 13 adiposity and cardiometabolic traits available across both cohorts: BMI, fasting glucose, fasting insulin, HDL, height, hip circumference, HOMA-IR, LDL, T2D, total cholesterol, triglycerides, waist circumference, and waist-hip ratio. We found 25 trait-loci associations that met genome-wide significance: 20 previously reported or known associations and 5 associations newly confirmed via meta-analysis. In particular, with improved statistical power, we were able to confirm the suspected association between the missense CREBRF variant with fasting glucose levels. The remaining 4 potentially novel loci-trait associations for BMI, LDL, and waist-hip ratio, however, were not replicated in multi-ethnic datasets from All-of-Us despite having reasonable power to replicate. The lack of Polynesian-enriched findings outside of the CREBRF locus informs the bounds of the effect sizes or frequency of any enriched variants, and suggests that further expansion of cohort sizes from this region of the world and improved imputation references specific to these populations are needed to identify more population-enriched associations.

Journal Article↗

The value of relatives with phenotypes but missing genotypes in association studies for quantitative traits.

The additional statistical power of association studies for quantitative traits was derived when ungenotyped relatives with phenotypes are included in the analysis. It was shown that the extra power is a simple function of the coefficient of additive genetic relationship and the phenotypic correlation coefficient between the genotyped and ungenotyped relatives. For close relatives, such as pairs of fullsibs and identical twin pairs, gains in power in the range of 10 to 30% are achieved if only one of the pair is genotyped. The theoretical results were verified by simulations. It was shown that ignoring the error in estimating the genotype of the ungenotyped relative has little impact on the estimates and on statistical power, consistent with results from quantitative trait loci (QTL) linkage studies. For genome-wide association studies in which not all relatives with phenotypes can be genotyped, our study provides a prediction of the additional power of an analysis that includes phenotypes on ungenotyped individuals, and can be used in experimental design. We show that a two-step procedure, in which missing genotypes are imputed and subsequently an association analysis is performed, is efficient and powerful.

Genetic Predisposition to Disease↗

On normality, ethnicity, and missing values in quantitative trait locus mapping.

BACKGROUND: This paper deals with the detection of significant linkage for quantitative traits using a variance components approach. Microsatellite markers were obtained for the Genetic Analysis Workshop 14 Collaborative Study on the Genetics of Alcoholism data. Ethnic heterogeneity, highly skewed quantitative measures, and a high rate of missing values are all present in this dataset and well known to impact upon linkage analysis. This makes it a good candidate for investigation. RESULTS: As expected, we observed a number of changes in LOD scores, especially for chromosomes 1, 7, and 18, along with the three factors studied. A dramatic example of such changes can be found in chromosome 7. Highly significant linkage to one of the quantitative traits became insignificant when a proper normalizing transformation of the trait was used and when analysis was carried out on an ethnically homogeneous subset of the original pedigrees. CONCLUSION: In agreement with existing literature, transforming a trait to ensure normality using a Box-Cox transformation is highly recommended in order to avoid false-positive linkages. Furthermore, pedigrees should be sorted by ethnic groups and analyses should be carried out separately. Finally, one should be aware that the inclusion of covariates with a high rate of missing values reduces considerably the number of subjects included in the model. In such a case, the loss in power may be large. Imputation methods are then recommended.

Chromosome Mapping↗