PubMed Health⌕ Search

Biomedical subjects

Xihong Lin

Publications and source records attributed to Xihong Lin.

13 recordsLinked to original sources

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans↗

Whole genome sequence analysis of low-density lipoprotein cholesterol across 246 K individuals.

BACKGROUND: Rare genetic variation provided by whole genome sequence datasets has been relatively less explored for its contributions to human traits. Meta-analysis of sequencing data offers advantages by integrating larger sample sizes from diverse cohorts, thereby increasing the likelihood of discovering novel insights into complex traits. Furthermore, emerging methods in genome-wide rare variant association testing further improve power and interpretability. RESULTS: Here, we conduct the largest meta-analysis of whole genome sequencing for low-density lipoprotein cholesterol (LDL-C), a therapeutic target for coronary artery disease, analyzing data from 246 K participants and integrating 1.23B variants from the UK Biobank and the Trans-Omics for Precision Medicine (TOPMed) program. We identify numerous rare coding and non-coding gene associations related to LDL-C, with replication across 86 K participants in All of Us. Our findings are based on single-variant analyses, rare coding and non-coding variant aggregation tests, and sliding window approaches. Through this comprehensive analysis, we identify 704 novel single-variant associations, 25 novel rare coding variant aggregates, 28 novel rare non-coding variant aggregates, and one novel sliding window aggregate. CONCLUSIONS: This study provides a meta-analysis framework for large-scale whole genome sequence association analyses from diverse population groups, yielding novel rare non-coding variant associations.

Humans↗

Causal Mediation Analysis for Integrating Exposure, Genomic, and Phenotype Data.

Causal mediation analysis provides an attractive framework for integrating diverse types of exposure, genomic, and phenotype data. Recently, this field has seen a surge of interest, largely driven by the increasing need for causal mediation analyses in health and social sciences. This article aims to provide a review of recent developments in mediation analysis, encompassing mediation analysis of a single mediator and a large number of mediators, as well as mediation analysis with multiple exposures and mediators. Our review focuses on the recent advancements in statistical inference for causal mediation analysis, especially in the context of high-dimensional mediation analysis. We delve into the complexities of testing mediation effects, especially addressing the challenge of testing a large number of composite null hypotheses. Through extensive simulation studies, we compare the existing methods across a range of scenarios. We also include an analysis of data from the Normative Aging Study, which examines DNA methylation CpG sites as potential mediators of the effect of smoking status on lung function. We discuss the pros and cons of these methods and future research directions.

causal inference↗

Whole-genome sequencing in 333,100 individuals reveals rare non-coding single variant and aggregate associations with height.

The role of rare non-coding variation in complex human phenotypes is still largely unknown. To elucidate the impact of rare variants in regulatory elements, we performed a whole-genome sequencing association analysis for height using 333,100 individuals from three datasets: UK Biobank (N&#x2009;=&#x2009;200,003), TOPMed (N&#x2009;=&#x2009;87,652) and All of Us (N&#x2009;=&#x2009;45,445). We performed rare (&#x2009;<&#x2009;0.1% minor-allele-frequency) single-variant and aggregate testing of non-coding variants in regulatory regions based on proximal-regulatory, intergenic-regulatory and deep-intronic annotation. We observed 29 independent variants associated with height at P&#x2009;<&#x2009;after conditioning on previously reported variants, with effect sizes ranging from -7cm to +4.7&#x2009;cm. We also identified and replicated non-coding aggregate-based associations proximal to HMGA1 containing variants associated with a 5&#x2009;cm taller height and of highly-conserved variants in MIR497HG on chromosome 17. We have developed an approach for identifying non-coding rare variants in regulatory regions with large effects from whole-genome sequencing data associated with complex traits.

Humans↗

A tobit variance-component method for linkage analysis of censored trait data.

Variance-component (VC) methods are flexible and powerful procedures for the mapping of genes that influence quantitative traits. However, traditional VC methods make the critical assumption that the quantitative-trait data within a family either follow or can be transformed to follow a multivariate normal distribution. Violation of the multivariate normality assumption can occur if trait data are censored at some threshold value. Trait censoring can arise in a variety of ways, including assay limitation or confounding due to medication. Valid linkage analyses of censored data require the development of a modified VC method that directly models the censoring event. Here, we present such a model, which we call the "tobit VC method." Using simulation studies, we compare and contrast the performance of the traditional and tobit VC methods for linkage analysis of censored trait data. For the simulation settings that we considered, our results suggest that (1) analyses of censored data by using the traditional VC method lead to severe bias in parameter estimates and a modest increase in false-positive linkage findings, (2) analyses with the tobit VC method lead to unbiased parameter estimates and type I error rates that reflect nominal levels, and (3) the tobit VC method has a modest increase in linkage power as compared with the traditional VC method. We also apply the tobit VC method to censored data from the Finland-United States Investigation of Non-Insulin-Dependent Diabetes Mellitus Genetics study and provide two examples in which the tobit VC method yields noticeably different results as compared with the traditional method.

Analysis of Variance↗

The role of choice in health education intervention trials: a review and case study.

Although the randomized, controlled trial (RCT) is considered the gold standard in research for determining the efficacy of health education interventions, such trials may be vulnerable to "preference effects"; that is, differential outcomes depending on whether an individual is randomized to his or her preferred treatment. In this study, we review theoretical and empirical literature regarding designs that account for such effects in medical research, and consider the appropriateness of these designs to health education research. To illustrate the application of a preference design to health education research, we present analyses using process data from a mixed RCT/preference trial comparing two formats (Group or Self-Directed) of the "Women take PRIDE" heart disease management program. Results indicate that being able to choose one's program format did not significantly affect the decision to participate in the study. However, women who chose the Group format were over 4 times as likely to attend at least one class and were twice as likely to attend a greater number of classes than those who were randomized to the Group format. Several predictors of format preference were also identified, with important implications for targeting disease-management education to this population.

Aged↗

Substance use and psychotherapeutic medications: a likely contributor to menstrual disorders in women who are seropositive for human immunodeficiency virus.

OBJECTIVE: The purpose of this study was to evaluate the impact of substance use and psychotherapeutic medications on menstrual characteristics in women who are human immunodeficiency virus seropositive and seronegative. STUDY DESIGN: Menstrual calendars were prospectively collected for 1075 women who were human immunodeficiency virus seropositive and seronegative and who were enrolled in the Women's Interagency Human Immunodeficiency Virus Study or the Human Immunodeficiency Virus Epidemiology Research Study; several of the women were substance users or recipients of psychotherapeutic medications. RESULTS: Women who received methadone maintenance and who used injection drugs had substantially increased odds of a cycle of >or=90 days (odds ratio, 2.28; 95% CI, 1.23-4.22; and odds ratio, 3.87; 95% CI, 2.16-6.95, respectively). The use of psychotherapeutic medications increased the odds of having very short cycles, <18 days, and cycles of >or=90 days (odds ratio, 1.69; 95% CI, 1.16-2.45; and odds ratio, 1.86; 95% CI, 1.03-3.36, respectively). CONCLUSION: Clinicians should evaluate substance use, participation in methadone maintenance programs, and the use of psychotherapeutic medications and consider the neuroendocrinologic effects of these medications as a potential cause of menstrual disruptions.

Adult↗

Hypothesis testing in semiparametric additive mixed models.

We consider testing whether the nonparametric function in a semiparametric additive mixed model is a simple fixed degree polynomial, for example, a simple linear function. This test provides a goodness-of-fit test for checking parametric models against nonparametric models. It is based on the mixed-model representation of the smoothing spline estimator of the nonparametric function and the variance component score test by treating the inverse of the smoothing parameter as an extra variance component. We also consider testing the equivalence of two nonparametric functions in semiparametric additive mixed models for two groups, such as treatment and placebo groups. The proposed tests are applied to data from an epidemiological study and a clinical trial and their performance is evaluated through simulations.

Anticonvulsants↗

Scaled marginal models for multiple continuous outcomes.

In studies that involve multivariate outcomes it is often of interest to test for a common exposure effect. For example, our research is motivated by a study of neurocognitive performance in a cohort of HIV-infected women. The goal is to determine whether highly active antiretroviral therapy affects different aspects of neurocognitive functioning to the same degree and if so, to test for the treatment effect using a more powerful one-degree-of-freedom global test. Since multivariate continuous outcomes are likely to be measured on different scales, such a common exposure effect has not been well defined. We propose the use of a scaled marginal model for testing and estimating this global effect when the outcomes are all continuous. A key feature of the model is that the effect of exposure is represented by a common effect size and hence has a well-understood, practical interpretation. Estimating equations are proposed to estimate the regression coefficients and the outcome-specific scale parameters, where the correct specification of the within-subject correlation is not required. These estimating equations can be solved by repeatedly calling standard generalized estimating equations software such as SAS PROC GENMOD. To test whether the assumption of a common exposure effect is reasonable, we propose the use of an estimating-equation-based score-type test. We study the asymptotic efficiency loss of the proposed estimators, and show that they generally have high efficiency compared to the maximum likelihood estimators. The proposed method is applied to the HIV data.

Antiretroviral Therapy, Highly Active↗

Testing the correlation for clustered categorical and censored discrete time-to-event data when covariates are measured without/with errors.

In the analysis of clustered categorical data, it is of common interest to test for the correlation within clusters, and the heterogeneity across different clusters. We address this problem by proposing a class of score tests for the null hypothesis that the variance components are zero in random effects models, for clustered nominal and ordinal categorical responses. We extend the results to accommodate clustered censored discrete time-to-event data. We next consider such tests in the situation where covariates are measured with errors. We propose using the SIMEX method to construct the score tests for the null hypothesis that the variance components are zero. Key advantages of the proposed score tests are that they can be easily implemented by fitting standard polytomous regression models and discrete failure time models, and that they are robust in the sense that no assumptions need to be made regarding the distributions of the random effects and the unobserved covariates. The asymptotic properties of the proposed tests are studied. We illustrate these tests by analyzing two data sets and evaluate their performance with simulations.

Biometry↗

Ascertainment-adjusted parameter estimates revisited.

Ascertainment-adjusted parameter estimates from a genetic analysis are typically assumed to reflect the parameter values in the original population from which the ascertained data were collected. Burton et al. (2000) recently showed that, given unmodeled parameter heterogeneity, the standard ascertainment adjustment leads to biased parameter estimates of the population-based values. This finding has important implications in complex genetic studies, because of the potential existence of unmodeled genetic parameter heterogeneity. The authors further stated the important point that, given unmodeled heterogeneity, the ascertainment-adjusted parameter estimates reflect the true parameter values in the ascertained subpopulation. They illustrated these statements with two examples. By revisiting these examples, we demonstrate that if the ascertainment scheme and the nature of the data can be correctly modeled, then an ascertainment-adjusted analysis returns population-based parameter estimates. We further demonstrate that if the ascertainment scheme and data cannot be modeled properly, then the resulting ascertainment-adjusted analysis produces parameter estimates that generally do not reflect the true values in either the original population or the ascertained subpopulation.

Genetic Diseases, Inborn↗

Interictal epileptiform discharges do not change before seizures during sleep.

PURPOSE: Whether interictal epileptiform discharges (IEDs) increase, decrease, or are unchanged before epileptic seizures has implications for the pathophysiology of epilepsy. Prior studies relating IEDs and seizures have not demonstrated a change in IEDs before seizures. However, they have not controlled for changes in the depth of sleep. Our objective was to test the hypothesis that IEDs are related to seizures during sleep while adjusting for log delta power (LDP), a continuous measure of sleep depth. METHODS: Twenty-two seizures during sleep were identified in 16 subjects with epilepsy admitted for presurgical monitoring. The IEDs that occurred in the hour of sleep before each seizure were used to test the relation between IEDs and seizure occurrence. Sleep depth was measured by LDP (quantity of 1- to 4-Hz activity in 30-s epochs), and records were scored visually for sleep staging and for IEDs. Multivariate logistic regression analyses were applied. RESULTS: Adjusting for LDP, number of seizures before the current seizure, quartile of the night, and total number of IEDs that occurred during the night, IED did not increase or decrease before seizures (p > 0.1). The rate of IEDs increased directly with LDP (p=0.0001), as shown in prior work. CONCLUSIONS: IEDs are not activated or suppressed before seizures during sleep, suggesting that different pathophysiologic processes underlie these two phenomena. These results corroborate prior studies, while providing a more advanced analysis by adjusting for sleep depth and applying multivariate logistic regression analyses.

Adult↗

Effects of Long-Term Video-electroencephalographic Monitoring on Mood in Epilepsy Patients.

This study examines changes in mood of 79 epilepsy patients who completed the Profile of Mood States during long-term video-electroencephalographic monitoring (LTM). Statistical linear models included the effects of age, gender, increased seizure frequency, sleep deprivation, and taper of antiepileptic drugs (AEDs) on mood. Sleep deprivation increased fatigue and decreased vigor from baseline to Day 3, but not from baseline to Day 8 or the final day of the protocol. Taper of AEDs did not adversely affect mood, with removal of phenytoin improving mood. Subjects who had seizures during LTM also improved in mood, becoming less depressed and less fatigued than those who did not have seizures. Overall, our data indicate that LTM does not adversely affect mood. However, in the first few days of LTM, sleep deprivation may produce fatigue and lack of vigor, and should be used only as needed to provoke seizures.

Journal Article↗