PubMed Health⌕ Search

Biomedical subjects

Nianjun Liu

Publications and source records attributed to Nianjun Liu.

14 recordsLinked to original sources

Regional admixture mapping and structured association testing: conceptual unification and an extensible general linear model.

Individual genetic admixture estimates, determined both across the genome and at specific genomic regions, have been proposed for use in identifying specific genomic regions harboring loci influencing phenotypes in regional admixture mapping (RAM). Estimates of individual ancestry can be used in structured association tests (SAT) to reduce confounding induced by various forms of population substructure. Although presented as two distinct approaches, we provide a conceptual framework in which both RAM and SAT are special cases of a more general linear model. We clarify which variables are sufficient to condition upon in order to prevent spurious associations and also provide a simple closed form "semiparametric" method of evaluating the reliability of individual admixture estimates. An estimate of the reliability of individual admixture estimates is required to make an inherent errors-in-variables problem tractable. Casting RAM and SAT methods as a general linear model offers enormous flexibility enabling application to a rich set of phenotypes, populations, covariates, and situations, including interaction terms and multilocus models. This approach should allow far wider use of RAM and SAT, often using standard software, in addressing admixture as either a confounder of association studies or a tool for finding loci influencing complex phenotypes in species as diverse as plants, humans, and nonhuman animals.

Computer Simulation↗

Effect of cardiac motion on body surface electrocardiographic potentials: an MRI-based simulation study.

This paper describes an electrical model of cardiac ventricles incorporating real geometry and motion. The heart anatomy and its motion through the cardiac cycle are obtained from segmentations of multiple-slice MRI time sequences; the special conduction system is constructed using an automated mapping procedure from an existing static heart model. The heart model is mounted in an anatomically realistic voxel model of the human body. The cardiac electrical source and surface potentials are determined numerically using both a finite-difference scheme and a boundary-element method with the incorporation of the motion of the heart. The electrocardiograms (ECG) and body surface potential maps are calculated and compared to the static simulation in the resting heart. The simulations demonstrate that introducing motion into the cardiac model modifies the ECG signals, with the most obvious change occurring during the T-wave at peak contraction of the ventricles. Body surface potential maps differ in some local positions during the T-wave, which may be of importance to a number of cardiac models, including those incorporating inverse methods.

Algorithms↗

PSMIX: an R package for population structure inference via maximum likelihood method.

BACKGROUND: Inference of population stratification and individual admixture from genetic markers is an integrative part of a study in diverse situations, such as association mapping and evolutionary studies. Bayesian methods have been proposed for population stratification and admixture inference using multilocus genotypes and widely used in practice. However, these Bayesian methods demand intensive computation resources and may run into convergence problem in Markov Chain Monte Carlo based posterior samplings. RESULTS: We have developed PSMIX, an R package based on maximum likelihood method using expectation-maximization algorithm, for inference of population stratification and individual admixture. CONCLUSION: Compared with software based on Bayesian methods (e.g., STRUCTURE), PSMIX has similar accuracy, but more efficient computations.PSMIX and its supplemental documents are freely available at http://bioinformatics.med.yale.edu/PSMIX.

Algorithms↗

Macrophage migration inhibitory factor promoter polymorphisms and the clinical expression of scleroderma.

OBJECTIVE: To investigate the potential association between functional polymorphisms in the gene for the innate mediator, macrophage migration inhibitory factor (MIF), and the clinical expression of systemic sclerosis (SSc). METHODS: Genomic DNA samples and clinical data were collected from the Scleroderma Family Registry and DNA Repository at the University of Texas Health Science Center at Houston. A total of 740 subjects were studied; 203 of them had diffuse cutaneous SSc (dcSSc), 283 had limited cutaneous SSc (lcSSc), and the remaining 254 healthy subjects served as controls. Association analyses were performed on the whole data set and on patient and sex subsets. Significant relationships were determined between clinical variables and MIF polymorphisms for each disease subtype in the studied groups. RESULTS: The frequency of the -173*C MIF allele, which was previously reported to be associated with high production of MIF, was lower in the lcSSc group (12.6%) than in the dcSSc (19.2%) or control (18.5%) groups (P = 0.010 and P = 0.011, respectively). Haplotype analysis for 2 closely linked polymorphisms in the MIF promoter showed that in white subjects with lcSSc or dcSSc, the lcSSc population had a significantly lower representation of the high-expression MIF haplotype defined by -173*C and -794 with 7 CATT repeats (C7) (P = 0.015, odds ratio 1.94 [95% confidence interval 1.14-3.32]). Fibroblasts encoding the C7 MIF haplotype were observed to produce more MIF upon in vitro stimulation than those with a non-C7 haplotype. CONCLUSION: Functional promoter polymorphisms in the MIF gene affect the clinical presentation of SSc. The proinflammatory haplotype defined by C7 is underrepresented in patients with lcSSc.

Female↗

Haplotype analysis in the presence of informatively missing genotype data.

It is common to have missing genotypes in practical genetic studies, but the exact underlying missing data mechanism is generally unknown to the investigators. Although some statistical methods can handle missing data, they usually assume that genotypes are missing at random, that is, at a given marker, different genotypes and different alleles are missing with the same probability. These include those methods on haplotype frequency estimation and haplotype association analysis. However, it is likely that this simple assumption does not hold in practice, yet few studies to date have examined the magnitude of the effects when this simplifying assumption is violated. In this study, we demonstrate that the violation of this assumption may lead to serious bias in haplotype frequency estimates, and haplotype association analysis based on this assumption can induce both false-positive and false-negative evidence of association. To address this limitation in the current methods, we propose a general missing data model to characterize missing data patterns across a set of two or more markers simultaneously. We prove that haplotype frequencies and missing data probabilities are identifiable if and only if there is linkage disequilibrium between these markers under our general missing data model. Simulation studies on the analysis of haplotypes consisting of two single nucleotide polymorphisms illustrate that our proposed model can reduce the bias both for haplotype frequency estimates and association analysis due to incorrect assumption on the missing data mechanism. Finally, we illustrate the utilities of our method through its application to a real data set.

Algorithms↗

A non-parametric approach to population structure inference using multilocus genotypes.

Inference of population structure from genetic markers is helpful in diverse situations, such as association and evolutionary studies. In this paper, we describe a two-stage strategy in inferring population structure using multilocus genotype data. In the first stage, we use dimension reduction methods such as singular value decomposition to reduce the dimension of the data, and in the second stage, we use clustering methods on the reduced data to identify population structure. The strategy has the ability to identify population structure and assign each individual to its corresponding subpopulation. The strategy does not depend on any population genetics assumptions (such as Hardy-Weinberg equilibrium and linkage equilibrium between loci within populations) and can be used with any genotype data. When applied to real and simulated data, the strategy is found to have similar or better performance compared with STRUCTURE, the most popular method in current use. Therefore, the proposed strategy provides a useful alternative to analyse population data.

Cluster Analysis↗

A Bayesian genome screening of maximum number of drinks as an alcoholism phenotype with the new Haseman-Elston method.

Common human disorders, such as alcoholism, may be the result of interactions of many genes as well as environmental risk factors. Therefore, it is important to incorporate gene x gene and gene x environment interactions in complex disease gene mapping. In this study, we applied a robust Bayesian genome screening method that can incorporate interaction effects to map genes underlying alcoholism through its application to the data of the Collaborative Studies on Genetics of Alcoholism provided by Genetic Analysis Workshop 14. Our Bayesian genome screening method uses the regression-based stochastic variable selection, coupled with the new Haseman-Elston method to identify markers linked to phenotypes of interest. Compared to traditional linkage methods based on single-gene disease models, our method allows for multilocus disease models for simultaneous screening including both main and interaction (epistatic) effects. It is conceptually simple and computationally efficient through the use of Gibbs sampler. We conducted genome-wide analysis and comparison between scans based on microsatellites and single-nucleotide polymorphisms. A total of 328 microsatellites and 11,560 single-nucleotide polymorphisms (by Affymetrix) on 22 autosomal chromosomes and sex chromosome were used.

Alcoholism↗

Whole-genome association studies on alcoholism comparing different phenotypes using single-nucleotide polymorphisms and microsatellites.

Alcoholism is a complex disease. As with other common diseases, genetic variants underlying alcoholism have been illusive, possibly due to the small effect from each individual susceptible variant, gene x environment and gene x gene interactions and complications in phenotype definition. We conducted association tests, the family-based association tests (FBAT) and the backward haplotype transmission association (BHTA), on the Collaborative Study of the Genetics of Alcoholism (COGA) data provided by Genetic Analysis Workshop (GAW) 14. Efron's local false discovery rate method was applied to control the proportion of false discoveries. For FBAT, we compared the results based on different types of genetic markers (single-nucleotide polymorphisms (SNPs) versus microsatellites) and different phenotype definitions (clinical diagnoses versus electrophysiological phenotypes). Significant association results were found only between SNPs and clinical diagnoses. In contrast, significant results were found only between microsatellites and electrophysiological phenotypes. In addition, we obtained the association results for SNPs and microsatellites using COGA diagnosis as phenotype based on BHTA. In this case, the results for SNPs and microsatellites are more consistent. Compared to FBAT, more significant markers are detected with BHTA.

Alcoholism↗

Comparison of single-nucleotide polymorphisms and microsatellites in inference of population structure.

Single-nucleotide polymorphisms (SNPs) are a class of attractive genetic markers for population genetic studies and for identifying genetic variations underlying complex traits. However, the usefulness and efficiency of SNPs in comparison to microsatellites in different scientific contexts, e.g., population structure inference or association analysis, still must be systematically evaluated through large empirical studies. In this article, we use the Collaborative Studies on Genetics of Alcoholism (COGA) data from Genetic Analysis Workshop 14 (GAW14) to compare the performance of microsatellites and SNPs in the whole human genome in the context of population structure inference. A total of 328 microsatellites and 15,840 SNPs are used to infer population structure in 236 unrelated individuals. We find that, on average, the informativeness of random microsatellites is four to twelve times that of random SNPs for various population comparisons, which is consistent with previous studies. Our results also indicate that for the combined set of microsatellites and SNPs, SNPs constitute the majority among the most informative markers and the use of these SNPs leads to better inference of population structure than the use of microsatellites. We also find that the inclusion of less informative markers may add noise and worsen the results.

Genetic Loci↗

Whole-genome linkage analysis in mapping alcoholism genes using single-nucleotide polymorphisms and microsatellites.

There is currently a great interest in using single-nucleotide polymorphisms (SNPs) in genetic linkage and association studies because of the abundance of SNPs as well as the availability of high-throughput genotyping technologies. In this study, we compared the performance of whole-genome scans using SNPs with microsatellites on 143 pedigrees from the Collaborative Studies on Genetics of Alcoholism provided by Genetic Analysis Workshop 14. A total of 315 microsatellites and 10,081 SNPs from Affymetrix on 22 autosomal chromosomes were used in our analyses. We found that the results from the two scans had good overall concordance. One region on chromosome 2 and two regions on chromosome 7 showed significant linkage signals (i.e., NPL >or= 2) for alcoholism from both the SNP and microsatellite scans. The different results observed between the two scans may be explained by the difference observed in information content between the SNPs and the microsatellites.

Alcoholism↗

Whole-genome association analysis to identify markers associated with recombination rates using single-nucleotide polymorphisms and microsatellites.

Recombination during meiosis is one of the most important biological processes, and the level of recombination rates for a given individual is under genetic control. In this study, we conducted genome-wide association studies to identify chromosomal regions associated with recombination rates. We analyzed genotype data collected on the pedigrees in the Collaborative Study on the Genetics on Alcoholism data provided by Genetic Analysis Workshop 14. A total of 315 microsatellites and 10,081 single-nucleotide polymorphisms from Affymetrix on 22 autosomal chromosomes were used in our association analysis. Genome-wide gender-specific recombination counts for family founders were inferred first and association analysis was performed using multiple linear regressions. We used the positive false discovery rate (pFDR) to account for multiple comparisons in the two genome-wide scans. Eight regions showed some evidence of association with recombination counts based on the single-nucleotide polymorphism analysis after adjusting for multiple comparisons. However, no region was found to be significant using microsatellites.

Female↗

Inferring protein-protein interactions through high-throughput interaction data from diverse organisms.

MOTIVATION: Identifying protein-protein interactions is critical for understanding cellular processes. Because protein domains represent binding modules and are responsible for the interactions between proteins, computational approaches have been proposed to predict protein interactions at the domain level. The fact that protein domains are likely evolutionarily conserved allows us to pool information from data across multiple organisms for the inference of domain-domain and protein-protein interaction probabilities. RESULTS: We use a likelihood approach to estimating domain-domain interaction probabilities by integrating large-scale protein interaction data from three organisms, Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. The estimated domain-domain interaction probabilities are then used to predict protein-protein interactions in S.cerevisiae. Based on a thorough comparison of sensitivity and specificity, Gene Ontology term enrichment and gene expression profiles, we have demonstrated that it may be far more informative to predict protein-protein interactions from diverse organisms than from a single organism. AVAILABILITY: The program for computing the protein-protein interaction probabilities and supplementary material are available at http://bioinformatics.med.yale.edu/interaction.

Binding Sites↗

Haplotype block structures show significant variation among populations.

Recent studies suggest that haplotypes tend to have block-like structures throughout the human genome. Several methods were proposed for haplotype block partitioning and for tagging single-nucleotide polymorphism (SNP) identification. In population genetics studies, several research groups compared block structures across human populations. However, the measures used to quantify population similarity are either less than satisfactory or nonexistent. In this article, we propose several similarity measures to facilitate the comparisons of haplotype structures, namely block boundaries and tagging SNPs, across populations. With these measures, we can more objectively compare haplotype block structures and tagging SNP sets between different populations. In addition, these measures allow us to compare the results of different methods for block partition and tagging SNP identification. When we applied these measures to a real data set on chromosome 10 in 16 worldwide populations, we found that in this genome region: 1) haplotype block boundaries vary among populations, with European and some African populations showing similar boundaries but other populations showing other patterns; 2) tagging SNP sets are generally similar for populations with similar haplotype block structures but differ if the block structures differ; and 3) all but one of the block finding methods we tested yield consistent results, although variations exist regarding consistency. Our tentative results show that at least in the genome region studied, it is unlikely that a common haplotype pattern exists for all human populations: many populations, even in the same geographical region, may have different haplotype patterns.

Chromosome Mapping↗

Addition of the alpha2-antagonist yohimbine to fluoxetine: effects on rate of antidepressant response.

Electrophysiological studies suggest that alpha2-adrenoceptors profoundly affect monoaminergic neurotransmission by enhancing noradrenergic tone and serotonergic firing rates. Recent reports suggest that alpha2-antagonism may hasten and improve the response to antidepressant medications. To test this hypothesis, a randomized double-blind controlled trial was undertaken to determine if the combination of an alpha2-antagonist (yohimbine) with a selective serotonin reuptake agent (SSRI) (fluoxetine) results in more rapid onset of antidepressant action than an SSRI agent alone. In all, 50 subjects with a DSM-IV diagnosis of major depressive disorder confirmed by SCID interview were randomly assigned to receive either fluoxetine 20 mg plus placebo (F/P) or fluxetine 20 mg plus a titrated dose of yohimbine (F/Y). The yohimbine dose was titrated based on blood pressure changes over the treatment period, in a blind-preserving manner. Hamilton depression scale ratings (HDRS) and clinical global impression (CGI) ratings were obtained weekly over a period of 6 weeks. The rate of achieving categorical positive responses was significantly more rapid in the F/Y group compared to the F/P group using both the HDRS and the CGI scales as outcome measures in a survival analysis using a log-rank test (chi2(1) = 5.86, p = 0.016 and chi2(1) = 5.29, p = 0.021, respectively). At the last observed visit, 18 (69%) of the 26 F/Y subjects met the response criteria for CGI compared to 10 (42%) of 24 F/P subjects. Using the HDRS criteria, 17 (65%) of 26 F/Y subject vs 10 (42%) of 24 F/P subjects were responders. The addition of the alpha2-antagonist yohimbine to fluoxetine appears to hasten the antidepressant response. There is also a trend suggesting an increased percentage of responders to the combined treatment at the end of the 6-week trial.

Adolescent↗