PubMed Health⌕ Search

Biomedical subjects

Qiuying Sha

Publications and source records attributed to Qiuying Sha.

7 recordsLinked to original sources

Phenotypic and Genetic Associations Between Cardiovascular Disease Subtypes and Alzheimer's Disease.

BACKGROUND: Cardiovascular disease (CVD) and Alzheimer's disease (AD) are major public health concerns that share overlapping risk factors and potential mechanistic pathways. While vascular contributions to cognitive decline are well-documented, the specific relationships between AD and different CVD subtypes remain poorly understood. METHODS: We examined associations between AD and 11 CVD subtypes using logistic regression models in two large biobanks: the UK Biobank (n = 502,133) and the All of Us Research Program (n = 287,011). Models were adjusted for demographic, lifestyle, and clinical covariates. We also explored genetic overlap between AD and CVD traits through colocalization of significant single nucleotide polymorphisms (SNPs) (p < 5&#xd7;10-8) using genome-wide association study (GWAS) data. RESULTS: Most CVD subtypes were significantly associated with AD in both cohorts. Hypotension had the strongest and most consistent association, followed by hypertension and cerebral infarction. Acute myocardial infarction was the only subtype not significantly linked to AD. Genetic analyses revealed shared loci between AD and CVD-related traits, particularly in regions near APOE, MAPT, and genes influencing myocardial structure and vascular function. CONCLUSIONS: This study identifies subtype-specific CVD associations with AD across two diverse cohorts and highlights shared genetic architecture underlying heart-brain interactions. These findings underscore the importance of vascular health in AD risk and suggest that certain CVD subtypes, especially hypotension, may play underrecognized roles in cognitive decline.

Alzheimer&#x2019;s disease↗

Analytical correction for multiple testing in admixture mapping.

Admixture mapping, using unrelated individuals from the admixture populations that result from recent mating between members of each parental population, is an efficient approach to localize disease-causing variants that differ in frequency between two or more historically separated populations. Recently, several methods have been proposed to test linkage between a susceptibility gene and a disease locus by using admixture-generated linkage disequilibrium (LD) for each of the genotyped markers. In a genome scan, admixture mapping usually tests 2,000 to 3,000 markers across the genome. Currently, either a very conservative Sidak (or Bonferroni) correction or a very time consuming simulation-based method is used to correct for the multiple tests and evaluate the overall p value. In this report, we propose a computationally efficient analytical approach for correction of the multiple tests and for calculating the overall p value for an admixture genome scan. Except for the Sidak (or Bonferroni) correction, our proposed method is the first analytical approach for correction of the multiple tests and for calculating the overall p value for a genome scan. Our simulation studies show that the proposed method gives correct overall type I error rates for genome scans in all cases, and is much more computationally efficient than simulation-based methods.

Computational Biology↗

A combinatorial searching method for detecting a set of interacting loci associated with complex traits.

Complex diseases are presumed to be the results of the interaction of several genes and environmental factors, with each gene only having a small effect on the disease. Mapping complex disease genes therefore becomes one of the greatest challenges facing geneticists. Most current approaches of association studies essentially evaluate one marker or one gene (haplotype approach) at a time. These approaches ignore the possibility that effects of multilocus functional genetic units may play a larger role than a single-locus effect in determining trait variability. In this article, we propose a Combinatorial Searching Method (CSM) to detect a set of interacting loci (may be unlinked) that predicts the complex trait. In the application of the CSM, a simple filter is used to filter all the possible locus-sets and retain the candidate locus-sets, then a new objective function based on the cross-validation and partitions of the multi-locus genotypes is proposed to evaluate the retained locus-sets. The locus-set with the largest value of the objective function is the final locus-set and a permutation procedure is performed to evaluate the overall p-value of the test for association between the final locus-set and the trait. The performance of the method is evaluated by simulation studies as well as by being applied to a real data set. The simulation studies show that the CSM has reasonable power to detect high-order interactions. When the CSM is applied to a real data set to detect the locus-set (among the 13 loci in the ACE gene) that predicts systolic blood pressure (SBP) or diastolic blood pressure (DBP), we found that a four-locus gene-gene interaction model best predicts SBP with an overall p-value = 0.033, and similarly a two-locus gene-gene interaction model best predicts DBP with an overall p-value = 0.045.

Blood Pressure↗

Haplotype sharing transmission/disequilibrium tests that allow for genotyping errors.

The present study introduces new Haplotype Sharing Transmission/Disequilibrium Tests (HS-TDTs) that allow for random genotyping errors. We evaluate the type I error rate and power of the new proposed tests under a variety of scenarios and perform a power comparison among the proposed tests, the HS-TDT and the single-marker TDT. The results indicate that the HS-TDT shows a significant increase in type I error when applied to data in which either Mendelian inconsistent trios are removed or Mendelian inconsistent markers are treated as missing genotypes, and the magnitude of the type I error increases both with an increase in sample size and with an increase in genotyping error rate. The results also show that a simple strategy, that is, merging each rare haplotype to a most similar common haplotype, can control the type I error inflation for a wide range of genotyping error rates, and after merging rare haplotypes, the power of the test is very similar to that without merging the rare haplotypes. Therefore, we conclude that a simple strategy may make the HS-TDT robust to genotyping errors. Our simulation results also show that this strategy may also be applicable to other haplotype-based TDTs.

Algorithms↗

Tests of association between quantitative traits and haplotypes in a reduced-dimensional space.

Candidate gene association tests are currently performed using several intragenic SNPs simultaneously, by testing SNP haplotype or genotype effects in multifactorial diseases or traits. The number of haplotypes drastically increases with an increase in the number of typed SNPs. As a result, large numbers of haplotypes will introduce large degrees of freedom in haplotype-based tests, and thus limit the power of the tests. In this study we propose using the principal component method to reduce the dimension, and then construct association tests on the lower-dimensional space to test the association between haplotypes and a quantitative trait using population-based samples. The proposed method allows ambiguous haplotypes. We use simulation studies to evaluate the type I error rate of the tests, and compare the power of the proposed tests with that of the tests without dimension reduction, and the tests with dimension reduction by merging rare haplotypes. The simulation results show that the proposed tests have correct type I error rates and are more powerful than other tests in most cases considered in our simulation studies.

Computer Simulation↗

Joint analysis of two microarray gene-expression data sets to select lung adenocarcinoma marker genes.

BACKGROUND: Due to the high cost and low reproducibility of many microarray experiments, it is not surprising to find a limited number of patient samples in each study, and very few common identified marker genes among different studies involving patients with the same disease. Therefore, it is of great interest and challenge to merge data sets from multiple studies to increase the sample size, which may in turn increase the power of statistical inferences. In this study, we combined two lung cancer studies using microarray GeneChip, employed two gene shaving methods and a two-step survival test to identify genes with expression patterns that can distinguish diseased from normal samples, and to indicate patient survival, respectively. RESULTS: In addition to common data transformation and normalization procedures, we applied a distribution transformation method to integrate the two data sets. Gene shaving (GS) methods based on Random Forests (RF) and Fisher's Linear Discrimination (FLD) were then applied separately to the joint data set for cancer gene selection. The two methods discovered 13 and 10 marker genes (5 in common), respectively, with expression patterns differentiating diseased from normal samples. Among these marker genes, 8 and 7 were found to be cancer-related in other published reports. Furthermore, based on these marker genes, the classifiers we built from one data set predicted the other data set with more than 98% accuracy. Using the univariate Cox proportional hazard regression model, the expression patterns of 36 genes were found to be significantly correlated with patient survival (p < 0.05). Twenty-six of these 36 genes were reported as survival-related genes from the literature, including 7 known tumor-suppressor genes and 9 oncogenes. Additional principal component regression analysis further reduced the gene list from 36 to 16. CONCLUSION: This study provided a valuable method of integrating microarray data sets with different origins, and new methods of selecting a minimum number of marker genes to aid in cancer diagnosis. After careful data integration, the classification method developed from one data set can be applied to the other with high prediction accuracy.

Adenocarcinoma↗

Transmission/disequilibrium test based on haplotype sharing for tightly linked markers.

Studies using haplotypes of multiple tightly linked markers are more informative than those using a single marker. However, studies based on multimarker haplotypes have some difficulties. First, if we consider each haplotype as an allele and use the conventional single-marker transmission/disequilibrium test (TDT), then the rapid increase in the degrees of freedom with an increasing number of markers means that the statistical power of the conventional tests will be low. Second, the parental haplotypes cannot always be unambiguously reconstructed. In the present article, we propose a haplotype-sharing TDT (HS-TDT) for linkage or association between a disease-susceptibility locus and a chromosome region in which several tightly linked markers have been typed. This method is applicable to both quantitative traits and qualitative traits. It is applicable to any size of nuclear family, with or without ambiguous phase information, and it is applicable to any number of alleles at each of the markers. The degrees of freedom (in a broad sense) of the test increase linearly as the number of markers considered increases but do not increase as the number of alleles at the markers increases. Our simulation results show that the HS-TDT has the correct type I error rate in structured populations and that, in most cases, the power of HS-TDT is higher than the power of the existing single-marker TDTs and haplotype-based TDTs.

Computer Simulation↗