PubMed Health⌕ Search

Biomedical subjects

Taesung Park

Publications and source records attributed to Taesung Park.

6 recordsLinked to original sources

Evaluation of normalization methods for microarray data.

BACKGROUND: Microarray technology allows the monitoring of expression levels for thousands of genes simultaneously. This novel technique helps us to understand gene regulation as well as gene by gene interactions more systematically. In the microarray experiment, however, many undesirable systematic variations are observed. Even in replicated experiment, some variations are commonly observed. Normalization is the process of removing some sources of variation which affect the measured gene expression levels. Although a number of normalization methods have been proposed, it has been difficult to decide which methods perform best. Normalization plays an important role in the earlier stage of microarray data analysis. The subsequent analysis results are highly dependent on normalization. RESULTS: In this paper, we use the variability among the replicated slides to compare performance of normalization methods. We also compare normalization methods with regard to bias and mean square error using simulated data. CONCLUSIONS: Our results show that intensity-dependent normalization often performs better than global normalization methods, and that linear and nonlinear normalization methods perform similarly. These conclusions are based on analysis of 36 cDNA microarrays of 3,840 genes obtained in an experiment to search for changes in gene expression profiles during neuronal differentiation of cortical stem cells. Simulation studies confirm our findings.

Analysis of Variance↗

Statistical tests for identifying differentially expressed genes in time-course microarray experiments.

MOTIVATION: Microarray technology allows the monitoring of expression levels for thousands of genes simultaneously. In time-course experiments in which gene expression is monitored over time, we are interested in testing gene expression profiles for different experimental groups. However, no sophisticated analytic methods have yet been proposed to handle time-course experiment data. RESULTS: We propose a statistical test procedure based on the ANOVA model to identify genes that have different gene expression profiles among experimental groups in time-course experiments. Especially, we propose a permutation test which does not require the normality assumption. For this test, we use residuals from the ANOVA model only with time-effects. Using this test, we detect genes that have different gene expression profiles among experimental groups. The proposed model is illustrated using cDNA microarrays of 3840 genes obtained in an experiment to search for changes in gene expression profiles during neuronal differentiation of cortical stem cells.

Algorithms↗

Risk analysis of aseptic meningitis after measles-mumps-rubella vaccination in Korean children by using a case-crossover design.

Epidemiologic study of a vaccine's adverse events is not easy; so many countries have no reliable data. Vaccines containing the Urabe or Hoshino strain have been withdrawn from use in several countries. However, the data are not strong enough to form the basis of a recommendation not to use specific strains. The authors used a case-crossover design to estimate the relative risk of aseptic meningitis in children after receiving the measles-mumps-rubella vaccine in Korea. Study subjects were hospitalized children aged 8-36 months who had aseptic meningitis in 1998. Cases were confirmed by hospital chart reviews using previously defined criteria. Through a telephone survey, the authors obtained vaccination date and place information from parents' vaccination records. Study results showed that no significant risk was associated with the Jeryl Lynn or Rubini strain of the vaccine (relative risk = 0.6, 95% confidence interval (CI): 0.18, 1.97). For the Urabe or Hoshino strain, the relative risk was 5.5 (95% CI: 2.6, 11.8); the risk increased in the third week after vaccination (relative risk = 15.6, 95% CI: 5.9, 41.2) and was elevated until the sixth week. The case-crossover design was useful in confirming the risk of acute adverse events after receiving vaccines.

Child, Preschool↗

Longitudinal data analysis in pedigree studies.

Longitudinal family studies provide a valuable resource for investigating genetic and environmental factors that influence long-term averages and changes over time in a complex trait. This paper summarizes 13 contributions to Genetic Analysis Workshop 13, which include a wide range of methods for genetic analysis of longitudinal data in families. The methods can be grouped into two basic approaches: 1) two-step modeling, in which repeated observations are first reduced to one summary statistic per subject (e.g., a mean or slope), after which this statistic is used in a standard genetic analysis, or 2) joint modeling, in which genetic and longitudinal model parameters are estimated simultaneously in a single analysis. In applications to Framingham Heart Study data, contributors collectively reported evidence for genes that affected trait mean on chromosomes 1, 2, 3, 5, 8, 9, 10, 13, and 17, but most did not find genes affecting slope. Applications to simulated data suggested that even for a gene that only affected slope, use of a mean-type statistic could provide greater power than a slope-type statistic for detecting that gene. We report on the results of a small experiment that sheds some light on this apparently paradoxical finding, and indicate how one might form a more powerful test for finding a slope-affecting gene. Several areas for future research are discussed.

Cardiovascular Diseases↗

A Bayesian hierarchical model for categorical data with nonignorable nonresponse.

Log-linear models have been shown to be useful for smoothing contingency tables when categorical outcomes are subject to nonignorable nonresponse. A log-linear model can be fit to an augmented data table that includes an indicator variable designating whether subjects are respondents or nonrespondents. Maximum likelihood estimates calculated from the augmented data table are known to suffer from instability due to boundary solutions. Park and Brown (1994, Journal of the American Statistical Association 89, 44-52) and Park (1998, Biometrics 54, 1579-1590) developed empirical Bayes models that tend to smooth estimates away from the boundary. In those approaches, estimates for nonrespondents were calculated using an EM algorithm by maximizing a posterior distribution. As an extension of their earlier work, we develop a Bayesian hierarchical model that incorporates a log-linear model in the prior specification. In addition, due to uncertainty in the variable selection process associated with just one log-linear model, we simultaneously consider a finite number of models using a stochastic search variable selection (SSVS) procedure due to George and McCulloch (1997, Statistica Sinica 7, 339-373). The integration of the SSVS procedure into a Markov chain Monte Carlo (MCMC) sampler is straightforward, and leads to estimates of cell frequencies for the nonrespondents that are averages resulting from several log-linear models. The methods are demonstrated with a data example involving serum creatinine levels of patients who survived renal transplants. A simulation study is conducted to investigate properties of the model.

Algorithms↗

Covariance models for nested repeated measures data: analysis of ovarian steroid secretion data.

We consider several covariance models for analysing repeated measures data from a study of ovarian steroid secretion in reproductive-aged women. Urinary oestradiol and serum oestrogen were repeatedly observed over three or four menstrual periods, each period separated by one year. For each menstrual period, daily first morning urine specimens were collected 8 to 18 times, and serum specimens 2 to 5 times. Thus, measurements were repeatedly observed over menstrual cycle days within menstrual periods. Owing to missing observations, the number of observations differed from subject to subject. In this study, there were two repeat factors: menstrual cycle day and menstrual period. The first repeat factor, cycle day, is nested within the second repeat factor, menstrual period. In analysing these nested repeated measures data, the correlation structure should be modelled that will account for both repeat factors. We present several covariance models for defining appropriate covariance structures for these data.

Adult↗