PubMed Health⌕ Search

Biomedical subjects

Joanna H Shih

Publications and source records attributed to Joanna H Shih.

17 recordsLinked to original sources

Appropriateness of some resampling-based inference procedures for assessing performance of prognostic classifiers derived from microarray data.

The goal of many gene-expression microarray profiling clinical studies is to develop a multivariate classifier to predict patient disease outcome from a gene-expression profile measured on some biological specimen from the patient. Often some preliminary validation of the predictive power of a profile-based classifier is carried out using the same data set that was used to derive the classifier. Techniques such as cross-validation or bootstrapping can be used in this setting to assess predictive power, and if applied correctly, can result in a less biased estimate of predictive accuracy of a classifier. However, some investigators have attempted to apply standard statistical inference procedures to assess the statistical significance of associations between true and cross-validated predicted outcomes. We demonstrate in this paper that naïve application of standard statistical inference procedures to these measures of association under null situations can result in greatly inflated testing type I error rates. Under alternatives of small to moderate associations, confidence interval coverage probabilities may be too low, although for very large associations coverage probabilities approach their intended values. Our results suggest that caution should be exercised in interpreting some of the claims of exceptional prognostic classifier performance that have been reported in prominent biomedical journals in the past few years.

Clinical Trials as Topic↗

Heterologous tissue culture expression signature predicts human breast cancer prognosis.

BACKGROUND: Cancer patients have highly variable clinical outcomes owing to many factors, among which are genes that determine the likelihood of invasion and metastasis. This predisposition can be reflected in the gene expression pattern of the primary tumor, which may predict outcomes and guide the choice of treatment better than other clinical predictors. METHODOLOGY/PRINCIPAL FINDINGS: We developed an mRNA expression-based model that can predict prognosis/outcomes of human breast cancer patients regardless of microarray platform and patient group. Our model was developed using genes differentially expressed in mouse plasma cell tumors growing in vivo versus those growing in vitro. The prediction system was validated using published data from three cohorts of patients for whom microarray and clinical data had been compiled. The model stratified patients into four independent survival groups (BEST, GOOD, BAD, and WORST: log-rank test p = 1.7x10(-8)). CONCLUSIONS: Our model significantly improved the survival prediction over other expression-based models and permitted recognition of patients with different prognoses within the estrogen receptor-positive group and within a single pathological tumor class. Basing our predictor on a dataset that originated in a different species and a different cell type may have rendered it less sensitive to proliferation differences and endowed it with wide applicability. SIGNIFICANCE: Prognosis prediction for patients with breast cancer is currently based on histopathological typing and estrogen receptor positivity. Yet both assays define groups that are heterogeneous in survival. Gene expression profiling allows subdivision of these groups and recognition of patients whose tumors are very unlikely to be lethal and those with much grimmer outlooks, which can augment the predictive power of conventional tumor analysis and aid the clinician in choosing relaxed vs. aggressive therapy.

Animals↗

Desmoglein 3 as a prognostic factor in lung cancer.

Desmoglein 3 is a desmosomal protein of the cadherin family. Our cDNA expression profile demonstrated that desmoglein 3 was highly expressed in squamous cell carcinoma of the lung but not detected in pulmonary adenocarcinoma or normal lung. To investigate the clinical significance of desmoglein 3 in lung cancer, we surveyed its expression in primary non-small-cell lung cancers and neuroendocrine tumors. We used immunohistochemical analysis to examine the expression of desmoglein 3 by using tissue microarrays containing samples from 300 surgical non-small-cell lung cancer and 183 lung neuroendocrine tumor. Staining status was determined based on the sum of the distribution score (0, 1, or 2) and the intensity score (0, 1, 2, or 3) of the staining signal. Follow-up was available for 346 cases (median follow-up of 2.8 years). We determined the survival statistical significance of desmoglein 3 by using the log-rank test, and we plotted Kaplan-Meier curves. Negative immunohistochemical staining with desmoglein 3 was associated with shorter survival for all lung cancer patients regardless of the histologic subtype (5-year survival of 20.9% versus 49.5%, P < .001) in our series. In patients with atypical carcinoid tumors, lacking desmoglein 3 expression showed a 5-year survival of 0% compared with 36.8% for desmoglein 3-positive cases (P < .001). Desmoglein 3 status indicated a poor prognosis in lung cancers and portends a more aggressive behavior for atypical carcinoid tumors.

Adenocarcinoma↗

Histological staining methods preparatory to laser capture microdissection significantly affect the integrity of the cellular RNA.

BACKGROUND: Gene expression profiling by microarray analysis of cells enriched by laser capture microdissection (LCM) faces several technical challenges. Frozen sections yield higher quality RNA than paraffin-imbedded sections, but even with frozen sections, the staining methods used for histological identification of cells of interest could still damage the mRNA in the cells. To study the contribution of staining methods to degradation of results from gene expression profiling of LCM samples, we subjected pellets of the mouse plasma cell tumor cell line TEPC 1165 to direct RNA extraction and to parallel frozen sectioning for LCM and subsequent RNA extraction. We used microarray hybridization analysis to compare gene expression profiles of RNA from cell pellets with gene expression profiles of RNA from frozen sections that had been stained with hematoxylin and eosin (H&E), Nissl Stain (NS), and for immunofluorescence (IF) as well as with the plasma cell-revealing methyl green pyronin (MGP) stain. All RNAs were amplified with two rounds of T7-based in vitro transcription and analyzed by two-color expression analysis on 10-K cDNA microarrays. RESULTS: The MGP-stained samples showed the least introduction of mRNA loss, followed by H&E and immunofluorescence. Nissl staining was significantly more detrimental to gene expression profiles, presumably owing to an aqueous step in which RNA may have been damaged by endogenous or exogenous RNAases. CONCLUSION: RNA damage can occur during the staining steps preparatory to laser capture microdissection, with the consequence of loss of representation of certain genes in microarray hybridization analysis. Inclusion of RNAase inhibitor in aqueous staining solutions appears to be important in protecting RNA from loss of gene transcripts.

Animals↗

Case-control and case-only designs with genotype and family history data: estimating relative risk, residual familial aggregation, and cumulative risk.

In case-control studies of inherited diseases, participating subjects (probands) are often interviewed to collect detailed data about disease history and age-at-onset information in their family members. Genotype data are typically collected from the probands, but not from their relatives. In this article, we introduce an approach that combines case-control analysis of data on the probands with kin-cohort analysis of disease history data on relatives. Assuming a marginally specified multivariate survival model for joint risk of disease among family members, we describe methods for estimating relative risk, cumulative risk, and residual familial aggregation. We also describe a variation of the methodology that can be used for kin-cohort analysis of the family history data from a sample of genotyped cases only. We perform simulation studies to assess performance of the proposed methodologies with correct and mis-specified models for familial aggregation. We illustrate the proposed methodologies by estimating the risk of breast cancer from BRCA1/2 mutations using data from the Washington Ashkenazi Study.

BRCA1 Protein↗

Case-cohort designs and analysis for clustered failure time data.

Case-cohort design is an efficient and economical design to study risk factors for infrequent disease in a large cohort. It involves the collection of covariate data from all failures ascertained throughout the entire cohort, and from the members of a random subcohort selected at the onset of follow-up. In the literature, the case-cohort design has been extensively studied, but was exclusively considered for univariate failure time data. In this article, we propose case-cohort designs adapted to multivariate failure time data. An estimation procedure with the independence working model approach is used to estimate the regression parameters in the marginal proportional hazards model, where the correlation structure between individuals within a cluster is left unspecified. Statistical properties of the proposed estimators are developed. The performance of the proposed estimators and comparisons of statistical efficiencies are investigated with simulation studies. A data example from the Translating Research into Action for Diabetes (TRIAD) study is used to illustrate the proposed methodology.

Biometry↗

Evaluation of two phosphorylation sites improves the prognostic significance of Akt activation in non-small-cell lung cancer tumors.

PURPOSE: Akt is a serine/threonine kinase that has been implicated in lung tumorigenesis and lung cancer therapeutic resistance. Full activation of Akt requires two phosphorylation events, but only one site of phosphorylation (S473) has been evaluated thus far in clinical non-small-cell lung cancer (NSCLC) specimens, which has resulted in conflicting results regarding the prognostic significance of Akt activation in NSCLC. In this study, we sought to determine whether evaluation of Akt phosphorylation at T308 would improve prognostic accuracy. PATIENTS AND METHODS: Phosphospecific antibodies against T308 and S473 were validated and used in an immunohistochemical analysis of tissue microarray slides containing NSCLC specimens (n = 300) and surrounding lung tissue specimens (n = 100). RESULTS: Phosphorylation of either S473 or T308 was positive in most NSCSLC specimens, but was detected rarely in surrounding normal tissues. When Akt activation was defined by using both sites of phosphorylation, Akt activation was specific for NSCLC tumors versus surrounding tissue (73.4% v 0%; P < .05), was higher in adenocarcinoma than in squamous cell carcinoma (78.1% v 68.5%; P = .040), and was associated with shorter overall survival for all stages of disease (log-rank P = .041). In multivariate analyses, increased phosphorylation of T308 alone was a poor prognostic factor for stage I patients or for tumors < 5 cm (log-rank P = .011 and P = .015, respectively). CONCLUSION: These results suggest that monitoring phosphorylation of Akt at T308 improves the assessment of Akt activation, and show that Akt activation is a poor prognostic factor for all stages of NSCLC.

Adenocarcinoma↗

Whole genome expression profiling of advance stage papillary serous ovarian cancer reveals activated pathways.

Ovarian cancer is the most lethal type of gynecologic cancer in the Western world. The high case fatality rate is due in part because most ovarian cancer patients present with advanced stage disease which is essentially incurable. In order to obtain a whole genome assessment of aberrant gene expression in advanced ovarian cancer, we used oligonucleotide microarrays comprising over 40,000 features to profile 37 advanced stage papillary serous primary carcinomas. We identified 1191 genes that were significantly (P < 0.001) differentially regulated between the ovarian cancer specimens and normal ovarian surface epithelium. The microarray data were validated using real time RT-PCR on 14 randomly selected differentially regulated genes. The list of differentially expressed genes includes ones that are involved in cell growth, differentiation, adhesion, apoptosis and migration. In addition, numerous genes whose function remains to be elucidated were also identified. The microarray data were imported into PathwayAssist software to identify signaling pathways involved in ovarian cancer tumorigenesis. Based on our expression results, a signaling pathway associated with tumor cell migration, spread and invasion was identified as being activated in advanced ovarian cancer. The data generated in this study represent a comprehensive list of genes aberrantly expressed in serous papillary ovarian adenocarcinoma and may be useful for the identification of potentially new and novel markers and therapeutic targets for ovarian cancer.

Carcinoma, Papillary↗

Gene expression profiling identifies a unique androgen-mediated inflammatory/immune signature and a PTEN (phosphatase and tensin homolog deleted on chromosome 10)-mediated apoptotic response specific to the rat ventral prostate.

Understanding androgen regulation of gene expression is critical for deciphering mechanisms responsible for the transition from androgen-responsive (AR) to androgen-independent (AI) prostate cancer (PCa). To identify genes differentially regulated by androgens in each prostate lobe, the rat castration model was used. Microarray analysis was performed to compare dorsolateral (DLP) and ventral prostate (VP) samples from sham-castrated, castrated, and testosterone-replenished castrated rats. Our data demonstrate that, after castration, the VP and the DLP differed in the number of genes with altered expression (1496 in VP vs. 256 in DLP) and the nature of pathways modulated. Gene signatures related to apoptosis and immune response specific to the ventral prostate were identified. Microarray and RT-PCR analyses demonstrated the androgen repression of IGF binding protein-3 and -5, CCAAT-enhancer binding protein-delta, and phosphatase and tensin homolog deleted on chromosome 10 (PTEN) genes, previously implicated in apoptosis. We show that PTEN protein was increased only in the luminal epithelial cells of the VP, suggesting that it may be a key mediator of VP apoptosis in the absence of androgens. The castration-induced immune/inflammatory gene cluster observed specifically in the VP included IL-15 and IL-18. Immunostaining of the VP, but not the DLP, showed an influx of T cells, macrophages, and mast cells, suggesting that these cells may be the source of the immune signature genes. Interestingly, IL-18 was localized mainly to the basal epithelial cells and the infiltrating macrophages in the regressing VP, whereas IL-15 was induced in the luminal epithelium. The VP castration model exhibits immune cell infiltration and loss of PTEN that is often observed in progressive PCa, thereby making this model useful for further delineation of androgen-regulated gene expression with relevance to PCa.

Androgens↗

Effects of pooling mRNA in microarray class comparisons.

MOTIVATION: In microarray experiments investigators sometimes wish to pool RNA samples before labeling and hybridization due to insufficient RNA from each individual sample or to reduce the number of arrays for the purpose of saving cost. The basic assumption of pooling is that the expression of an mRNA molecule in the pool is close to the average expression from individual samples. Recently, a method for studying the effect of pooling mRNA on statistical power in detecting differentially expressed genes between classes has been proposed, but the different sources of variation arising in microarray experiments were not distinguished. Another paper recently did take different sources of variation into account, but did not address power and sample size for class comparison. In this paper, we study the implication of pooling in detecting differential gene expression taking into account different sources of variation and check the basic assumption of pooling using data from both the cDNA and Affymetrix GeneChip microarray experiments. RESULTS: We present formulas for the required number of subjects and arrays to achieve a desired power at a specified significance level. We show that due to the loss of degrees of freedom for a pooled design, a large increase in the number of subjects may be required to achieve a power comparable to that of a non-pooled design. The added expense of additional samples for the pooled design may outweigh the benefit of saving on microarray cost. The microarray data from both platforms show that the major assumption of pooling may not hold. SUPPLEMENTARY INFORMATION: Supplementary material referenced in the text is available at http://linus.nci.nih.gov/brb/TechReport.htm.

Algorithms↗

Chromatin remodeling factors and BRM/BRG1 expression as prognostic indicators in non-small cell lung cancer.

We immunohistochemically examined 12 core proteins involved in the chromatin remodeling machinery using a tissue microarray composed of 150 lung adenocarcinoma (AD) and 150 squamous cell carcinoma (SCC) cases. Most of the proteins showed nuclear staining, whereas some also showed cytoplasmic or membranous staining. When the expression patterns of all tested antigens were considered, proteins with nuclear staining clustered into two major groups. Nuclear signals of BRM, Ini-1, retinoblastoma, mSin3A, HDAC1, and HAT1 clustered together, whereas nuclear signals of BRG1, BAF155, HDAC2, BAF170, and RbAP48 formed a second cluster. Additionally, two thirds of the cases on the lung tissue array had follow-up information, and survival analysis was performed for each of the tested proteins. Positive nuclear BRM (N-BRM) staining correlated with a favorable prognosis in SCC and AD patients with a 5 year-survival of 53.5% compared with 32.3% for those whose tumors were negative for N-BRM (P = 0.015). Furthermore, patients whose tumors stained positive for both N-BRM and nuclear BRG1 had a 5 year-survival of 72% compared with 33.6% (P = 0.013) for those whose tumors were positive for either or negative for both markers. In contrast, membranous BRM (M-BRM) staining correlated with a poorer prognosis in AD patients with a 5 year-survival of 16.7% compared with those without M-BRM staining (38.1%; P = 0.016). These results support the notion that BRM and BRG1 participate in two distinct chromosome remodeling complexes that are functionally complementary and that the nuclear presence of BRM, its coexpression with nuclear BRG1, and the altered cellular localization of BRM (M-BRM) are useful markers for non-small cell lung cancer prognosis.

Adenocarcinoma↗

Predicting survival in patients with metastatic kidney cancer by gene-expression profiling in the primary tumor.

To identify potential molecular determinants of tumor biology and possible clinical outcomes, global gene-expression patterns were analyzed in the primary tumors of patients with metastatic renal cell cancer by using cDNA microarrays. We used grossly dissected tumor masses that included tumor, blood vessels, connective tissue, and infiltrating immune cells to obtain a gene-expression "profile" from each primary tumor. Two patterns of gene expression were found within this uniformly staged patient population, which correlated with a significant difference in overall survival between the two patient groups. Subsets of genes most significantly associated with survival were defined, and vascular cell adhesion molecule-1 (VCAM-1) was the gene most predictive for survival. Therefore, despite the complex biological nature of metastatic cancer, basic clinical behavior as defined by survival may be determined by the gene-expression patterns expressed within the compilation of primary gross tumor cells. We conclude that survival in patients with metastatic renal cell cancer can be correlated with the expression of various genes based solely on the expression profile in the primary kidney tumor.

Adult↗

Statistical issues in the design and analysis of gene expression microarray studies of animal models.

Appropriate statistical design and analysis of gene expression microarray studies is critical in order to draw valid and useful conclusions from expression profiling studies of animal models. In this paper, several aspects of study design are discussed, including the number of animals that need to be studied to ensure sufficiently powered studies, usefulness of replication and pooling, and allocation of samples to arrays. Data preprocessing methods for both cDNA dual-label spotted arrays and Affymetrix-style oligonucleotide arrays are reviewed. High-level analysis strategies are briefly discussed for each of the types of study aims, namely class comparison, class discovery, and class prediction. For class comparison, methods are discussed for identifying genes differentially expressed between classes while guarding against unacceptably high numbers of false positive findings. Various clustering methods are discussed for class discovery aims. Class prediction methods are briefly reviewed, and reference is made to the importance of proper validation of predictors.

Animals↗

Modeling tumor growth with random onset.

The longitudinal assessment of tumor volume is commonly used as an endpoint in small animal studies in cancer research. Groups of genetically identical mice are injected with mutant cells from clones developed with different mutations. The interest is on comparing tumor onset (i.e., the time of tumor detection) and tumor growth after onset, between mutation groups. This article proposes a class of linear and nonlinear growth models for jointly modeling tumor onset and growth in this situation. Our approach allows for interval-censored time of onset and missing-at-random dropout due to early sacrifice, which are common situations in animal research. We show that our approach has good small-sample properties for testing and is robust to some key unverifiable modeling assumptions. We illustrate this methodology with an application examining the effect of different mutations on tumorigenesis.

Animals↗

Analysis of survival data from case-control family studies.

In case-control family studies with survival endpoint, age of onset of diseases can be used to assess the familial aggregation of the disease and the relationship between the disease and genetic or environmental risk factors. Because of the retrospective nature of the case--control study, methods for analyzing prospectively collected correlated failure time data do not apply directly. In this article, we propose a semiparametric quasi-partial-likelihood approach to simultaneously estimate the effect of covariates on the age of onset and the association of ages of onset among family members that does not require specification of the baseline marginal distribution. We conducted a simulation study to evaluate the performance of the proposed approach and compare it with the existing semiparametric ones. Simulation results demonstrate that the proposed approach has better performance in terms of consistency and efficiency. We illustrate the methodology using a subset of data from the Washington Ashkenazi Study.

Age of Onset↗