PubMed Health⌕ Search

Biomedical subjects

Robert Tibshirani

Publications and source records attributed to Robert Tibshirani.

At least 19 recordsLinked to original sources

The use of plasma surface-enhanced laser desorption/ionization time-of-flight mass spectrometry proteomic patterns for detection of head and neck squamous cell cancers.

PURPOSE: Our study was undertaken to determine the utility of plasma proteomic profiling using surface-enhanced laser desorption/ionization time-of-flight (SELDI-TOF) mass spectrometry for the detection of head and neck squamous cell carcinomas (HNSCCs). EXPERIMENTAL DESIGN: Pretreatment plasma samples from HNSCC patients or controls without known neoplastic disease were analyzed on the Protein Biology System IIc SELDI-TOF mass spectrometer (Ciphergen Biosystems, Fremont, CA). Proteomic spectra of mass:charge ratio (m/z) were generated by the application of plasma to immobilized metal-affinity-capture (IMAC) ProteinChip arrays activated with copper. A total of 37356 data points were generated for each sample. A training set of spectra from 56 cancer patients and 52 controls were applied to the "Lasso" technique to identify protein profiles that can distinguish cancer from noncancer, and cross-validation was used to determine test errors in this training set. The discovery pattern was then used to classify a separate masked test set of 57 cancer and 52 controls. In total, we analyzed the proteomic spectra of 113 cancer patients and 104 controls. RESULTS: The Lasso approach identified 65 significant data points for the discrimination of normal from cancer profiles. The discriminatory pattern correctly identified 39 of 57 HNSCC patients and 40 of 52 noncancer controls in the masked test set. These results yielded a sensitivity of 68% and specificity of 73%. Subgroup analyses in the test set of four different demographic factors (age, gender, and cigarette and alcohol use) that can potentially confound the interpretation of the results suggest that this model tended to overpredict cancer in control smokers. CONCLUSIONS: Plasma proteomic profiling with SELDI-TOF mass spectrometry provides moderate sensitivity and specificity in discriminating HNSCC. Further improvement and validation of this approach is needed to determine its usefulness in screening for this disease.

Adult↗

Gene expression profiles at diagnosis in de novo childhood AML patients identify FLT3 mutations with good clinical outcomes.

Fms-like tyrosine kinase 3 (FLT3) mutations are associated with unfavorable outcomes in children with acute myeloid leukemia (AML). We used DNA microarrays to identify gene expression profiles related to FLT3 status and outcome in childhood AML. Among 81 diagnostic specimens, 36 had FLT3 mutations (FLT3-MUs), 24 with internal tandem duplications (ITDs) and 12 with activating loop mutations (ALMs). In addition, 8 of 19 specimens from patients with relapses had FLT3-MUs. Predictive analysis of microarrays (PAM) identified genes that differentiated FLT3-ITD from FLT3-ALM and FLT3 wild-type (FLT3-WT) cases. Among the 42 specimens with FLT3-MUs, PAM identified 128 genes that correlated with clinical outcome. Event-free survival (EFS) in FLT3-MU patients with a favorable signature was 45% versus 5% for those with an unfavorable signature (P = .018). Among FLT3-MU specimens, high expression of the RUNX3 gene and low expression of the ATRX gene were associated with inferior outcome. The ratio of RUNX3 to ATRX expression was used to classify FLT3-MU cases into 3 EFS groups: 70%, 37%, and 0% for low, intermediate, and high ratios, respectively (P < .0001). Thus, gene expression profiling identified AML patients with divergent prognoses within the FLT3-MU group, and the RUNX3 to ATRX expression ratio should be a useful prognostic indicator in these patients.

Acute Disease↗

Sample classification from protein mass spectrometry, by 'peak probability contrasts'.

MOTIVATION: Early cancer detection has always been a major research focus in solid tumor oncology. Early tumor detection can theoretically result in lower stage tumors, more treatable diseases and ultimately higher cure rates with less treatment-related morbidities. Protein mass spectrometry is a potentially powerful tool for early cancer detection. We propose a novel method for sample classification from protein mass spectrometry data. When applied to spectra from both diseased and healthy patients, the 'peak probability contrast' technique provides a list of all common peaks among the spectra, their statistical significance and their relative importance in discriminating between the two groups. We illustrate the method on matrix-assisted laser desorption and ionization mass spectrometry data from a study of ovarian cancers. RESULTS: Compared to other statistical approaches for class prediction, the peak probability contrast method performs as well or better than several methods that require the full spectra, rather than just labelled peaks. It is also much more interpretable biologically. The peak probability contrast method is a potentially useful tool for sample classification from protein mass spectrometry data.

Algorithms↗

Toxicity from radiation therapy associated with abnormal transcriptional responses to DNA damage.

Toxicity from radiation therapy is a grave problem for cancer patients. We hypothesized that some cases of toxicity are associated with abnormal transcriptional responses to radiation. We used microarrays to measure responses to ionizing and UV radiation in lymphoblastoid cells derived from 14 patients with acute radiation toxicity. The analysis used heterogeneity-associated transformation of the data to account for a clinical outcome arising from more than one underlying cause. To compute the risk of toxicity for each patient, we applied nearest shrunken centroids, a method that identifies and cross-validates predictive genes. Transcriptional responses in 24 genes predicted radiation toxicity in 9 of 14 patients with no false positives among 43 controls (P = 2.2 x 10(-7)). The responses of these nine patients displayed significant heterogeneity. Of the five patients with toxicity and normal responses, two were treated with protocols that proved to be highly toxic. These results may enable physicians to predict toxicity and tailor treatment for individual patients.

Adult↗

Use of gene-expression profiling to identify prognostic subclasses in adult acute myeloid leukemia.

BACKGROUND: In patients with acute myeloid leukemia (AML), the presence or absence of recurrent cytogenetic aberrations is used to identify the appropriate therapy. However, the current classification system does not fully reflect the molecular heterogeneity of the disease, and treatment stratification is difficult, especially for patients with intermediate-risk AML with a normal karyotype. METHODS: We used complementary-DNA microarrays to determine the levels of gene expression in peripheral-blood samples or bone marrow samples from 116 adults with AML (including 45 with a normal karyotype). We used unsupervised hierarchical clustering analysis to identify molecular subgroups with distinct gene-expression signatures. Using a training set of samples from 59 patients, we applied a novel supervised learning algorithm to devise a gene-expression-based clinical-outcome predictor, which we then tested using an independent validation group comprising the 57 remaining patients. RESULTS: Unsupervised analysis identified new molecular subtypes of AML, including two prognostically relevant subgroups in AML with a normal karyotype. Using the supervised learning algorithm, we constructed an optimal 133-gene clinical-outcome predictor, which accurately predicted overall survival among patients in the independent validation group (P=0.006), including the subgroup of patients with AML with a normal karyotype (P=0.046). In multivariate analysis, the gene-expression predictor was a strong independent prognostic factor (odds ratio, 8.8; 95 percent confidence interval, 2.6 to 29.3; P<0.001). CONCLUSIONS: The use of gene-expression profiling improves the molecular classification of adult AML.

Acute Disease↗

Semi-supervised methods to predict patient survival from gene expression data.

An important goal of DNA microarray research is to develop tools to diagnose cancer more accurately based on the genetic profile of a tumor. There are several existing techniques in the literature for performing this type of diagnosis. Unfortunately, most of these techniques assume that different subtypes of cancer are already known to exist. Their utility is limited when such subtypes have not been previously identified. Although methods for identifying such subtypes exist, these methods do not work well for all datasets. It would be desirable to develop a procedure to find such subtypes that is applicable in a wide variety of circumstances. Even if no information is known about possible subtypes of a certain form of cancer, clinical information about the patients, such as their survival time, is often available. In this study, we develop some procedures that utilize both the gene expression data and the clinical data to identify subtypes of cancer and use this knowledge to diagnose future patients. These procedures were successfully applied to several publicly available datasets. We present diagnostic procedures that accurately predict the survival of future patients based on the gene expression profile and survival times of previous patients. This has the potential to be a powerful tool for diagnosing and treating cancer.

Breast Neoplasms↗

Cancer characterization and feature set extraction by discriminative margin clustering.

BACKGROUND: A central challenge in the molecular diagnosis and treatment of cancer is to define a set of molecular features that, taken together, distinguish a given cancer, or type of cancer, from all normal cells and tissues. RESULTS: Discriminative margin clustering is a new technique for analyzing high dimensional quantitative datasets, specially applicable to gene expression data from microarray experiments related to cancer. The goal of the analysis is find highly specialized sub-types of a tumor type which are similar in having a small combination of genes which together provide a unique molecular portrait for distinguishing the sub-type from any normal cell or tissue. Detection of the products of these genes can then, in principle, provide a basis for detection and diagnosis of a cancer, and a therapy directed specifically at the distinguishing constellation of molecular features can, in principle, provide a way to eliminate the cancer cells, while minimizing toxicity to any normal cell. CONCLUSIONS: The new methodology yields highly specialized tumor subtypes which are similar in terms of potential diagnostic markers.

Cluster Analysis↗

Gene expression profiling identifies clinically relevant subtypes of prostate cancer.

Prostate cancer, a leading cause of cancer death, displays a broad range of clinical behavior from relatively indolent to aggressive metastatic disease. To explore potential molecular variation underlying this clinical heterogeneity, we profiled gene expression in 62 primary prostate tumors, as well as 41 normal prostate specimens and nine lymph node metastases, using cDNA microarrays containing approximately 26,000 genes. Unsupervised hierarchical clustering readily distinguished tumors from normal samples, and further identified three subclasses of prostate tumors based on distinct patterns of gene expression. High-grade and advanced stage tumors, as well as tumors associated with recurrence, were disproportionately represented among two of the three subtypes, one of which also included most lymph node metastases. To further characterize the clinical relevance of tumor subtypes, we evaluated as surrogate markers two genes differentially expressed among tumor subgroups by using immunohistochemistry on tissue microarrays representing an independent set of 225 prostate tumors. Positive staining for MUC1, a gene highly expressed in the subgroups with "aggressive" clinicopathological features, was associated with an elevated risk of recurrence (P = 0.003), whereas strong staining for AZGP1, a gene highly expressed in the other subgroup, was associated with a decreased risk of recurrence (P = 0.0008). In multivariate analysis, MUC1 and AZGP1 staining were strong predictors of tumor recurrence independent of tumor grade, stage, and preoperative prostate-specific antigen levels. Our results suggest that prostate tumors can be usefully classified according to their gene expression patterns, and these tumor subtypes may provide a basis for improved prognostication and treatment stratification.

Biomarkers, Tumor↗

Efficient quadratic regularization for expression arrays.

Gene expression arrays typically have 50 to 100 samples and 1000 to 20,000 variables (genes). There have been many attempts to adapt statistical models for regression and classification to these data, and in many cases these attempts have challenged the computational resources. In this article we expose a class of techniques based on quadratic regularization of linear models, including regularized (ridge) regression, logistic and multinomial regression, linear and mixture discriminant analysis, the Cox model and neural networks. For all of these models, we show that dramatic computational savings are possible over naive implementations, using standard transformations in numerical linear algebra.

Data Interpretation, Statistical↗

Gene expression patterns in ovarian carcinomas.

We used DNA microarrays to characterize the global gene expression patterns in surface epithelial cancers of the ovary. We identified groups of genes that distinguished the clear cell subtype from other ovarian carcinomas, grade I and II from grade III serous papillary carcinomas, and ovarian from breast carcinomas. Six clear cell carcinomas were distinguished from 36 other ovarian carcinomas (predominantly serous papillary) based on their gene expression patterns. The differences may yield insights into the worse prognosis and therapeutic resistance associated with clear cell carcinomas. A comparison of the gene expression patterns in the ovarian cancers to published data of gene expression in breast cancers revealed a large number of differentially expressed genes. We identified a group of 62 genes that correctly classified all 125 breast and ovarian cancer specimens. Among the best discriminators more highly expressed in the ovarian carcinomas were PAX8 (paired box gene 8), mesothelin, and ephrin-B1 (EFNB1). Although estrogen receptor was expressed in both the ovarian and breast cancers, genes that are coregulated with the estrogen receptor in breast cancers, including GATA-3, LIV-1, and X-box binding protein 1, did not show a similar pattern of coexpression in the ovarian cancers.

Adenocarcinoma↗

Statistical significance for genomewide studies.

With the increase in genomewide experiments and the sequencing of multiple genomes, the analysis of large data sets has become commonplace in biology. It is often the case that thousands of features in a genomewide data set are tested against some null hypothesis, where a number of features are expected to be significant. Here we propose an approach to measuring statistical significance in these genomewide studies based on the concept of the false discovery rate. This approach offers a sensible balance between the number of true and false positives that is automatically calibrated and easily interpreted. In doing so, a measure of statistical significance called the q value is associated with each tested feature. The q value is similar to the well known p value, except it is a measure of significance in terms of the false discovery rate rather than the false positive rate. Our approach avoids a flood of false positive results, while offering a more liberal criterion than what has been used in genome scans for linkage.

Algorithms↗

Repeated observation of breast tumor subtypes in independent gene expression data sets.

Characteristic patterns of gene expression measured by DNA microarrays have been used to classify tumors into clinically relevant subgroups. In this study, we have refined the previously defined subtypes of breast tumors that could be distinguished by their distinct patterns of gene expression. A total of 115 malignant breast tumors were analyzed by hierarchical clustering based on patterns of expression of 534 "intrinsic" genes and shown to subdivide into one basal-like, one ERBB2-overexpressing, two luminal-like, and one normal breast tissue-like subgroup. The genes used for classification were selected based on their similar expression levels between pairs of consecutive samples taken from the same tumor separated by 15 weeks of neoadjuvant treatment. Similar cluster analyses of two published, independent data sets representing different patient cohorts from different laboratories, uncovered some of the same breast cancer subtypes. In the one data set that included information on time to development of distant metastasis, subtypes were associated with significant differences in this clinical feature. By including a group of tumors from BRCA1 carriers in the analysis, we found that this genotype predisposes to the basal tumor subtype. Our results strongly support the idea that many of these breast tumor subtypes represent biologically distinct disease entities.

Breast Neoplasms↗

HGAL is a novel interleukin-4-inducible gene that strongly predicts survival in diffuse large B-cell lymphoma.

We have cloned and characterized a novel human gene, HGAL (human germinal center-associated lymphoma), which predicts outcome in patients with diffuse large B-cell lymphoma (DLBCL). The HGAL gene comprises 6 exons and encodes a cytoplasmic protein of 178 amino acids that contains an immunoreceptor tyrosine-based activation motif (ITAM). It is highly expressed in germinal center (GC) lymphocytes and GC-derived lymphomas and is homologous to the mouse GC-specific gene M17. Expression of the HGAL gene is specifically induced in B cells by interleukin-4 (IL-4). Patients with DLBCL expressing high levels of HGAL mRNA demonstrate significantly longer overall survival than do patients with low HGAL expression. This association was independent of the clinical international prognostic index. High HGAL mRNA expression should be used as a prognostic factor in DLBCL.

B-Lymphocytes↗

Characterization of variant patterns of nodular lymphocyte predominant hodgkin lymphoma with immunohistologic and clinical correlation.

Nodular lymphocyte predominant Hodgkin lymphoma (NLPHL) has traditionally been recognized as having two morphologic patterns, nodular and diffuse, and the current WHO definition of NLPHL requires at least a partial nodular pattern. Variant patterns have not been well documented. We analyzed retrospectively the morphologic and immunophenotypic patterns of NLPHL from 118 patients (total of 137 biopsy samples). Histology plus antibodies directed against CD20, CD3, and CD21 were used to evaluate the immunoarchitecture. We identified six distinct immunoarchitectural patterns in our cases of NLPHL: "classic" (B-cell-rich) nodular, serpiginous/interconnected nodular, nodular with prominent extranodular L&H cells, T-cell-rich nodular, diffuse with a T-cell-rich background (T-cell-rich B-cell lymphoma [TCRBCL]-like), and a (diffuse) B-cell-rich pattern. Small germinal centers within neoplastic nodules were found in approximately 15% of cases, a finding not previously emphasized in NLPHL. Prominent sclerosis was identified in approximately 20% of cases and was frequently seen in recurrent disease. Clinical follow-up was obtained on 56 patients, including 26 patients who had not had recurrence of disease and 30 patients who had recurrence. The follow-up period was 5 months to 16 years (median 2.5 years). The presence of a diffuse (TCRBCL-like) pattern was significantly more common in patients with recurrent disease than those without recurrence. Furthermore, the presence of a diffuse pattern (TCRBCL-like) was shown to be an independent predictor of recurrent disease (P = 0.00324). In addition, there is a tendency for progression to an increasingly more diffuse pattern over time. Analysis of sequential biopsies from patients with recurrent disease suggests that the presence of prominent extranodular L&H cells might represent early evolution to a diffuse (TCRBCL-like) pattern. We also report three patients who presented initially with diffuse large B-cell lymphoma and later developed NLPHL.

Adolescent↗

Microarray analysis reveals a major direct role of DNA copy number alteration in the transcriptional program of human breast tumors.

Genomic DNA copy number alterations are key genetic events in the development and progression of human cancers. Here we report a genome-wide microarray comparative genomic hybridization (array CGH) analysis of DNA copy number variation in a series of primary human breast tumors. We have profiled DNA copy number alteration across 6,691 mapped human genes, in 44 predominantly advanced, primary breast tumors and 10 breast cancer cell lines. While the overall patterns of DNA amplification and deletion corroborate previous cytogenetic studies, the high-resolution (gene-by-gene) mapping of amplicon boundaries and the quantitative analysis of amplicon shape provide significant improvement in the localization of candidate oncogenes. Parallel microarray measurements of mRNA levels reveal the remarkable degree to which variation in gene copy number contributes to variation in gene expression in tumor cells. Specifically, we find that 62% of highly amplified genes show moderately or highly elevated expression, that DNA copy number influences gene expression across a wide range of DNA copy number alterations (deletion, low-, mid- and high-level amplification), that on average, a 2-fold change in DNA copy number is associated with a corresponding 1.5-fold change in mRNA levels, and that overall, at least 12% of all the variation in gene expression among the breast tumors is directly attributable to underlying variation in gene copy number. These findings provide evidence that widespread DNA copy number alteration can lead directly to global deregulation of gene expression, which may contribute to the development or progression of cancer.

Breast Neoplasms↗

Transcriptional programs activated by exposure of human prostate cancer cells to androgen.

BACKGROUND: Androgens are required for both normal prostate development and prostate carcinogenesis. We used DNA microarrays, representing approximately 18,000 genes, to examine the temporal program of gene expression following treatment of the human prostate cancer cell line LNCaP with a synthetic androgen. RESULTS: We observed statistically significant changes in levels of transcripts of more than 500 genes. Many of these genes were previously reported androgen targets, but most were not previously known to be regulated by androgens. The androgen-induced expression programs in three additional androgen-responsive human prostate cancer cell lines, and in four androgen-independent subclones derived from LNCaP, shared many features with those observed in LNCaP, but some differences were observed. A remarkable fraction of the genes induced by androgen appeared to be related to production of seminal fluid and these genes included many with roles in protein folding, trafficking, and secretion. CONCLUSIONS: Prostate cancer cell lines retain features of androgen responsiveness that reflect normal prostatic physiology. These results provide a broad view of the effect of androgen signaling on the transcriptional program in these cancer cells, and a foundation for further studies of androgen action.

Androgens↗