PubMed Health⌕ Search

Biomedical subjects

Sudhir Varma

Publications and source records attributed to Sudhir Varma.

6 recordsLinked to original sources

SCLC TumorMiner: A genomics platform for small cell lung cancer precision oncology.

Small cell lung cancer (SCLC) is among the most aggressive malignancies. Unlike many other cancers, it is not represented in The Cancer Genome Atlas, and available datasets are fragmented across institutions, disease stages, and treatment settings. RNA sequencing provides a powerful and cost-effective approach, but the high dimensionality of transcriptomic data and the heterogeneity of patient cohorts pose significant challenges. To address such challenges, we developed SCLC TumorMiner (https://discover.nci.nih.gov/SclcTumorMinerCDB/), which includes 50 tumor samples from relapsed patients at the National Cancer Institute (NCI) and 154 samples from untreated patients at the University of Cologne and Tongji University. SCLC TumorMiner enables molecular classification, genomic pathway analyses, risk stratification, identification of predictive cell-surface biomarkers such as DLL3 or TROP2, and drug-response biomarkers such as SLFN11. SCLC TumorMiner illustrates profound differences between untreated and relapsed patient samples. Additionally, "MyPatient", one of SCLC TumorMiner's modules, is presented as a medical assistant application prototype.

SCLC↗

Gene expression profiling reveals a massive, aneuploidy-dependent transcriptional deregulation and distinct differences between lymph node-negative and lymph node-positive colon carcinomas.

To characterize patterns of global transcriptional deregulation in primary colon carcinomas, we did gene expression profiling of 73 tumors [Unio Internationale Contra Cancrum stage II (n = 33) and stage III (n = 40)] using oligonucleotide microarrays. For 30 of the tumors, expression profiles were compared with those from matched normal mucosa samples. We identified a set of 1,950 genes with highly significant deregulation between tumors and mucosa samples (P < 1e-7). A significant proportion of these genes mapped to chromosome 20 (P = 0.01). Seventeen genes had a >5-fold average expression difference between normal colon mucosa and carcinomas, including up-regulation of MYC and of HMGA1, a putative oncogene. Furthermore, we identified 68 genes that were significantly differentially expressed between lymph node-negative and lymph node-positive tumors (P < 0.001), the functional annotation of which revealed a preponderance of genes that play a role in cellular immune response and surveillance. The microarray-derived gene expression levels of 20 deregulated genes were validated using quantitative real-time reverse transcription-PCR in >40 tumor and normal mucosa samples with good concordance between the techniques. Finally, we established a relationship between specific genomic imbalances, which were mapped for 32 of the analyzed colon tumors by comparative genomic hybridization, and alterations of global transcriptional activity. Previously, we had conducted a similar analysis of primary rectal carcinomas. The systematic comparison of colon and rectal carcinomas revealed a significant overlap of genomic imbalances and transcriptional deregulation, including activation of the Wnt/beta-catenin signaling cascade, suggesting similar pathogenic pathways.

Adenocarcinoma↗

Bias in error estimation when using cross-validation for model selection.

BACKGROUND: Cross-validation (CV) is an effective method for estimating the prediction error of a classifier. Some recent articles have proposed methods for optimizing classifiers by choosing classifier parameter values that minimize the CV error estimate. We have evaluated the validity of using the CV error estimate of the optimized classifier as an estimate of the true error expected on independent data. RESULTS: We used CV to optimize the classification parameters for two kinds of classifiers; Shrunken Centroids and Support Vector Machines (SVM). Random training datasets were created, with no difference in the distribution of the features between the two classes. Using these "null" datasets, we selected classifier parameter values that minimized the CV error estimate. 10-fold CV was used for Shrunken Centroids while Leave-One-Out-CV (LOOCV) was used for the SVM. Independent test data was created to estimate the true error. With "null" and "non null" (with differential expression between the classes) data, we also tested a nested CV procedure, where an inner CV loop is used to perform the tuning of the parameters while an outer CV is used to compute an estimate of the error. The CV error estimate for the classifier with the optimal parameters was found to be a substantially biased estimate of the true error that the classifier would incur on independent data. Even though there is no real difference between the two classes for the "null" datasets, the CV error estimate for the Shrunken Centroid with the optimal parameters was less than 30% on 18.5% of simulated training data-sets. For SVM with optimal parameters the estimated error rate was less than 30% on 38% of "null" data-sets. Performance of the optimized classifiers on the independent test set was no better than chance. The nested CV procedure reduces the bias considerably and gives an estimate of the error that is very close to that obtained on the independent testing set for both Shrunken Centroids and SVM classifiers for "null" and "non-null" data distributions. CONCLUSION: We show that using CV to compute an error estimate for a classifier that has itself been tuned using CV gives a significantly biased estimate of the true error. Proper use of CV for estimating true error of a classifier developed using a well defined algorithm requires that all steps of the algorithm, including classifier parameter tuning, be repeated in each CV loop. A nested CV procedure provides an almost unbiased estimate of the true error.

Algorithms↗

Aneuploidy-dependent massive deregulation of the cellular transcriptome and apparent divergence of the Wnt/beta-catenin signaling pathway in human rectal carcinomas.

To identify genetic alterations underlying rectal carcinogenesis, we used global gene expression profiling of a series of 17 locally advanced rectal adenocarcinomas and 20 normal rectal mucosa biopsies on oligonucleotide arrays. A total of 351 genes were differentially expressed (P < 1.0e-7) between normal rectal mucosa and rectal carcinomas, 77 genes had a >5-fold difference, and 85 genes always had at least a 2-fold change in all of the matched samples. Twelve genes satisfied all three of these criteria. Altered expression of genes such as PTGS2 (COX-2), WNT1, TGFB1, VEGF, and MYC was confirmed, whereas our data for other genes, like PPARD and LEF1, were inconsistent with previous reports. In addition, we found deregulated expression of many genes whose involvement in rectal carcinogenesis has not been reported. By mapping the genomic imbalances in the tumors using comparative genomic hybridization, we could show that DNA copy number gains of recurrently aneuploid chromosome arms 7p, 8q, 13q, 18q, 20p, and 20q correlated significantly with their average chromosome arm expression profile. Taken together, our results show that both the high-level, significant transcriptional deregulation of specific genes and general modification of the average transcriptional activity of genes residing on aneuploid chromosomes coexist in rectal adenocarcinomas.

Adenocarcinoma↗

Effectiveness of gene expression profiling for response prediction of rectal adenocarcinomas to preoperative chemoradiotherapy.

PURPOSE: There is a wide spectrum of tumor responsiveness of rectal adenocarcinomas to preoperative chemoradiotherapy ranging from complete response to complete resistance. This study aimed to investigate whether parallel gene expression profiling of the primary tumor can contribute to stratification of patients into groups of responders or nonresponders. PATIENTS AND METHODS: Pretherapeutic biopsies from 30 locally advanced rectal carcinomas were analyzed for gene expression signatures using microarrays. All patients were participants of a phase III clinical trial (CAO/ARO/AIO-94, German Rectal Cancer Trial) and were randomized to receive a preoperative combined-modality therapy including fluorouracil and radiation. Class comparison was used to identify a set of genes that were differentially expressed between responders and nonresponders as measured by T level downsizing and histopathologic tumor regression grading. RESULTS: In an initial set of 23 patients, responders and nonresponders showed significantly different expression levels for 54 genes (P < .001). The ability to predict response to therapy using gene expression profiles was rigorously evaluated using leave-one-out cross-validation. Tumor behavior was correctly predicted in 83% of patients (P = .02). Sensitivity (correct prediction of response) was 78%, and specificity (correct prediction of nonresponse) was 86%, with a positive and negative predictive value of 78% and 86%, respectively. CONCLUSION: Our results suggest that pretherapeutic gene expression profiling may assist in response prediction of rectal adenocarcinomas to preoperative chemoradiotherapy. The implementation of gene expression profiles for treatment stratification and clinical management of cancer patients requires validation in large, independent studies, which are now warranted.

Adenocarcinoma↗

Iterative class discovery and feature selection using Minimal Spanning Trees.

BACKGROUND: Clustering is one of the most commonly used methods for discovering hidden structure in microarray gene expression data. Most current methods for clustering samples are based on distance metrics utilizing all genes. This has the effect of obscuring clustering in samples that may be evident only when looking at a subset of genes, because noise from irrelevant genes dominates the signal from the relevant genes in the distance calculation. RESULTS: We describe an algorithm for automatically detecting clusters of samples that are discernable only in a subset of genes. We use iteration between Minimal Spanning Tree based clustering and feature selection to remove noise genes in a step-wise manner while simultaneously sharpening the clustering. Evaluation of this algorithm on synthetic data shows that it resolves planted clusters with high accuracy in spite of noise and the presence of other clusters. It also shows a low probability of detecting spurious clusters. Testing the algorithm on some well known micro-array data-sets reveals known biological classes as well as novel clusters. CONCLUSIONS: The iterative clustering method offers considerable improvement over clustering in all genes. This method can be used to discover partitions and their biological significance can be determined by comparing with clinical correlates and gene annotations. The MATLAB programs for the iterative clustering algorithm are available from http://linus.nci.nih.gov/supplement.html

Acute Disease↗