PubMed Health⌕ Search

Biomedical subjects

Michael J Korenberg

Publications and source records attributed to Michael J Korenberg.

7 recordsLinked to original sources

On the advantages of multi-input single-output parallel cascade classifiers.

Parallel Cascade Identification (PCI) has been successfully applied to build dynamic nonlinear systems that address diverse challenges in the field of bioinformatics. PCI may be used to identify either single-input single-output (SISO) or multi-input single-output (MISO) models. Although SISO PCI models have typically sufficed, it has been suggested that MISO PCI systems could also be used to form bioinformatics classifiers, and indeed they were successfully applied in one study. This paper reports on the first systematic comparison of MISO and SISO PCI classifiers. Motivation for using the MISO structure is given. The construction of MISO parallel cascade models is also briefly reviewed. In order to compare the accuracy of SISO and MISO PCI classifiers, genetic algorithms are applied to optimize the model architecture on a number of equivalent single-input and multi-input biological training datasets. Through evaluation of both model structures on independent test datasets, we establish that MISO PCI is capable of building classifiers of equal accuracy to those resulting from SISO PCI models. Moreover, we discuss and illustrate the benefits of the MISO approach, including significant reduction in training and testing times, and the ability to adjust automatically the weighting of individual inputs according to information content.

Algorithms↗

Gene expression monitoring accurately predicts medulloblastoma positive and negative clinical outcomes.

Prediction of medulloblastoma clinical outcome is crucial to personalizing treatment, both to identify high-risk patients for aggressive or alternative therapy and to spare those at low risk from excessive treatment. The best predictors [Pomeroy et al. (2002) Nature 415, 436-442], based on gene expression monitoring at diagnosis, have shown much less accuracy in recognizing patients with eventual failed outcomes - <50% for the predictor making fewest total errors - than those who would survive, while a single gene predictor exhibited reverse asymmetry. Such inaccuracy in recognizing one of the outcomes is a problem for clinical use. We hypothesized that a non-linear model could be built to significantly improve prediction of medulloblastoma outcome, thereby promoting use of gene-expression-based predictors in a clinical setting. In fact, this approach resulted in fewer errors and much less asymmetry in prediction, and bidirectional accuracy of about 80% could be obtained via its combination with other methods. Indeed, three combinations of methods were identified that yielded significantly better predictions of clinical outcome than previously attained, making feasible predictors of medulloblastoma treatment response with greatly improved bidirectional accuracy essential for clinical use.

Cerebellar Neoplasms↗

Recognition of adenosine triphosphate binding sites using parallel cascade system identification.

Parallel cascade identification (PCI) is a method for approximating the behavior of a nonlinear system, from input/output training data, by constructing a parallel array of cascaded dynamic linear and static nonlinear elements. PCI has previously been shown to provide an effective means for classifying protein sequences into structure/function families. In the present study, PCI is used to distinguish proteins that are binding to adenosine triphosphate or guanine triphosphate molecules from those that are nonbinding. Classification accuracy of 87.1% using the hydrophobicity scale of Rose et al. (Hydrophobicity of amino acid residues in globular proteins. Science 229:834-838, 1985), and 88.8% using Korenberg's SARAH1 scale, are obtained, as measured by tenfold cross-validation testing. Nearest-neighbor and K-nearest-neighbor (KNN) classifiers are constructed, and the resulting accuracy is, respectively, 88.0% and 90.8% on the SARAH1-encoded test data set, as measured by the above testing protocol. Significantly improved classification accuracy is achieved by combining PCI and KNN classifiers using quadratic discriminant analysis: accuracy rises from 87.9% (PCI) and 87.4% (KNN) to 96.5% for the combination, as measured by twofold cross-validation testing on the SARAH1-encoded test data set.

Adenosine Triphosphate↗

Using the fast orthogonal search with first term reselection to find subharmonic terms in spectral analysis.

The fast orthogonal search (FOS) algorithm has been shown to accurately model various types of time series by implicitly creating a specialized orthogonal basis set to fit the desired time series. When the data contain periodic components, FOS can find frequencies with a resolution greater than the discrete Fourier transform (DFT) algorithm. Frequencies with less than one period in the record length, called subharmonic frequencies, and frequencies between the bins of a DFT, can be resolved. This paper considers the resolution of subharmonic frequencies using the FOS algorithm. A new criterion for determining the number of non-noise terms in the model is introduced. This new criterion does not assume the first model term fitted is a dc component as did the previous stopping criterion. An iterative FOS algorithm called FOS first-term reselection (FOS-FTR), is introduced. FOS-FTR reduces the mean-square error of the sinusoidal model and selects the subharmonic frequencies more accurately than does the unmodified FOS algorithm.

Algorithms↗

Parallel cascade recognition of exon and intron DNA sequences.

Many of the current procedures for detecting coding regions on human DNA sequences combine a number of individual techniques such as discriminant analysis and neural net methods. Recent papers have used techniques from nonlinear systems identification, in particular, parallel cascade identification (PCI), as one means for classifying protein sequences into their structure/function groups. In the present paper, PCI is used in a pilot study to distinguish exon (coding) from intron (noncoding; interspersed within genes) human DNA sequences. Only the first exon and first intron sequences with known boundaries in genomic DNA from the beta T-cell receptor locus were used for training. Then, the parallel cascade classifiers were able to achieve classification rates of about 89% on novel sequences in a test set, and averaged about 82% when results of a blind test were included. In testing over a much wider range of human nucleotide sequences, PCI classifiers averaged 83.6% correct classifications. These results indicate that parallel cascade classifiers may be useful components in future coding region detection programs.

Algorithms↗

Prediction of treatment response using gene expression profiles.

This paper concerns prediction of clinical outcome from gene expression profiles using work in a different area, nonlinear system identification. In particular, the approach can predict long-term treatment response from data of a landmark article by Golub et al. (Golub, T. R.; Slonim, D. K.; Tamayo, P.; Huard, C.; Gaasenbeek, M.; Mesirov, J. P. et al. Science 1999, 286, 531-537) that has not previously been achieved with these data. The present paper shows that, for these data, gene expression profiles taken at time of diagnosis of acute myeloid leukemia contain information predictive of eventual response to chemotherapy. This was not evident in previous work; indeed, the Golub et al. article did not find a set of genes strongly correlated with clinical outcome. However, the present approach can accurately predict outcome class of gene expression profiles even when the genes do not have large differences in expression levels between the classes.

Antibiotics, Antineoplastic↗

On predicting medulloblastoma metastasis by gene expression profiling.

Accurately predicting clinical outcome or metastatic status from gene expression profiles remains one of the biggest hurdles facing the adoption of predictive medicine. Recently, MacDonald et al. (Nat. Genet. 2001, 29, 143-152) used gene expression profiles, from samples taken at diagnosis, to distinguish between clinically designated metastatic and nonmetastatic primary medulloblastomas, helping to elucidate the genetic mechanisms underlying metastasis and suggesting novel therapeutic targets. The obtained accuracy of predicting metastatic status does not, however, reach statistical significance on Fisher's exact test, although 22 training samples were used to make each prediction via leave-one-out testing. This paper introduces readily implemented nonlinear filters to transform sequences of gene expression levels into output signals that are significantly easier to classify and predict metastasis. It is shown that when only 3 exemplars each from the metastatic and nonmetastatic classes were assumed known, a predictor was constructed whose accuracy is statistically significant over the remaining profiles set aside as a test set. The predictor was as effective in recognizing metastatic as nonmetastatic medulloblastomas, and may be helpful in deciding which patients require more aggressive therapy. The same predictor was similarly effective on an independent set of 5 nonmetastatic tumors and 3 metastatic cell lines also used by MacDonald et al.

Gene Expression Profiling↗