PubMed Health⌕ Search

Biomedical subjects

Liat Ein-Dor

Publications and source records attributed to Liat Ein-Dor.

3 recordsLinked to original sources

Thousands of samples are needed to generate a robust gene list for predicting outcome in cancer.

Predicting at the time of discovery the prognosis and metastatic potential of cancer is a major challenge in current clinical research. Numerous recent studies searched for gene expression signatures that outperform traditionally used clinical parameters in outcome prediction. Finding such a signature will free many patients of the suffering and toxicity associated with adjuvant chemotherapy given to them under current protocols, even though they do not need such treatment. A reliable set of predictive genes also will contribute to a better understanding of the biological mechanism of metastasis. Several groups have published lists of predictive genes and reported good predictive performance based on them. However, the gene lists obtained for the same clinical types of patients by different groups differed widely and had only very few genes in common. This lack of agreement raised doubts about the reliability and robustness of the reported predictive gene lists, and the main source of the problem was shown to be the small number of samples that were used to generate the gene lists. Here, we introduce a previously undescribed mathematical method, probably approximately correct (PAC) sorting, for evaluating the robustness of such lists. We calculate for several published data sets the number of samples that are needed to achieve any desired level of reproducibility. For example, to achieve a typical overlap of 50% between two predictive lists of genes, breast cancer studies would need the expression profiles of several thousand early discovery patients.

Breast Neoplasms↗

Outcome signature genes in breast cancer: is there a unique set?

MOTIVATION: Predicting the metastatic potential of primary malignant tissues has direct bearing on the choice of therapy. Several microarray studies yielded gene sets whose expression profiles successfully predicted survival. Nevertheless, the overlap between these gene sets is almost zero. Such small overlaps were observed also in other complex diseases, and the variables that could account for the differences had evoked a wide interest. One of the main open questions in this context is whether the disparity can be attributed only to trivial reasons such as different technologies, different patients and different types of analyses. RESULTS: To answer this question, we concentrated on a single breast cancer dataset, and analyzed it by a single method, the one which was used by van't Veer et al. to produce a set of outcome-predictive genes. We showed that, in fact, the resulting set of genes is not unique; it is strongly influenced by the subset of patients used for gene selection. Many equally predictive lists could have been produced from the same analysis. Three main properties of the data explain this sensitivity: (1) many genes are correlated with survival; (2) the differences between these correlations are small; (3) the correlations fluctuate strongly when measured over different subsets of patients. A possible biological explanation for these properties is discussed. CONTACT: eytan.domany@weizmann.ac.il SUPPLEMENTARY INFORMATION: http://www.weizmann.ac.il/physics/complex/compphys/downloads/liate/

Biomarkers, Tumor↗

Low autocorrelated multiphase sequences.

The interplay between the ground-state energy of the generalized Bernasconi model to multiphase, and the minimal value of the maximal autocorrelation function, C(max)=max(K)/C(K)/, K=1,...,N-1, is examined analytically in the thermodynamic limit where the main results are (a) For the binary case, the minimal value of C(max) over all sequences of length N, minC(max), is 0.435sqrt[N], significantly smaller than the typical value for random sequences O(sqrt[log N]sqrt[N]). (b) A new method to approximate F(max) is obtained using the observation of data collapse. (c) minC(max) is obtained in an energy which is about 30% above the ground-state energy of the generalized Bernasconi model, independent of the number of phases m. (d) For a given m, minC(max) proportional sqrt[N/m] indicating that for m=N, minC(max)=1, i.e., a generalized Barker code exists. The analytical results are confirmed by simulations.

Journal Article↗