PubMed Health⌕ Search

Biomedical subjects

Yudi Pawitan

Publications and source records attributed to Yudi Pawitan.

At least 19 recordsLinked to original sources

Protein Biomarkers in Risk and Prognosis of Amyotrophic Lateral Sclerosis.

BACKGROUND: Plasma and cerebrospinal fluid (CSF) protein biomarkers in amyotrophic lateral sclerosis (ALS) may provide insight into disease mechanisms and yield clinically useful biomarkers. METHODS: Overall, 363 proteins in plasma and CSF from 198 patients with ALS and 125 matched controls were profiled using Olink assays. Associations with disease status, survival, and functional decline, as well as longitudinal biomarker stability across the disease course were assessed, together with network and enrichment analyses. ALS risk-associated biomarkers were externally validated in the UK Biobank (UKB). RESULTS: Overall, 125 proteins were significantly associated with at least one outcome (i.e., case status, risk, survival, or functional decline), and 21 were associated with three or more outcomes. NEFL was the most robust biomarker in plasma and CSF, alongside TNFRSF12A in plasma and CSF, EDA2R in plasma, and FABP4 in plasma and CSF. Most biomarkers remained stable longitudinally across the disease course. ALS risk-associated biomarkers were replicated in UKB, in which > 3000 plasma proteins were measured in 52,990 participants, including 298 with ALS. Network and enrichment analyses highlighted their roles in immune response and extracellular-matrix remodeling, and their enrichments in the brain and T-cell subsets. Construction of an ALS risk-prediction model achieved an ROC-AUC of 0.72 in the UKB validation cohort. CONCLUSIONS: These findings suggest candidate protein biomarkers for ALS risk stratification, early detection, and clinical therapeutic monitoring.

Humans↗

Gene expression in 16q is associated with survival and differs between Sørlie breast cancer subtypes.

We have investigated the relationship between gene expression and chromosomal positions in 402 breast cancer patients. Using an overrepresentation approach based on Fisher's exact test, we identified disproportionate contributions of specific chromosomal positions to genes associated with survival. Our major finding is that the gene expression in the long arm of chromosome 16 stands out in its relationship to survival. This arm contributes 36 (18%) and 55 (11%) genes to lists negatively associated with recurrence-free survival (set to sizes 200 and 500). This is a highly disproportionate contribution from the 313 (2%) genes in this arm represented on the used Affymetrix U133A and B microarray platforms (Bonferroni corrected Fisher test: P < 2.2 x 10(-16)). We also demonstrate differential expression in 16q across tumor subtypes, which suggests that the ERBB2, basal, and luminal B tumors progress along a high grade-poor prognosis path, while luminal A and normal-like tumors progress along a low grade-good prognosis path, in accordance with a previously proposed model of tumor progression. We conclude that important biological information can be extracted from gene expression data in breast cancer by studying non-random connections between chromosomal positions and gene expression. This article contains Supplementary Material available at http://www.interscience.wiley.com/jpages/1045-2257/suppmat.

Breast Neoplasms↗

Genetic reclassification of histologic grade delineates new clinical subtypes of breast cancer.

Histologic grading of breast cancer defines morphologic subtypes informative of metastatic potential, although not without considerable interobserver disagreement and clinical heterogeneity particularly among the moderately differentiated grade 2 (G2) tumors. We posited that a gene expression signature capable of discerning tumors of grade 1 (G1) and grade 3 (G3) histology might provide a more objective measure of grade with prognostic benefit for patients with G2 disease. To this end, we studied the expression profiles of 347 primary invasive breast tumors analyzed on Affymetrix microarrays. Using class prediction algorithms, we identified 264 robust grade-associated markers, six of which could accurately classify G1 and G3 tumors, and separate G2 tumors into two highly discriminant classes (termed G2a and G2b genetic grades) with patient survival outcomes highly similar to those with G1 and G3 histology, respectively. Statistical analysis of conventional clinical variables further distinguished G2a and G2b subtypes from each other, but also from histologic G1 and G3 tumors. In multivariate analyses, genetic grade was consistently found to be an independent prognostic indicator of disease recurrence comparable with that of lymph node status and tumor size. When incorporated into the Nottingham prognostic index, genetic grade enhanced detection of patients with less harmful tumors, likely to benefit little from adjuvant therapy. Our findings show that a genetic grade signature can improve prognosis and therapeutic planning for breast cancer patients, and support the view that low- and high-grade disease, as defined genetically, reflect independent pathobiological entities rather than a continuum of cancer progression.

Algorithms↗

Estimation of false discovery proportion under general dependence.

MOTIVATION: Wide-scale correlations between genes are commonly observed in gene expression data, due to both biological and technical reasons. These correlations increase the variability of the standard estimate of the false discovery rate (FDR). We highlight the false discovery proportion (FDP, instead of the FDR) as the suitable quantity for assessing differential expression in microarray data, demonstrate the deleterious effects of correlation on FDP estimation and propose an improved estimation method that accounts for the correlations. METHODS: We analyse the variation pattern of the distribution of test statistics under permutation using the singular value decomposition. The results suggest a latent FDR model that accounts for the effects of correlation, and is statistically closer to the FDP. We develop a procedure for estimating the latent FDR (ELF) based on a Poisson regression model. RESULTS: For simulated data based on the correlation structure of real datasets, we find that ELF performs substantially better than the standard FDR approach in estimating the FDP. We illustrate the use of ELF in the analysis of breast cancer and lymphoma data. AVAILABILITY: R code to perform ELF is available in http://www.meb.ki.se/~yudpaw.

Algorithms↗

Parental age and risk of childhood cancers: a population-based cohort study from Sweden.

BACKGROUND: Frequent germ line cells mutations were previously demonstrated to be associated with aging. This suggests a higher incidence of childhood cancer among children of older parents. A population-based cohort study of parental ages and other prenatal risk factors for five main childhood cancers was performed with the use of a linkage between several national-based registries. METHODS: In total, about 4.3 million children with their parents, born between 1961 and 2000, were included in the study. Multivariate Poisson regression was used to obtain the incidence rate ratios (IRR) and 95% confidence interval (CI). Children <5 years of age and children 5-14 years of age were analysed independently. RESULTS: There was no significant result for children 5-14 years of age. For children <5 years of age, maternal age were associated with elevated risk of retinoblastoma (oldest age group's IRR = 2.39, 95%CI = 1.17-4.85) and leukaemia (oldest age group's IRR = 1.44, 95%CI = 1.01-2.05). Paternal age was significantly associated with leukaemia (oldest age group's IRR = 1.31, 95%CI = 1.04-1.66). For central nervous system cancer, the effect of paternal age was found to be significant (oldest age group's IRR = 1.69, 95%CI = 1.21-2.35) when maternal age was included in the analysis. CONCLUSION: Our findings indicate that advanced parental age might be associated with an increased risk of early childhood cancers.

Adolescent↗

Hormone-replacement therapy influences gene expression profiles and is associated with breast-cancer prognosis: a cohort study.

BACKGROUND: Postmenopausal hormone-replacement therapy (HRT) increases breast-cancer risk. The influence of HRT on the biology of the primary tumor, however, is not well understood. METHODS: We obtained breast-cancer gene expression profiles using Affymetrix human genome U133A arrays. We examined the relationship between HRT-regulated gene profiles, tumor characteristics, and recurrence-free survival in 72 postmenopausal women. RESULTS: HRT use in patients with estrogen receptor (ER) protein positive tumors (n = 72) was associated with an altered regulation of 276 genes. Expression profiles based on these genes clustered ER-positive tumors into two molecular subclasses, one of which was associated with HRT use and had significantly better recurrence free survival despite lower ER levels. A comparison with external data suggested that gene regulation in tumors associated with HRT was negatively correlated with gene regulation induced by short-term estrogen exposure, but positively correlated with the effect of tamoxifen. CONCLUSION: Our findings suggest that post-menopausal HRT use is associated with a distinct gene expression profile related to better recurrence-free survival and lower ER protein levels. Tentatively, HRT-associated gene expression in tumors resembles the effect of tamoxifen exposure on MCF-7 cells.

Breast Neoplasms↗

Tobacco use, body mass index and the risk of malignant lymphomas--a nationwide cohort study in Sweden.

In the search for risk factors involved in the etiology of lymphoproliferative malignancies there is still inconsistent evidence regarding effects of smoking tobacco, and the role of smokeless tobacco is poorly investigated. New evidence indicates that excess body weight increases the risk of NHL and HD. To determine if tobacco use of various forms and high Body Mass Index (BMI) affect the occurrence of these neoplasms, we conducted a prospective cohort study on over 330,000 Swedish construction workers included in the Construction Industry Working Environment and Health program. Information on smoking, snuff dipping, height and weight was gathered by self administered questionnaires together with personal interviews. Cancer incidence was ascertained through the year 2000 by record linkage to the nationwide Swedish Cancer Registry, Migration Registry and Cause of Death Registry. At the end of follow up, 1,309 subjects had been diagnosed with NHL (including chronic lymphocytic leukemia) and 205 with HD respectively. Age adjusted incidence rate ratios were computed using Cox proportional Hazard regression modeling. Smoking cigarette, pipe or cigar was not associated with NHL or HD. There was no evidence indicating a relation between quantity and duration of smoking and NHL or HD risk. No link was found between NHL and usage of smokeless tobacco. Having a BMI of 30 or higher did not convey excess risk of developing NHL or HD compared to normal weight (BMI 18.6-24.9). We conclude that tobacco smoking and high BMI do not entail an increased risk of NHL and HD. Our findings of a relation between the duration of snuff dipping and HD need further investigation.

Adolescent↗

Finding regions of significance in SELDI measurements for identifying protein biomarkers.

MOTIVATION: There is a well-recognized potential of protein expression profiling using the surface-enhanced laser desorption and ionization technology for discovering biomarkers that can be applied in clinical diagnosis, prognosis and therapy prediction. The pre-processing of the raw data, however, is still problematic. METHODS: We focus on the peak detection step, where the standard method is marked by poor specificity. Currently, scientists need to inspect individual spectra visually and laboriously in order to verify that spectral peaks identified by the standard method are real. Motivated by this multi-spectral process, we investigate an analytical approach-called RS for 'regions of significance'-that reduces the data to a single spectrum of F-statistics capturing significant variability between spectra. To account for multiple testing, we use a false discovery rate criterion for identifying potentially interesting proteins. RESULTS: We show that RS has better operating characteristics than several existing methods and demonstrate routine applications on a number of large datasets.

Algorithms↗

Annotated regions of significance of SELDI-TOF-MS spectra for detecting protein biomarkers.

Peak detection is a key step in the analysis of SELDI-TOF-MS spectra, but the current default method has low specificity and poor peak annotation. To improve data quality, scientists still have to validate the identified peaks visually, a tedious and time-consuming process, especially for large data sets. Hence, there is a genuine need for methods that minimize manual validation. We have previously reported a multi-spectral signal detection method, called RS for 'region of significance', with improved specificity. Here we extend it to include a peak quantification algorithm based on annotated regions of significance (ARS). For each spectral region flagged as significant by RS, we first identify a dominant spectrum for determining the number of peaks and the m/z region of these peaks. From each m/z region of peaks, a peak template is extracted from all spectra via the principal component analysis. Finally, with the template, we estimate the amplitude and location of the peak in each spectrum with the least-squares method and refine the estimation of the amplitude via the mixture model. We have evaluated the ARS algorithm on patient samples from a clinical study. Comparison with the standard method shows that ARS (i) inherits the superior specificity of RS, and (ii) gives more accurate peak annotations than the standard method. In conclusion, we find that ARS alleviates the main problems in the preprocessing of SELDI-TOF spectra. The R-package ProSpect that implements ARS is freely available for academic use at http://www.meb.ki.se/ yudpaw.

Adenocarcinoma↗

Familial aggregation of small-for-gestational-age births: the importance of fetal genetic effects.

OBJECTIVE: This study was undertaken to disentangle the maternal genetic, fetal genetic, and environmental effects for the risk of having small-for-gestational-age (SGA) offspring. STUDY DESIGN: By cross-linking the population-based Swedish Multi-Generation and Medical Birth Registers, we extracted 2,193,142 births between 1973 and 2001 with both parents identified. Odds ratios (OR) were calculated to estimate the relative risks, and generalized linear mixed models were used to estimate the contribution of genetic and environmental effects. RESULTS: Women whose full sisters had an offspring born SGA had a significantly increased risk of having a SGA offspring themselves (OR = 1.8, 95% CI 1.7-1.9), whereas the corresponding risk for brothers was lower (OR = 1.3, 95% CI 1.2-1.4). Thirty-seven percent of the liability was explained by fetal (including both maternal and paternal) genetic effects and 9% by maternal genetic effects. CONCLUSION: Genetic factors account for almost half of the liability to have SGA births. These effects are primarily caused by fetal genes.

Female↗

Genomic instability and prognosis in breast carcinomas.

BACKGROUND: We recently reported that DNA content of breast adenocarcinomas, cytometrically assessed by diploid (D), tetraploid (T), and aneuploid (A) categories, can be further divided into genomically stable and unstable subtypes by means of the stemline scatter index (SSI). The aim of the present study was to survey the clinical correlates and the prognostic value of the SSI in a consecutive series of 890 breast cancer patients. RESULTS: Genomically stable subtype had a significantly better survival compared with the unstable subtype within each ploidy category: D (P = 0.04), T (P = 0.008), and A (P = 0.004). By contrast, no statistically significant difference in survival was observed between the D, T, and A categories within the stable (P = 0.23) and unstable subtypes (P = 0.12). Among A tumors, the unstable subtype tended to be larger, more frequently estrogen- and progesterone-receptor negative, and to be of higher grade compared with the stable subtype. Stable D tumors tended to have lower grade than the unstable subtype, but among the D and T tumors, genomic instability was not associated with receptor status. Within the Elston grade 3, lymph node-positive or estrogen receptor-positive subgroups, patients with stable tumors had significantly better survival compared with unstable tumors (P = 0.01, 0.002, and 7.2E-5, respectively). CONCLUSIONS: The SSI contributes supplementary biological and clinical information in addition to ploidy information alone. Objective classification of breast adenocarcinomas into stable and unstable subtypes is a useful prognostic indicator independent of established clinical factors.

Adult↗

Intrinsic molecular signature of breast cancer in a population-based cohort of 412 patients.

BACKGROUND: Molecular markers and the rich biological information they contain have great potential for cancer diagnosis, prognostication and therapy prediction. So far, however, they have not superseded routine histopathology and staging criteria, partly because the few studies performed on molecular subtyping have had little validation and limited clinical characterization. METHODS: We obtained gene expression and clinical data for 412 breast cancers obtained from population-based cohorts of patients from Stockholm and Uppsala, Sweden. Using the intrinsic set of approximately 500 genes derived in the Norway/Stanford breast cancer data, we validated the existence of five molecular subtypes--basal-like, ERBB2, luminal A/B and normal-like--and characterized these subtypes extensively with the use of conventional clinical variables. RESULTS: We found an overall 77.5% concordance between the centroid prediction of the Swedish cohort by using the Norway/Stanford signature and the k-means clustering performed internally within the Swedish cohort. The highest rate of discordant assignments occurred between the luminal A and luminal B subtypes and between the luminal B and ERBB2 subtypes. The subtypes varied significantly in terms of grade (p < 0.001), p53 mutation (p < 0.001) and genomic instability (p = 0.01), but surprisingly there was little difference in lymph-node metastasis (p = 0.31). Furthermore, current users of hormone-replacement therapy were strikingly over-represented in the normal-like subgroup (p < 0.001). Separate analyses of the patients who received endocrine therapy and those who did not receive any adjuvant therapy supported the previous hypothesis that the basal-like subtype responded to adjuvant treatment, whereas the ERBB2 and luminal B subtypes were poor responders. CONCLUSION: We found that the intrinsic molecular subtypes of breast cancer are broadly present in a diverse collection of patients from a population-based cohort in Sweden. The intrinsic gene set, originally selected to reveal stable tumor characteristics, was shown to have a strong correlation with progression-related properties such as grade, p53 mutation and genomic instability.

Breast Neoplasms↗

Multidimensional local false discovery rate for microarray studies.

MOTIVATION: The false discovery rate (fdr) is a key tool for statistical assessment of differential expression (DE) in microarray studies. Overall control of the fdr alone, however, is not sufficient to address the problem of genes with small variance, which generally suffer from a disproportionally high rate of false positives. It is desirable to have an fdr-controlling procedure that automatically accounts for gene variability. METHODS: We generalize the local fdr as a function of multiple statistics, combining a common test statistic for assessing DE with its standard error information. We use a non-parametric mixture model for DE and non-DE genes to describe the observed multi-dimensional statistics, and estimate the distribution for non-DE genes via the permutation method. We demonstrate this fdr2d approach for simulated and real microarray data. RESULTS: The fdr2d allows objective assessment of DE as a function of gene variability. We also show that the fdr2d performs better than commonly used modified test statistics. AVAILABILITY: An R-package OCplus containing functions for computing fdr2d() and other operating characteristics of microarray data is available at http://www.meb.ki.se/~yudpaw.

Algorithms↗

Gene expression profiling spares early breast cancer patients from adjuvant therapy: derived and validated in two population-based cohorts.

INTRODUCTION: Adjuvant breast cancer therapy significantly improves survival, but overtreatment and undertreatment are major problems. Breast cancer expression profiling has so far mainly been used to identify women with a poor prognosis as candidates for adjuvant therapy but without demonstrated value for therapy prediction. METHODS: We obtained the gene expression profiles of 159 population-derived breast cancer patients, and used hierarchical clustering to identify the signature associated with prognosis and impact of adjuvant therapies, defined as distant metastasis or death within 5 years. Independent datasets of 76 treated population-derived Swedish patients, 135 untreated population-derived Swedish patients and 78 Dutch patients were used for validation. The inclusion and exclusion criteria for the studies of population-derived Swedish patients were defined. RESULTS: Among the 159 patients, a subset of 64 genes was found to give an optimal separation of patients with good and poor outcomes. Hierarchical clustering revealed three subgroups: patients who did well with therapy, patients who did well without therapy, and patients that failed to benefit from given therapy. The expression profile gave significantly better prognostication (odds ratio, 4.19; P = 0.007) (breast cancer end-points odds ratio, 10.64) compared with the Elston-Ellis histological grading (odds ratio of grade 2 vs 1 and grade 3 vs 1, 2.81 and 3.32 respectively; P = 0.24 and 0.16), tumor stage (odds ratio of stage 2 vs 1 and stage 3 vs 1, 1.11 and 1.28; P = 0.83 and 0.68) and age (odds ratio, 0.11; P = 0.55). The risk groups were consistent and validated in the independent Swedish and Dutch data sets used with 211 and 78 patients, respectively. CONCLUSION: We have identified discriminatory gene expression signatures working both on untreated and systematically treated primary breast cancer patients with the potential to spare them from adjuvant therapy.

Adult↗

Fold-change estimation of differentially expressed genes using mixture mixed-model.

Microarray experiments produce expression measurements for thousands of genes simultaneously, though usually for a small number of RNA samples. The most common problem is the identification of genes that are differentially expressed between different groups of samples or biological conditions. As the number of genes far exceeds the number of RNA samples, the inherent multiplicity poses a severe problem in both hypothesis testing and effect estimation. While much of the recent literature is focused on the hypothesis aspects, we concentrate in this paper on effect estimation as a tool for the identification of differentially expressed genes. We propose a linear mixed model where the random effects are assumed to follow a mixture distribution, and study in detail the case of three normals, corresponding to genes that are down-, up- or non regulated. Our approach leads to a new type of non-linear shrinkage estimation, where a proportion of estimates is shrunk to zero, while the rest follows standard linear shrinkage. This allows us to estimate the log fold-change of the genes involved and to identify those that are differentially expressed within the same model framework. We investigate the operating characteristics of our method using simulation and spike-in studies, and illustrate its application to real data using a breast-cancer dataset.

Journal Article↗

An expression signature for p53 status in human breast cancer predicts mutation status, transcriptional effects, and patient survival.

Perturbations of the p53 pathway are associated with more aggressive and therapeutically refractory tumors. However, molecular assessment of p53 status, by using sequence analysis and immunohistochemistry, are incomplete assessors of p53 functional effects. We posited that the transcriptional fingerprint is a more definitive downstream indicator of p53 function. Herein, we analyzed transcript profiles of 251 p53-sequenced primary breast tumors and identified a clinically embedded 32-gene expression signature that distinguishes p53-mutant and wild-type tumors of different histologies and outperforms sequence-based assessments of p53 in predicting prognosis and therapeutic response. Moreover, the p53 signature identified a subset of aggressive tumors absent of sequence mutations in p53 yet exhibiting expression characteristics consistent with p53 deficiency because of attenuated p53 transcript levels. Our results show the primary importance of p53 functional status in predicting clinical breast cancer behavior.

Breast Neoplasms↗

Bias in the estimation of false discovery rate in microarray studies.

MOTIVATION: The false discovery rate (FDR) provides a key statistical assessment for microarray studies. Its value depends on the proportion pi(0) of non-differentially expressed (non-DE) genes. In most microarray studies, many genes have small effects not easily separable from non-DE genes. As a result, current methods often overestimate pi(0) and FDR, leading to unnecessary loss of power in the overall analysis. METHODS: For the common two-sample comparison we derive a natural mixture model of the test statistic and an explicit bias formula in the standard estimation of pi(0). We suggest an improved estimation of pi(0) based on the mixture model and describe a practical likelihood-based procedure for this purpose. RESULTS: The analysis shows that a large bias occurs when pi(0) is far from 1 and when the non-centrality parameters of the distribution of the test statistic are near zero. The theoretical result also explains substantial discrepancies between non-parametric and model-based estimates of pi(0). Simulation studies indicate mixture-model estimates are less biased than standard estimates. The method is applied to breast cancer and lymphoma data examples. AVAILABILITY: An R-package OCplus containing functions to compute pi(0) based on the mixture model, the resulting FDR and other operating characteristics of microarray data, is freely available at http://www.meb.ki.se/~yudpaw CONTACT: yudi.pawitan@meb.ki.se and alexander.ploner@meb.ki.se.

Computer Simulation↗

False discovery rate, sensitivity and sample size for microarray studies.

MOTIVATION: In microarray data studies most researchers are keenly aware of the potentially high rate of false positives and the need to control it. One key statistical shift is the move away from the well-known P-value to false discovery rate (FDR). Less discussion perhaps has been spent on the sensitivity or the associated false negative rate (FNR). The purpose of this paper is to explain in simple ways why the shift from P-value to FDR for statistical assessment of microarray data is necessary, to elucidate the determining factors of FDR and, for a two-sample comparative study, to discuss its control via sample size at the design stage. RESULTS: We use a mixture model, involving differentially expressed (DE) and non-DE genes, that captures the most common problem of finding DE genes. Factors determining FDR are (1) the proportion of truly differentially expressed genes, (2) the distribution of the true differences, (3) measurement variability and (4) sample size. Many current small microarray studies are plagued with large FDR, but controlling FDR alone can lead to unacceptably large FNR. In evaluating a design of a microarray study, sensitivity or FNR curves should be computed routinely together with FDR curves. Under certain assumptions, the FDR and FNR curves coincide, thus simplifying the choice of sample size for controlling the FDR and FNR jointly.

Algorithms↗