PubMed Health⌕ Search

Biomedical subjects

Olli Yli-Harja

Publications and source records attributed to Olli Yli-Harja.

At least 19 recordsLinked to original sources

Evaluating the performance of microarray segmentation algorithms.

MOTIVATION: Although numerous algorithms have been developed for microarray segmentation, extensive comparisons between the algorithms have acquired far less attention. In this study, we evaluate the performance of nine microarray segmentation algorithms. Using both simulated and real microarray experiments, we overcome the challenges in performance evaluation, arising from the lack of ground-truth information. The usage of simulated experiments allows us to analyze the segmentation accuracy on a single pixel level as is commonly done in traditional image processing studies. With real experiments, we indirectly measure the segmentation performance, identify significant differences between the algorithms, and study the characteristics of the resulting gene expression data. RESULTS: Overall, our results show clear differences between the algorithms. The results demonstrate how the segmentation performance depends on the image quality, which algorithms operate on significantly different performance levels, and how the selection of a segmentation algorithm affects the identification of differentially expressed genes. AVAILABILITY: Supplementary results and the microarray images used in this study are available at the companion web site http://www.cs.tut.fi/sgn/csb/spotseg/

Algorithms↗

Iterated maps for annealed Boolean networks.

Boolean networks are used to study the large-scale properties of nonlinear systems and are mainly applied to model genetic regulatory networks. A statistical method called the annealed approximation is commonly used to examine the dynamical properties of randomly generated Boolean networks that are created with selected statistical features. However, in the literature there are several variations of the annealed approximation. These approximations cannot be interchangeably used in all cases due to different background assumptions. In this paper, we present the so-called four-state model, derive the different approximations from this model, and make the differences and connections between these approximations explicit. As an application of the presented results, we study the properties of the Boolean networks that are constructed with random functions, canalizing functions, and regulatory functions found in the biological literature.

Journal Article↗

Differential gene expression in non-malignant tumour microenvironment is associated with outcome in follicular lymphoma patients treated with rituximab and CHOP.

Rituximab in combination with chemotherapy (immunochemotherapy) is one of the most effective treatments available for follicular lymphoma (FL). This study aimed to determine whether differences in gene expression in FL tissue correlate with outcome in response to rituximab and CHOP (cyclophosphamide, doxorubicin, vincristine, prednisone) chemotherapy (R-CHOP). We divided 24 patients into long- [time to treatment failure (TTF) >35 months] and short-term (TTF <23 months) responders, and analysed the gene expression profiles of lymphoma tissue using oligonucleotide microarrays. We used a supervised learning technique to identify genes correlating with outcome, and confirmed the expression of selected genes with quantitative polymerase chain reaction (qPCR) and immunohistochemistry. Among the transcripts with a high correlation between microarray and qPCR analyses, we identified EPHA1, a tyrosine kinase involved in transepithelial migration, SMAD1, a transcription factor and a mediator of bone morphogenetic protein and transforming growth factor-beta signalling, and MARCO, a scavenger receptor on macrophages. According to Kaplan-Meier estimates, high EPHA1, and low SMAD1 and MARCO expression were associated with better progression-free survival (PFS). Immunohistochemistry showed that EphA1 was primarily localised in granulocytes. In addition, both EphA1 and Smad1 were expressed in vascular endothelia. However, no difference in vasculature was detected between long- and short-term responders. In a validation set of 40 patients, a trend towards a better PFS was observed among patients with high EphA1 expression. We conclude that gene expression in non-malignant cells contributes to clinical outcome in R-CHOP-treated FL patients.

Adult↗

Simulation of microarray data with realistic characteristics.

BACKGROUND: Microarray technologies have become common tools in biological research. As a result, a need for effective computational methods for data analysis has emerged. Numerous different algorithms have been proposed for analyzing the data. However, an objective evaluation of the proposed algorithms is not possible due to the lack of biological ground truth information. To overcome this fundamental problem, the use of simulated microarray data for algorithm validation has been proposed. RESULTS: We present a microarray simulation model which can be used to validate different kinds of data analysis algorithms. The proposed model is unique in the sense that it includes all the steps that affect the quality of real microarray data. These steps include the simulation of biological ground truth data, applying biological and measurement technology specific error models, and finally simulating the microarray slide manufacturing and hybridization. After all these steps are taken into account, the simulated data has realistic biological and statistical characteristics. The applicability of the proposed model is demonstrated by several examples. CONCLUSION: The proposed microarray simulation model is modular and can be used in different kinds of applications. It includes several error models that have been proposed earlier and it can be used with different types of input data. The model can be used to simulate both spotted two-channel and oligonucleotide based single-channel microarrays. All this makes the model a valuable tool for example in validation of data analysis algorithms.

Algorithms↗

Unsupervised analysis uncovers changes in histopathologic diagnosis in supervised genomic studies.

Human gastrointestinal stromal tumors (GIST) have recently emerged as a distinct mesenchymal tumor type that has a unique phenotype characterized by a gain of function mutations in c-kit. In contrast, leiomyosarcomas (LMS) of the gastrointestinal tract or retroperitoneum, which were previously classified together with GISTs as gastrointestinal sarcomas, have much less frequent mutations of c-kit. We performed microarray analyses to gain a comprehensive understanding of the difference between the two types of soft-tissue sarcomas at the level of gene expression. Microarray experiments were performed on 30 GISTs and 30 LMSs that were collected at the time of surgical resection. These tumors were categorized based on the histopathologic diagnosis recorded in our institutional database. Prior to our search for genes that are differentially expressed between these two types of cancers, we first carried out an unsupervised analysis using multidimensional scaling (MDS) to determine whether the two groups have marked overall differences in gene expression. Initially, the MDS did not reveal a good separation between the two groups. We then re-reviewed the histopathology of these tumors and realized that some of the cases included in our study were acquired 10 years ago when the diagnosis of gastrointestinal sarcoma was made according to histopathologic criteria alone without immunohistochemistry for c-kit. An experienced pathologist reviewed all of the specimens and this revealed that a number of the GIST cases were classified as LMS in the clinical database. Correction of the histopathologic diagnosis and relabeling of the samples resulted in a much more pronounced separation of GIST and LMS in the MDS analysis. This study underscores the need to re-review histopathology as reclassification occurs. While updating the clinical database may be desired, this is usually impractical. For molecular studies that use archival samples, it is critical to have the archival samples re-reviewed by a pathologist. Further, unsupervised analysis often proves to be a critical quality control step in identifying structural problems that may exist. Finally, MDS analysis further supports that GIST is a distinct type of sarcoma.

Biomarkers, Tumor↗

Analysis of angiogenesis using in vitro experiments and stochastic growth models.

The global properties of vascular networks grown with an in vitro angiogenesis assay are compared quantitatively, using automated image analysis, with the global properties of networks obtained with discrete, stochastic growth models. The model classes that are investigated are invasion percolation and diffusion limited aggregation. By matching global properties to experimental data, one can infer which model classes and parameters are most reflective of angiogenesis in experimental cells. This sheds light on large-scale emergent properties of angiogenesis from a systems perspective. It is found that invasion percolation is better than diffusion limited aggregation at matching experimental data. We also present evidence that the distribution of the lengths of real tubule complexes follows a power law.

Animals↗

Quantification of vesicles in differentiating human SH-SY5Y neuroblastoma cells by automated image analysis.

A new automated image analysis method for quantification of fluorescent dots is presented. This method facilitates counting the number of fluorescent puncta in specific locations of individual cells and also enables estimation of the number of cells by detecting the labeled nuclei. The method is here used for counting the AM1-43 labeled fluorescent puncta in human SH-SY5Y neuroblastoma cells induced to differentiate with all-trans retinoic acid (RA), and further stimulated with high potassium (K+) containing solution. The automated quantification results correlate well with the results obtained manually through visual inspection. The manual method has the disadvantage of being slow, labor-intensive, and subjective, and the results may not be reproducible even in the intra-observer case. The automated method, however, has the advantage of allowing fast quantification with explicitly defined methods, with no user intervention. This ensures objectivity of the quantification. In addition to the number of fluorescent dots, further development of the method allows its use for quantification of several other parameters, such as intensity, size, and shape of the puncta, that are difficult to quantify manually.

Algorithms↗

Global gene expression profile of human cord blood-derived CD133+ cells.

Human cord blood (CB)-derived CD133+ cells carry characteristics of primitive hematopoietic cells and proffer an alternative for CD34+ cells in hematopoietic stem cell (HSC) transplantation. To characterize the CD133+ cell population on a genetic level, a global expression analysis of CD133+ cells was performed using oligonucleotide microarrays. CD133+ cells were purified from four fresh CB units by immunomagnetic selection. All four CD133+ samples showed significant similarity in their gene expression pattern, whereas they differed clearly from the CD133- control samples. In all, 690 transcripts were differentially expressed between CD133+ and CD133- cells. Of these, 393 were increased and 297 were decreased in CD133+ cells. The highest overexpression was noted in genes associated with metabolism, cellular physiological processes, cell communication, and development. A set of 257 transcripts expressed solely in the CD133+ cell population was identified. Colony-forming unit (CFU) assay was used to detect the clonal progeny of precursors present in the studied cell populations. The results demonstrate that CD133+ cells express primitive markers and possess clonogenic progenitor capacity. This study provides a gene expression profile for human CD133+ cells. It presents a set of genes that may be used to unravel the properties of the CD133+ cell population, assumed to be highly enriched in HSCs.

AC133 Antigen↗

Tracking perturbations in Boolean networks with spectral methods.

In this paper we present a method for predicting the spread of perturbations in Boolean networks. The method is applicable to networks that have no regular topology. The prediction of perturbations can be performed easily by using a presented result which enables the efficient computation of the required iterative formulas. This result is based on abstract Fourier transform of the functions in the network. In this paper the method is applied to show the spread of perturbations in networks containing a distribution of functions found from biological data. The advances in the study of the spread of perturbations can directly be applied to enable ways of quantifying chaos in Boolean networks. Derrida plots over an arbitrary number of time steps can be computed and thus distributions of functions compared with each other with respect to the amount of order they create in random networks.

Journal Article↗

Robust detection of periodic time series measured from biological systems.

BACKGROUND: Periodic phenomena are widespread in biology. The problem of finding periodicity in biological time series can be viewed as a multiple hypothesis testing of the spectral content of a given time series. The exact noise characteristics are unknown in many bioinformatics applications. Furthermore, the observed time series can exhibit other non-idealities, such as outliers, short length and distortion from the original wave form. Hence, the computational methods should preferably be robust against such anomalies in the data. RESULTS: We propose a general-purpose robust testing procedure for finding periodic sequences in multiple time series data. The proposed method is based on a robust spectral estimator which is incorporated into the hypothesis testing framework using a so-called g-statistic together with correction for multiple testing. This results in a robust testing procedure which is insensitive to heavy contamination of outliers, missing-values, short time series, nonlinear distortions, and is completely insensitive to any monotone nonlinear distortions. The performance of the methods is evaluated by performing extensive simulations. In addition, we compare the proposed method with another recent statistical signal detection estimator that uses Fisher's test, based on the Gaussian noise assumption. The results demonstrate that the proposed robust method provides remarkably better robustness properties. Moreover, the performance of the proposed method is preferable also in the standard Gaussian case. We validate the performance of the proposed method on real data on which the method performs very favorably. CONCLUSION: As the time series measured from biological systems are usually short and prone to contain different kinds of non-idealities, we are very optimistic about the multitude of possible applications for our proposed robust statistical periodicity detection method.

Algorithms↗

In silico microdissection of microarray data from heterogeneous cell populations.

BACKGROUND: Very few analytical approaches have been reported to resolve the variability in microarray measurements stemming from sample heterogeneity. For example, tissue samples used in cancer studies are usually contaminated with the surrounding or infiltrating cell types. This heterogeneity in the sample preparation hinders further statistical analysis, significantly so if different samples contain different proportions of these cell types. Thus, sample heterogeneity can result in the identification of differentially expressed genes that may be unrelated to the biological question being studied. Similarly, irrelevant gene combinations can be discovered in the case of gene expression based classification. RESULTS: We propose a computational framework for removing the effects of sample heterogeneity by "microdissecting" microarray data in silico. The computational method provides estimates of the expression values of the pure (non-heterogeneous) cell samples. The inversion of the sample heterogeneity can be facilitated by providing accurate estimates of the mixing percentages of different cell types in each measurement. For those cases where no such information is available, we develop an optimization-based method for joint estimation of the mixing percentages and the expression values of the pure cell samples. We also consider the problem of selecting the correct number of cell types. CONCLUSION: The efficiency of the proposed methods is illustrated by applying them to a carefully controlled cDNA microarray data obtained from heterogeneous samples. The results demonstrate that the methods are capable of reconstructing both the sample and cell type specific expression values from heterogeneous mixtures and that the mixing percentages of different cell types can also be estimated. Furthermore, a general purpose model selection method can be used to select the correct number of cell types.

Algorithms↗

Stability of functions in Boolean models of gene regulatory networks.

Boolean networks are used to model large nonlinear systems such as gene regulatory networks. We will present results that can be used to understand how the choice of functions affects the network dynamics. The so called bias-map and its fixed points depict much of the function's dynamical role in the network. We define the concept of stabilizing functions and show that many Post and canalizing functions are also stabilizing functions. Boolean networks constructed using the same type of stabilizing functions are always stable regardless of the average in-degree of network functions. We derive the number of all stabilizing functions and find it to be much larger than the number of Post and canalizing functions. We also discuss the implementation of functions and apply the presented results to biological data that give an approximation of the distribution of regulatory functions in eucaryotic cells. We find that the obtained theoretical results on the number of active genes are biologically plausible. Finally, based on the presented results, we discuss why canalizing and Post regulatory functions seem to be common in cells.

Adaptation, Physiological↗

Robust quantification of in vitro angiogenesis through image analysis.

An automated image analysis method for quantification of in vitro angiogenesis is presented. The method is designed for in vitro angiogenesis assays that are based on co-culturing endothelial cells with fibroblasts. Such assays are used in many current studies in which anti-angiogenic agents for the treatment of cancer are being sought. This search requires accurate quantification of the stimulatory and inhibitory effects of the different agents. The quantification method gives lengths and sizes of the tubule complexes as well as the numbers of junctions in each of them. The method is tested with a set of test images obtained with a commercially available in vitro angiogenesis assay. The results correctly indicate the inhibitory effect of suramin and the stimulatory effect of vascular endothelial growth factor. Moreover, the image analysis method is shown to be robust against variations in illumination. We have implemented a software package that utilizes the methods. The software as well as a set of test images are available at http://www.cs.tut.fi/sgn/csb/angioquant/.

Algorithms↗

Software for quantification of labeled bacteria from digital microscope images by automated image analysis.

Automated image analysis software, CellC, was developed and validated for quantification of bacterial cells from digital microscope images. CellC enables automated enumeration of bacterial cells, comparison of total count and specific count images [e.g., 4',6-diamino-2-phenylindole (DAPI) and fluorescence in situ hybridization (FISH) images], and provides quantitative estimates of cell morphology. The software includes an intuitive graphical user interface that enables easy usage as well as sequential analysis of multiple images without user intervention. Validation of enumeration reveals correlation to be better than 0.98 when total bacterial counts by CellC are compared with manual enumeration, with all validated image types. The software is freely available and modifiable: the executable files and MATLAB source codes can be obtained at www. cs. tut.fi/sgn/csb/cellc.

Bioreactors↗

Simulation tools for biochemical networks: evaluation of performance and usability.

MOTIVATION: Simulation of dynamic biochemical systems is receiving considerable attention due to increasing availability of experimental data of complex cellular functions. Numerous simulation tools have been developed for numerical simulation of the behavior of a system described in mathematical form. However, there exist only a few evaluation studies of these tools. Knowledge of the properties and capabilities of the simulation tools would help bioscientists in building models based on experimental data. RESULTS: We examine selected simulation tools that are intended for the simulation of biochemical systems. We choose four of them for more detailed study and perform time series simulations using a specific pathway describing the concentration of the active form of protein kinase C. We conclude that the simulation results are convergent between the chosen simulation tools. However, the tools differ in their usability, support for data transfer to other programs and support for automatic parameter estimation. From the experimentalists' point of view, all these are properties that need to be emphasized in the future.

Animals↗

Distinguishing key biological pathways between primary breast cancers and their lymph node metastases by gene function-based clustering analysis.

In order to identify key biological pathways that can distinguish between primary breast cancers and their lymph node metastases, we employed gene expression profiling together with gene function-based clustering analysis. We first acquired gene expression profiles of 9 matched primary tumors and the corresponding metastases that contained at least 75% of tumor cells. Then, we applied a clustering algorithm to the preprocessed data. In order to focus on the most informative genes, we ranked all the genes individually based on their abilities to separate the primary breast tumor and metastases samples. Further, we separated these genes into six functional groups according to the Stanford SOURCE database: 'cell cycle,' 'apoptosis,' 'metabolism,' 'cell adhesion and migration,' 'signal transduction,' and 'transcriptional factor and DNA binding molecules.' Unsupervised clustering analysis using all of the 2,303 genes on the microarrays was not able to separate the primary and metastases samples. Clustering analysis using the most informative genes revealed that primary tumors were more tightly clustered, whereas the metastases samples were relatively heterogeneous. The clustering analysis with the genes belonging to different functional groups showed that different functional gene sets varied in their abilities to separate primary tumors and their metastases. Marked separations were found with genes involved in metabolism, signal transduction, cell cycle, and transcriptional factor and DNA binding molecules. In contrast, apoptosis and cell adhesion and migration genes did not provide a clear separation of the two groups of samples. These results suggest that metastatic cells have different metabolism and signal transduction activities, regulated by transcriptional events, from the primary tumor cells. The results also suggest that the altered cell adhesion and migration potentials that are required for tumors to metastasize already exist in the primary tumors as a whole.

Biomarkers, Tumor↗

A novel strategy for microarray quality control using Bayesian networks.

MOTIVATION: High-throughput microarray technologies enable measurements of the expression levels of thousands of genes in parallel. However, microarray printing, hybridization and washing may create substantial variability in the quality of the data. As erroneous measurements may have a drastic impact on the results by disturbing the normalization schemes and by introducing expression patterns that lead to incorrect conclusions, it is crucial to discard low quality observations in the early phases of a microarray experiment. A typical microarray experiment consists of tens of thousands of spots on a microarray, making manual extraction of poor quality spots impossible. Thus, there is a need for a reliable and general microarray spot quality control strategy. RESULTS: We suggest a novel strategy for spot quality control by using Bayesian networks, which contain many appealing properties in the spot quality control context. We illustrate how a non-linear least squares based Gaussian fitting procedure can be used in order to extract features for a spot on a microarray. The features we used in this study are: spot intensity, size of the spot, roundness of the spot, alignment error, background intensity, background noise, and bleeding. We conclude that Bayesian networks are a reliable and useful model for microarray spot quality assessment. SUPPLEMENTARY INFORMATION: http://sigwww.cs.tut.fi/TICSP/SpotQuality/.

Algorithms↗

CGH-Plotter: MATLAB toolbox for CGH-data analysis.

CGH-Plotter is a MATLAB toolbox with a graphical user interface for the analysis of comparative genomic hybridization (CGH) microarray data. CGH-Plotter provides a tool for rapid visualization of CGH-data according to the locations of the genes along the genome. In addition, the CGH-Plotter identifies regions of amplifications and deletions, using k-means clustering and dynamic programming. The application offers a convenient way to analyze CGH-data and can also be applied for the analysis of cDNA microarray expression data. CGH-Plotter toolbox is platform independent and requires MATLAB 6.1 or higher to operate.

Cluster Analysis↗