PubMed Health⌕ Search

Biomedical subjects

Erkki Oja

Publications and source records attributed to Erkki Oja.

10 recordsLinked to original sources

Exploratory analysis of climate data using source separation methods.

We present an example of exploratory data analysis of climate measurements using a recently developed denoising source separation (DSS) framework. We analyzed a combined dataset containing daily measurements of three variables: surface temperature, sea level pressure and precipitation around the globe, for a period of 56 years. Components exhibiting slow temporal behavior were extracted using DSS with linear denoising. The first component, most prominent in the interannual time scale, captured the well-known El Niño-Southern Oscillation (ENSO) phenomenon and the second component was close to the derivative of the first one. The slow components extracted in a wider frequency range were further rotated using a frequency-based separation criterion implemented by DSS with nonlinear denoising. The rotated sources give a meaningful representation of the slow climate variability as a combination of trends, interannual oscillations, the annual cycle and slowly changing seasonal variations. Again, components related to the ENSO phenomenon emerge very clearly among the found sources.

Atmospheric Pressure↗

The Evolving Tree--analysis and applications.

In this paper, we enhance and analyze the Evolving Tree (ETree) data analysis algorithm. The suggested improvements aim to make the system perform better while still maintaining the simple nature of the basic algorithm. We also examine the system's behavior with many different kinds of tests, measurements and visualizations. We compare the ETree's performance against classical data analysis methods and very similar modern systems. We find that the ETree is a suitable method for unsupervised analysis of huge data sets.

Algorithms↗

Efficient variant of algorithm FastICA for independent component analysis attaining the Cramér-Rao lower bound.

FastICA is one of the most popular algorithms for independent component analysis (ICA), demixing a set of statistically independent sources that have been mixed linearly. A key question is how accurate the method is for finite data samples. We propose an improved version of the FastICA algorithm which is asymptotically efficient, i.e., its accuracy given by the residual error variance attains the Cramér-Rao lower bound (CRB). The error is thus as small as possible. This result is rigorously proven under the assumption that the probability distribution of the independent signal components belongs to the class of generalized Gaussian (GG) distributions with parameter alpha, denoted GG(alpha) for alpha > 2. We name the algorithm efficient FastICA (EFICA). Computational complexity of a Matlab implementation of the algorithm is shown to be only slightly (about three times) higher than that of the standard symmetric FastICA. Simulations corroborate these claims and show superior performance of the algorithm compared with algorithm JADE of Cardoso and Souloumiac and nonparametric ICA of Boscolo et al. on separating sources with distribution GG (alpha) with arbitrary alpha, as well as on sources with bimodal distribution, and a good performance in separating linearly mixed speech signals.

Algorithms↗

The fastICA algorithm revisited: convergence analysis.

The fast independent component analysis (FastICA) algorithm is one of the most popular methods to solve problems in ICA and blind source separation. It has been shown experimentally that it outperforms most of the commonly used ICA algorithms in convergence speed. A rigorous local convergence analysis has been presented only for the so-called one-unit case, in which just one of the rows of the separating matrix is considered. However, in the FastICA algorithm, there is also an explicit normalization step, and it may be questioned whether the extra rotation caused by the normalization will affect the convergence speed. The purpose of this paper is to show that this is not the case and the good convergence properties of the one-unit case are also shared by the full algorithm with symmetrical normalization. A local convergence analysis is given for the general case, and the global behavior is illustrated numerically for two sources and two mixtures in several typical cases.

Algorithms↗

A "nonnegative PCA" algorithm for independent component analysis.

We consider the task of independent component analysis when the independent sources are known to be nonnegative and well-grounded, so that they have a nonzero probability density function (pdf) in the region of zero. We propose the use of a "nonnegative principal component analysis (nonnegative PCA)" algorithm, which is a special case of the nonlinear PCA algorithm, but with a rectification nonlinearity, and we conjecture that this algorithm will find such nonnegative well-grounded independent sources, under reasonable initial conditions. While the algorithm has proved difficult to analyze in the general case, we give some analytical results that are consistent with this conjecture and some numerical simulations that illustrate its operation.

Algorithms↗

Nonlinear dynamical factor analysis for state change detection.

Changes in a dynamical process are often detected by monitoring selected indicators directly obtained from the process observations, such as the mean values or variances. Standard change detection algorithms such as the Shewhart control charts or the cumulative sum (CUSUM) algorithm are often based on such first- and second-order statistics. Much better results can be obtained if the dynamical process is properly modeled, for example by a nonlinear state-space model, and then the accuracy of the model is monitored over time. The success of the latter approach depends largely on the quality of the model. In practical applications like industrial processes, the state variables, dynamics, and observation mapping are rarely known accurately. Learning from data must be used; however, methods for the simultaneous estimation of the state and the unknown nonlinear mappings are very limited. We use a novel method of learning a nonlinear state-space model, the nonlinear dynamical factor analysis (NDFA) algorithm. It takes a set of multivariate observations over time and fits blindly a generative dynamical latent variable model, resembling nonlinear independent component analysis. We compare the performance of the model in process change detection to various traditional methods. It is shown that NDFA outperforms the classical methods by a wide margin in a variety of cases where the underlying process dynamics changes.

Factor Analysis, Statistical↗

Blind separation of positive sources by globally convergent gradient search.

The instantaneous noise-free linear mixing model in independent component analysis is largely a solved problem under the usual assumption of independent nongaussian sources and full column rank mixing matrix. However, with some prior information on the sources, like positivity, new analysis and perhaps simplified solution methods may yet become possible. In this letter, we consider the task of independent component analysis when the independent sources are known to be nonnegative and well grounded, which means that they have a nonzero pdf in the region of zero. It can be shown that in this case, the solution method is basically very simple: an orthogonal rotation of the whitened observation vector into nonnegative outputs will give a positive permutation of the original sources. We propose a cost function whose minimum coincides with nonnegativity and derive the gradient algorithm under the whitening constraint, under which the separating matrix is orthogonal. We further prove that in the Stiefel manifold of orthogonal matrices, the cost function is a Lyapunov function for the matrix gradient flow, implying global convergence. Thus, this algorithm is guaranteed to find the nonnegative well-grounded independent sources. The analysis is complemented by a numerical simulation, which illustrates the algorithm.

Algorithms↗

Class distributions on SOM surfaces for feature extraction and object retrieval.

A Self-Organizing Map (SOM) is typically trained in unsupervised mode, using a large batch of training data. If the data contain semantically related object groupings or classes, subsets of vectors belonging to such user-defined classes can be mapped on the SOM by finding the best matching unit for each vector in the set. The distribution of the data vectors over the map forms a two-dimensional discrete probability density. Even from the same data, qualitatively different distributions can be obtained by using different feature extraction techniques. We used such feature distributions for comparing different classes and different feature representations of the data in the context of our content-based image retrieval system PicSOM. The information-theoretic measures of entropy and mutual information are suggested to evaluate the compactness of a distribution and the independence of two distributions. Also, the effect of low-pass filtering the SOM surfaces prior to the calculation of the entropy is studied.

Artificial Intelligence↗

Independent component analysis for artefact separation in astrophysical images.

In this paper, we demonstrate that independent component analysis, a novel signal processing technique, is a powerful method for separating artefacts from astrophysical image data. When studying far-out galaxies from a series of consequent telescope images, there are several sources for artefacts that influence all the images, such as camera noise, atmospheric fluctuations and disturbances, cosmic rays, and stars in our own galaxy. In the analysis of astrophysical image data it is very important to implement techniques which are able to detect them with great accuracy, to avoid the possible physical events from being eliminated from the data along with the artefacts. For this problem, the linear ICA model holds very accurately because such artefacts are all theoretically independent of each other and of the physical events. Using image data on the M31 Galaxy, it is shown that several artefacts can be detected and recognized based on their temporal pixel luminosity profiles and independent component images. The obtained separation is good and the method is very fast. It is also shown that ICA outperforms principal component analysis in this task. For these reasons, ICA might provide a very useful pre-processing technique for the large amounts of available telescope image data.

Artifacts↗