PubMed Health⌕ Search

Biomedical subjects

Jianbo Gao

Publications and source records attributed to Jianbo Gao.

4 recordsLinked to original sources

Assessment of long-range correlation in time series: how to avoid pitfalls.

Due to the ubiquity of time series with long-range correlation in many areas of science and engineering, analysis and modeling of such data is an important problem. While the field seems to be mature, three major issues have not been satisfactorily resolved. (i) Many methods have been proposed to assess long-range correlation in time series. Under what circumstances do they yield consistent results? (ii) The mathematical theory of long-range correlation concerns the behavior of the correlation of the time series for very large times. A measured time series is finite, however. How can we relate the fractal scaling break at a specific time scale to important parameters of the data? (iii) An important technique in assessing long-range correlation in a time series is to construct a random walk process from the data, under the assumption that the data are like a stationary noise process. Due to the difficulty in determining whether a time series is stationary or not, however, one cannot be 100% sure whether the data should be treated as a noise or a random walk process. Is there any penalty if the data are interpreted as a noise process while in fact they are a random walk process, and vice versa? In this paper, we seek to gain important insights into these issues by examining three model systems, the autoregressive process of order 1, on-off intermittency, and Lévy motions, and considering an important engineering problem, target detection within sea-clutter radar returns. We also provide a few rules of thumb to safeguard against misinterpretations of long-range correlation in a time series, and discuss relevance of this study to pattern recognition.

Journal Article↗

Analysis of biomedical signals by the lempel-Ziv complexity: the effect of finite data size.

The Lempel-Ziv (LZ) complexity and its variants are popular metrics for characterizing biological signals. Proper interpretation of such analyses, however, has not been thoroughly addressed. In this letter, we study the the effect of finite data size. We derive analytic expressions for the LZ complexity for regular and random sequences, and employ them to develop a normalization scheme. To gain further understanding, we compare the LZ complexity with the correlation entropy from chaos theory in the context of epileptic seizure detection from EEG data, and discuss advantages of the normalized LZ complexity over the correlation entropy.

Algorithms↗

Reliability of the 0-1 test for chaos.

In time series analysis, it has been considered of key importance to determine whether a complex time series measured from the system is regular, deterministically chaotic, or random. Recently, Gottwald and Melbourne have proposed an interesting test for chaos in deterministic systems. Their analyses suggest that the test may be universally applicable to any deterministic dynamical system. In order to fruitfully apply their test to complex experimental data, it is important to understand the mechanism for the test to work, and how it behaves when it is employed to analyze various types of data, including those not from clean deterministic systems. We find that the essence of their test can be described as to first constructing a random walklike process from the data, then examining how the variance of the random walk scales with time. By applying the test to three sets of data, corresponding to (i) 1/falpha noise with long-range correlations, (ii) edge of chaos, and (iii) weak chaos, we show that the test mis-classifies (i) both deterministic and weakly stochastic edge of chaos and weak chaos as regular motions, and (ii) strongly stochastic edge of chaos and weak chaos, as well as 1/falpha noise as deterministic chaos. Our results suggest that, while the test may be effective to discriminate regular motion from fully developed deterministic chaos, it is not useful for exploratory purposes, especially for the analysis of experimental data with little a priori knowledge. A few speculative comments on the future of multiscale nonlinear time series analysis are made.

Journal Article↗

Protein coding sequence identification by simultaneously characterizing the periodic and random features of DNA sequences.

Most codon indices used today are based on highly biased nonrandom usage of codons in coding regions. The background of a coding or noncoding DNA sequence, however, is fairly random, and can be characterized as a random fractal. When a gene-finding algorithm incorporates multiple sources of information about coding regions, it becomes more successful. It is thus highly desirable to develop new and efficient codon indices by simultaneously characterizing the fractal and periodic features of a DNA sequence. In this paper, we describe a novel way of achieving this goal. The efficiency of the new codon index is evaluated by studying all of the 16 yeast chromosomes. In particular, we show that the method automatically and correctly identifies which of the three reading frames is the one that contains a gene.

Journal Article↗