PubMed Health⌕ Search

Biomedical subjects

Yinhe Cao

Publications and source records attributed to Yinhe Cao.

6 recordsLinked to original sources

Assessment of long-range correlation in time series: how to avoid pitfalls.

Due to the ubiquity of time series with long-range correlation in many areas of science and engineering, analysis and modeling of such data is an important problem. While the field seems to be mature, three major issues have not been satisfactorily resolved. (i) Many methods have been proposed to assess long-range correlation in time series. Under what circumstances do they yield consistent results? (ii) The mathematical theory of long-range correlation concerns the behavior of the correlation of the time series for very large times. A measured time series is finite, however. How can we relate the fractal scaling break at a specific time scale to important parameters of the data? (iii) An important technique in assessing long-range correlation in a time series is to construct a random walk process from the data, under the assumption that the data are like a stationary noise process. Due to the difficulty in determining whether a time series is stationary or not, however, one cannot be 100% sure whether the data should be treated as a noise or a random walk process. Is there any penalty if the data are interpreted as a noise process while in fact they are a random walk process, and vice versa? In this paper, we seek to gain important insights into these issues by examining three model systems, the autoregressive process of order 1, on-off intermittency, and Lévy motions, and considering an important engineering problem, target detection within sea-clutter radar returns. We also provide a few rules of thumb to safeguard against misinterpretations of long-range correlation in a time series, and discuss relevance of this study to pattern recognition.

Journal Article↗

Reliability of the 0-1 test for chaos.

In time series analysis, it has been considered of key importance to determine whether a complex time series measured from the system is regular, deterministically chaotic, or random. Recently, Gottwald and Melbourne have proposed an interesting test for chaos in deterministic systems. Their analyses suggest that the test may be universally applicable to any deterministic dynamical system. In order to fruitfully apply their test to complex experimental data, it is important to understand the mechanism for the test to work, and how it behaves when it is employed to analyze various types of data, including those not from clean deterministic systems. We find that the essence of their test can be described as to first constructing a random walklike process from the data, then examining how the variance of the random walk scales with time. By applying the test to three sets of data, corresponding to (i) 1/falpha noise with long-range correlations, (ii) edge of chaos, and (iii) weak chaos, we show that the test mis-classifies (i) both deterministic and weakly stochastic edge of chaos and weak chaos as regular motions, and (ii) strongly stochastic edge of chaos and weak chaos, as well as 1/falpha noise as deterministic chaos. Our results suggest that, while the test may be effective to discriminate regular motion from fully developed deterministic chaos, it is not useful for exploratory purposes, especially for the analysis of experimental data with little a priori knowledge. A few speculative comments on the future of multiscale nonlinear time series analysis are made.

Journal Article↗

Protein coding sequence identification by simultaneously characterizing the periodic and random features of DNA sequences.

Most codon indices used today are based on highly biased nonrandom usage of codons in coding regions. The background of a coding or noncoding DNA sequence, however, is fairly random, and can be characterized as a random fractal. When a gene-finding algorithm incorporates multiple sources of information about coding regions, it becomes more successful. It is thus highly desirable to develop new and efficient codon indices by simultaneously characterizing the fractal and periodic features of a DNA sequence. In this paper, we describe a novel way of achieving this goal. The efficiency of the new codon index is evaluated by studying all of the 16 yeast chromosomes. In particular, we show that the method automatically and correctly identifies which of the three reading frames is the one that contains a gene.

Journal Article↗

Recurrence time statistics: versatile tools for genomic DNA sequence analysis.

With the completion of the human and a few model organisms' genomes, and with the genomes of many other organisms waiting to be sequenced, it has become increasingly important to develop faster computational tools which are capable of easily identifying the structures and extracting features from DNA sequences. One of the more important structures in a DNA sequence is repeat-related. Often they have to be masked before protein coding regions along a DNA sequence are to be identified or redundant expressed sequence tags (ESTs) are to be sequenced. Here we report a novel recurrence time-based method for sequence analysis. The method can conveniently study all kinds of periodicity and exhaustively find all repeat-related features from a genomic DNA sequence. An efficient codon index is also derived from the recurrence time statistics, which has the salient features of being largely species-independent and working well on very short sequences. Efficient codon indices are key elements of successful gene finding algorithms, and are particularly useful for determining whether a suspected EST belongs to a coding or non-coding region. We illustrate the power of the method by studying the genomes of E. coli, the yeast S. cervisivae, the nematode worm C. elegans, and the human, Homo sapiens. Our method requires approximately 6 . N byte memory and a computational time of N log N to extract all the repeat-related and periodic or quasi-periodic features from a sequence of length N without any prior knowledge on the consensus sequence of those features, hence enables us to carry out sequence analysis on the whole genomic scale by a PC.

Algorithms↗

Detecting dynamical changes in time series using the permutation entropy.

Timely detection of unusual and/or unexpected events in natural and man-made systems has deep scientific and practical relevance. We show that the recently proposed conceptually simple and easily calculated measure of permutation entropy can be effectively used to detect qualitative and quantitative dynamical changes. We illustrate our results on two model systems as well as on clinically characterized brain wave data from epileptic patients.

Journal Article↗

Recurrence time statistics: versatile tools for genomic DNA sequence analysis.

With the completion of the human and a few model organisms' genomes, and the genomes of many other organisms waiting to be sequenced, it has become increasingly important to develop faster computational tools which are capable of easily identifying the structures and extracting features from DNA sequences. One of the more important structures in a DNA sequence is repeat-related. Often they have to be masked before protein coding regions along a DNA sequence are to be identified or redundant expressed sequence tags (ESTs) are to be sequenced. Here we report a novel recurrence time based method for sequence analysis. The method can conveniently study all kinds of periodicity and exhaustively find all repeat-related features from a genomic DNA sequence. An efficient codon index is also derived from the recurrence time statistics, which has the salient features of being largely species-independent and working well on very short sequences. Efficient codon indices are key elements of successful gene finding algorithms, and are particularly useful for determining whether a suspected EST belongs to a coding or non-coding region. We illustrate the power of the method by studying the genomes of E. coli, the yeast S. cervisivae, the nematode worm C. elegans, and the human, Homo sapiens. Computationally, our method is very efficient. It allows us to carry out analysis of genomes on the whole genomic scale by a PC.

Algorithms↗