PubMed Health⌕ Search

Biomedical subjects

Zhen-Su She

Publications and source records attributed to Zhen-Su She.

8 recordsLinked to original sources

Tree and rate estimation by local evaluation of heterochronous nucleotide data.

MOTIVATION: Heterochronous gene sequence data is important for characterizing the evolutionary processes of fast-evolving organisms such as RNA viruses. A limited set of algorithms exists for estimating the rate of nucleotide substitution and inferring phylogenetic trees from such data. The authors here present a new method, Tree and Rate Estimation by Local Evaluation (TREBLE) that robustly calculates the rate of nucleotide substitution and phylogeny with several orders of magnitude improvement in computational time. METHODS: For the basis of its rate estimation TREBLE novelly utilizes a geometric interpretation of the molecular clock assumption to deduce a local estimate of the rate of nucleotide substitution for triplets of dated sequences. Averaging the triplet estimates via a variance weighting yields a global estimate of the rate. From this value, an iterative refinement procedure relying on statistical properties of the triplets then generates a final estimate of the global rate of nucleotide substitution. The estimated global rate is then utilized to find the tree from the pairwise distance matrix via an UPGMA-like algorithm. RESULTS: Simulation studies show that TREBLE estimates the rate of nucleotide substitution with point estimates comparable with the best of available methods. Confidence intervals are comparable with that of BEAST. TREBLE's phylogenetic reconstruction is significantly improved over the other distance matrix method but not as accurate as the Bayesian algorithm. Compared with three other algorithms, TREBLE reduces computational time by a minimum factor of 3000. Relative to the algorithm with the most accurate estimates for the rate of nucleotide substitution (i.e. BEAST), TREBLE is over 10,000 times more computationally efficient. AVAILABILITY: jdobrien.bol.ucla.edu/TREBLE.html

Chromosome Mapping↗

Hierarchical structure analysis describing abnormal base composition of genomes.

Abnormal base compositional patterns of genomic DNA sequences are studied in the framework of a hierarchical structure (HS) model originally proposed for the study of fully developed turbulence [She and Lévêque, Phys. Rev. Lett. 72, 336 (1994)]. The HS similarity law is verified over scales between 10(3)bp and 10(5)bp, and the HS parameter beta is proposed to describe the degree of heterogeneity in the base composition patterns. More than one hundred bacteria, archaea, virus, yeast, and human genome sequences have been analyzed and the results show that the HS analysis efficiently captures abnormal base composition patterns, and the parameter beta is a characteristic measure of the genome. Detailed examination of the values of beta reveals an intriguing link to the evolutionary events of genetic material transfer. Finally, a sequence complexity (S) measure is proposed to characterize gradual increase of organizational complexity of the genome during the evolution. The present study raises several interesting issues in the evolutionary history of genomes.

Animals↗

Hierarchical structure description of spatiotemporal chaos.

We develop a hierarchical structure (HS) analysis for quantitative description of statistical states of spatially extended systems. Examples discussed here include an experimental reaction-diffusion system with Belousov-Zhabotinsky kinetics, the two-dimensional complex Ginzburg-Landau equation, and the modified FitzHugh-Nagumon equation, which all show complex dynamics of spirals and defects. We demonstrate that the spatial-temporal fluctuation fields in the above-mentioned systems all display the HS similarity property originally proposed for the study of fully developed turbulence [Phys. Rev. Lett. 72, 336 (1994)]]. The derived values of a HS parameter beta from experimental and numerical data in various physical regimes exhibit consistent trends and characterize the degree of turbulence in the systems near the transition, and the degree of heterogeneity of multiple disorders far from the transition. It is suggested that the HS analysis offers a useful quantitative description for the complex dynamics of two-dimensional spatiotemporal patterns.

Journal Article↗

Scaling and hierarchical structures in DNA sequences.

A method of analyzing DNA correlation structure is introduced. Density fluctuations of nucleotides are shown to display an extended self-similarity scaling when the scale varies between 100 and 8000 base pairs. The scaling is accurately described by a hierarchical structure model of She and Leveque [Phys. Rev. Lett. 72, 336 (1994)]]. The derived model parameter beta is able to quantify moderately large-scale correlations which exist in a true DNA sequence but are absent in its randomly shuffled sequence and in a simulated model sequence by an evolution model of Hsieh et al. [Phys. Rev. Lett. 90, (2003)]]. Finally, it is shown that beta varies with the evolution category and measures the organizational complexity of the genome.

Animals↗

Accuracy improvement for identifying translation initiation sites in microbial genomes.

MOTIVATION: At present the computational gene identification methods in microbial genomes have a high prediction accuracy of verified translation termination site (3' end), but a much lower accuracy of the translation initiation site (TIS, 5' end). The latter is important to the analysis and the understanding of the putative protein of a gene and the regulatory machinery of the translation. Improving the accuracy of prediction of TIS is one of the remaining open problems. RESULTS: In this paper, we develop a four-component statistical model to describe the TIS of prokaryotic genes. The model incorporates several features with biological meanings, including the correlation between translation termination site and TIS of genes, the sequence content around the start codon; the sequence content of the consensus signal related to ribosomal binding sites (RBSs), and the correlation between TIS and the upstream consensus signal. An entirely non-supervised training system is constructed, which takes as input a set of annotated coding open reading frames (ORFs) by any gene finder, and gives as output a set of organism-specific parameters (without any prior knowledge or empirical constants and formulas). The novel algorithm is tested on a set of reliable datasets of genes from Escherichia coli and Bacillus subtillis. MED-Start may correctly predict 95.4% of the start sites of 195 experimentally confirmed E.coli genes, 96.6% of 58 reliable B.subtillis genes. Moreover, the test results indicate that the algorithm gives higher accuracy for more reliable datasets, and is robust to the variation of gene length. MED-Start may be used as a postprocessor for a gene finder. After processing by our program, the improvement of gene start prediction of gene finder system is remarkable, e.g. the accuracy of TIS predicted by MED 1.0 increases from 61.7 to 91.5% for 854 E.coli verified genes, while that by GLIMMER 2.02 increases from 63.2 to 92.0% for the same dataset. These results show that our algorithm is one of the most accurate methods to identify TIS of prokaryotic genomes. AVAILABILITY: The program MED-Start can be accessed through the website of CTB at Peking University: http://ctb.pku.edu.cn/main/SheGroup/MED_Start.htm.

Algorithms↗

Multivariate entropy distance method for prokaryotic gene identification.

A new simple method is found for efficient and accurate identification of coding sequences in prokaryotic genome. The method employs a Shannon description of artificial language for DNA sequences. It consists in translating a DNA sequence into a pseudo-amino acid sequence with 20 fundamental words according to the universal genetic code. With an entropy-density profile (EDP), the method maps a sequence of finite length to a vector and then analyzes its position in the 20-dimensional phase space depending on its nature. It is found that the ratio of the relative distance to an averaged coding and non-coding EDP over a small number (up to one) of open reading frames (ORFs) can serve as a good coding potential. An iterative algorithm is designed for finding a set of "root" sequences using this coding potential. A multivariate entropy distance (MED) algorithm is then proposed for the identification of prokaryotic genes; it has a feature to combine the use of a coding potential and an EDP-based sequence similarity analysis. The current version of MED is unsupervised, parameter-free and simple to implement. It is demonstrated to be able to detect 95-99% genes with 10-30% of additional genes when tested against the RefSeq database of NCBI and to detect 97.5-99.8% of confirmed genes with known functions. It is also shown to be able to find a set of (functionally known) genes that are missed by other well-known gene finding algorithms. All measurements show that the MED algorithm reaches a similar performance level as the algorithms like GeneMark and Glimmer for prokaryotic gene prediction.

Algorithms↗

Extended self-similarity and hierarchical structure in turbulence.

We show that a generalization of the She-Leveque hierarchical structure [Z.S. She and E. Leveque, Phys. Rev. Lett. 72, 336 (1994)] together with a constant maximum magnitude of the velocity difference give rise to the extended self-similarity (ESS) [R. Benzi et al., Phy. Rev. E 48, R29 (1993)]. Our analysis thus suggests that the ESS measured in turbulent flows is an indication of the most intense structures being shocklike. Analyses of velocity measurements in a turbulent pipe flow support our conjecture.

Journal Article↗

Anomalous self-similarity in a turbulent rapidly rotating fluid.

Our velocity measurements on quasi-two-dimensional turbulent flow in a rapidly rotating annulus yield self-similar (scale-independent) probability distribution functions for longitudinal velocity differences, deltav(l) = v(x+l)-v(x). These distribution functions are strongly non-Gaussian, suggesting that the coherent vortices play a significant role. The structure functions <[deltav(l)](p)> approximately l(zeta)p exhibit anomalous scaling: zeta(p) = p / 2 rather than the expected zeta(p) = p / 3. Correspondingly, the energy spectrum is described by E(k) approximately k(-2) rather than the expected E(k) approximately k(-5/3).

Journal Article↗