PubMed Health⌕ Search

Biomedical subjects

Wei-Mou Zheng

Publications and source records attributed to Wei-Mou Zheng.

3 recordsLinked to original sources

In-phase implies large likelihood for independent codon model: distinguishing coding from non-coding sequences.

It is proven that under the independent codon model, the likelihood of a DNA coding sequence read according to the correct frame is asymptotically larger than that read with an incorrect frame. Based on this proposition, a single set of probabilities of the codon usage is enough for discriminating the six frames of coding sequences under the independent codon model. The direct coding sequence of Escherichia coli genome is taken as an example to examine the codon independency by using the mutual information and chi2 analysis. The contrast between the coding frame and the two offset frames is evident. A self-learning approach for generating training set is proposed to estimate probability parameters.

Codon↗

Distances and classification of amino acids for different protein secondary structures.

Window profiles of amino acids in protein sequences are used to describe the amino acid environment. The relative entropy or Kullback-Leibler distance derived from these profiles is used as a measure of dissimilarity for comparison of amino acids and secondary structure conformations. Distance matrices of amino acid pairs at different conformations are obtained, which display a non-negligible dependence of amino acid similarity on conformations. Based on the conformation specific distances, a clustering analysis for amino acids is conducted.

Algorithms↗

Simplified amino acid alphabets based on deviation of conditional probability from random background.

The primitive data for deducing the Miyazawa-Jernigan contact energy or blocks substitution matrix (BLOSUM) consists of pair frequency counts. Each amino acid corresponds to a conditional probability distribution. Based on the deviation of such a conditional probability from random background, a scheme for the reduction of the amino acid alphabet is proposed. It is observed that an evident discrepancy exists between the reduced alphabets obtained from the raw data of the Miyazawa-Jernigan's and BLOSUM's residue pair counts. Taking a homologous sequence database SCOP40 as a test set, we detect homology with the obtained coarse-grained substitution matrices. It is verified that the reduced alphabets obtained well preserve information contained in the original 20-letter alphabet.

Algorithms↗