PubMed HealthSearch

Biomedical subjects

C T Zhang

Publications and source records attributed to C T Zhang.

15 recordsLinked to original sources

A weighting method for predicting protein structural class from amino acid composition.

A protein is generally classified into one of the following four structural classes: all alpha, all beta, alpha+beta and alpha/beta. In this paper, based on the weighting to the 20 constituent amino acids, a new method is proposed for predicting the structural class of a protein according to its amino acid composition. The 20 weighting parameters, which reflect the different properties of the 20 constituent amino acids, have been obtained from a training set of proteins through the linear-programming approach. The rate of correct prediction for a training set of proteins by means of the new method was 100%, whereas the highest rate of previous methods was 82.8%. Furthermore, the results showed that the more numerous training proteins, the more effective the new method.

Amino Acids

A correlation-coefficient method to predicting protein-structural classes from amino acid compositions.

A protein is usually classified into one of the following four structural classes: all alpha, all beta, (alpha + beta) and alpha/beta. In this paper, based on the maximum correlation-coefficient principle, a new formulation is proposed for predicting the structural class of a protein according to its amino acid composition. Calculations have been made for a development set of proteins from which the amino acid compositions for the standard structural classes were derived, and an independent set of proteins which are outside the development set. The former can test the self consistency of a method and the latter can test its extrapolating effectiveness. In both cases, the results showed that the new method gave a considerably higher rate of correct prediction than any of the previous methods, implying that a significant improvement has been achieved by implementing the maximum-correlation-coefficient principle in the new method.

Algorithms

An optimization approach to predicting protein structural class from amino acid composition.

Proteins are generally classified into four structural classes: all-alpha proteins, all-beta proteins, alpha + beta proteins, and alpha/beta proteins. In this article, a protein is expressed as a vector of 20-dimensional space, in which its 20 components are defined by the composition of its 20 amino acids. Based on this, a new method, the so-called maximum component coefficient method, is proposed for predicting the structural class of a protein according to its amino acid composition. In comparison with the existing methods, the new method yields a higher general accuracy of prediction. Especially for the all-alpha proteins, the rate of correct prediction obtained by the new method is much higher than that by any of the existing methods. For instance, for the 19 all-alpha proteins investigated previously by P.Y. Chou, the rate of correct prediction by means of his method was 84.2%, but the correct rate when predicted with the new method would be 100%! Furthermore, the new method is characterized by an explicable physical picture. This is reflected by the process in which the vector representing a protein to be predicted is decomposed into four component vectors, each of which corresponds to one of the norms of the four protein structural classes.

Amino Acids

Monte Carlo simulation studies on the prediction of protein folding types from amino acid composition.

In the methodology development for statistical prediction of protein structures, the founders of different methods usually selected different sets of proteins to test their predicted results. Therefore, it is hard to make a fair comparison according to the results they reported. Even if the predictions by different methods are performed for the same set of proteins, there is still such a problem: a method better that the other for one set of proteins would not necessarily remain so when applied to another set of proteins. To tackle this problem, a Monte Carlo simulation method is proposed to establish an objective criterion to measure the accuracy of prediction for the protein folding type. Such an objective accuracy is actually corresponding to the asymptotical limit genereated during the Monte Carlo simulation process. Based on that, it has been found that the average objective accuracy for predicting the all-alpha, all-beta, alpha + beta, and alpha/beta proteins by the least Euclid's distance method (Nakashima, H., K. Nishikawa, and T. Ooi. 1986. J. Biochem. 99:152-162) is 73.0% and that by the least Minkowski's distance method (Chou, P.Y. 1989. Prediction in Protein Structure and the Principles of Protein Conformation. Plenum Press. New York. 549-586) is 70.9%, indicating that the former is better than the latter. However, according to the original reports, the latter claimed a rate of correct prediction with 79.7% but the former with only 70.2%, leading to a completely opposite conclusion. This indicates the necessity of establishing an objective criterion, and a comparison is meaningful only when it is based on the objective criterion. The simulation method and the idea developed here also can be applied to examine any other statistical prediction methods.

Amino Acids

Diagrammatization of codon usage in 339 human immunodeficiency virus proteins and its biological implication.

The occurrence frequencies of bases A (adenine), C (cytosine, G (guanine), and T (thymine) occurring in the 1st, 2nd, and 3rd codon positions in the codon usage table of viral genes for the 339 human immunodeficiency virus (HIV) proteins compiled recently have been calculated and diagrammatized. For comparison, the corresponding diagrammatic representations for the 2681 human proteins from the codon usage table for primate genes are also presented. The analyzed results based on these characteristic diagrams indicate that considerably similar features have been found between HIV and human proteins for the 1st and 2nd codon positions; i.e., they are all occupied predominantly by purine, especially base A. However, a significant difference in the 3rd codon position between HIV and human proteins has been observed; i.e., human proteins are of high C + G content and low A + G content in the 3rd codon position, whereas the case is just the opposite for HIV proteins. The biological implication of such a duality on the codon bias of HIV against human proteins is discussed. It is suggested that the 1st and 2nd codon positions can be termed as the structure-determining position, and the 3rd codon position termed as the species-determining position. The diagrammatic representation and analysis method described here possess a great potential for the study of molecular evolution from the viewpoint of the genetic code for which data have been accumulated rapidly and will continue to grow at a much faster pace.

Base Composition

Analysis of distribution of bases in the coding sequences by a diagrammatic technique.

The frequencies of occurrence of four bases in the first, second and third codon positions and in the total coding sequences have been calculated by the codon usage table published in 1990 by Ikemura et al. The distribution of frequencies are further analysed in detail by a graphic technique presented recently by us. Formulas expressing the frequencies of four bases in the first and second codon positions in terms of frequencies of amino acids have been given. It is shown by the graphic analysis that for 90 species, in the first codon position the purine bases are dominant and in most cases G is the most dominant base. In the second codon position A is the most dominant base, while G is the least dominant base. In the third codon position the G + C content varies from 0.1 to 0.9, keeping the A + C content equal to 1/2 and G content equal to that of C, approximately. If the frequencies for bases A, C, G and U in the total coding sequences are denoted by a, c, g and u, respectively, it is found that the unequal formula: a2 + c2 + g2 + u2 less than 1/3, is valid for each of the 90 species including the human and E.coli etc.

Base Sequence

The study of stacking energy for natural DNA sequences.

The stacking energies between bases of DNA for both A and B conformation have been calculated by Aida & Nagata (1986, Int. J. Quantum Chem. 29, 1253-1261). For naturally occurring DNA sequences, by assuming that the average stacking energy of A conformation is completely identical with that of B, the lowest average stacking energy for double strand DNA has been calculated and found to be equal to -7.38 kcal M-1. This conclusion has been confirmed by the data of stacking energy of 112 promoter sequences of Escherichia coli and some other natural DNA sequences. Our result shows that the natural DNA sequences are bi-stable. One stable-state is the A conformation, the other is B. It is shown that the result is only applicable to the transcription processes that regulate the gene expression. Through the interaction with RNA polymerase in the processes of transcription, the conformation of DNA might undergo a transition of B----A----B.

Animals

Diagrammatic representation of the distribution of DNA bases and its applications.

The frequencies of occurrence of the bases A, C, G and T (or U) in a DNA or mRNA sequence are denoted by a, c, g, t (or u), respectively. Since a + c + g + t = 1, the four real numbers are mapped into a point within a regular tetrahedron. The mapping points are then projected to some planes for further study. We have given several examples of application of this technique. As a by-product of an application we have found an empirical formula on the limit distribution of DNA bases for different kind of organisms, i.e. 1/4 less than or equal to a2 + c2 + g2 + t2 less than 1/3.

Adenine

Analysis of patterns of twist angles in DNA double helix.

We have analysed theoretically the patterns of twist angles of B-DNA by the Tung-Harvey model mainly. It is shown that for a sequence of twist angles a smaller twist angle tends to follow a larger one and vice versa. Therefore the sequence of twist angles always takes a gentle zig-zag form. For simplicity we convert the sequence of twist angles to a symbolic sequence consisting of L and S, where L or S represents a large or a small angle, respectively. The -10 and -35 regions of 112 well-defined promoters for E. coli RNA polymerase, which were compiled by Hawley and McClure, have been analysed in terms of LS sequences in detail. The results shows that the number of LS sequences for promoters is considerably limited and the promoter mutations do not change the patterns of LS sequences in most cases. Several new ideas, which are believed to be useful in the further study, have been presented.

Base Sequence

The study of DNA sequences by their sequences of twist angles.

To study the properties of DNA sequences we have transformed the sequences of bases into the sequences of twist angles along the chain of DNA double helix by using the Dickerson sum function. The Fourier transform and the auto-correlation function of the twist angles sequences have been used to study the periodicity and randomness of the original DNA sequences. Basing on the correlation coefficient, a "distance" between two DNA fragments has been defined and used to compare some realistic DNA sequences. It is hoped that the techniques developed here could be used to analyze more realistic DNA sequences.

Base Sequence

Upper limit for the variances of some helical parameters in DNA double helix.

Assuming that the realistic DNA chains are random sequences of purines and pyrimidines and, by using the Dickerson sum functions, it is shown that there exists an upper limit for the variance of helical twist angles, the base-plane roll angles and the other helical parameters respectively in the DNA sequences. The estimates of the variances of all the above helical parameters for the DNA sequences in the Los Alamos data base have been performed and found to be in good agreement with the theoretical results obtained in this paper.

Animals

Analysis of sequences of twist angles in DNA double helix.

By assuming that the realistic DNA chains are random sequence of bases and using the Tung-Harvey formula for the prediction of twist angles, it is shown that the mean value of the sequence of twist angles is almost sequence-independent. In general the variance for the A, T-rich sequence is larger than that of G, C-rich sequence. There exists an upper bound for the variance of all possible sequences, i.e., the variance is not greater than 27 deg2. It is pointed out that the large conformational deviation from ideal DNA is an important factor for the recognition of DNA with protein/enzyme.

Base Sequence

Thalassemia in the Silk Road region of China.

This paper summarizes data obtained during a screening program involving 11,563 persons from the Silk Road region in China. The mean incidence of thalassemia is 1.62% with an increase from east to west. The incidence in the Hui population (3.01%) is higher than in Kazaks (2.92%), and in Uygur (2.22%). The Han population also has a higher incidence (0.98%) than seen for other regions in Northern China. The thalassemias observed are classified into seven groups; beta-thalassemia accounts for 89.48% of the total, and alpha-thalassemia for 10.52%.

China