PubMed Health⌕ Search

Biomedical subjects

D F Feng

Publications and source records attributed to D F Feng.

21 records · Page 2Linked to original sources

Aligning amino acid sequences: comparison of commonly used methods.

We examined two extensive families of protein sequences using four different alignment schemes that employ various degrees of "weighting" in order to determine which approach is most sensitive in establishing relationships. All alignments used a similarity approach based on a general algorithm devised by Needleman and Wunsch. The approaches included a simple program, UM (unitary matrix), whereby only identities are scored; a scheme in which the genetic code is used as a basis for weighting (GC); another that employs a matrix based on structural similarity of amino acids taken together with the genetic basis of mutation (SG); and a fourth that uses the empirical log-odds matrix (LOM) developed by Dayhoff on the basis of observed amino acid replacements. The two sequence families examined were (a) nine different globins and (b) nine different tyrosine kinase-like proteins. It was assumed a priori that all members of a family share common ancestry. In cases where two sequences were more than 30% identical, alignments by all four methods were almost always the same. In cases where the percentage identity was less than 20%, however, there were often significant differences in the alignments. On the average, the Dayhoff LOM approach was the most effective in verifying distant relationships, as judged by an empirical "jumbling test." This was not universally the case, however, and in some instances the simple UM was actually as good or better. Trees constructed on the basis of the various alignments differed with regard to their limb lengths, but had essentially the same branching orders. We suggest some reasons for the different effectivenesses of the four approaches in the two different sequence settings, and offer some rules of thumb for assessing the significance of sequence relationships.

Amino Acid Sequence↗

Computer-based characterization of epidermal growth factor precursor.

The cDNA sequence of the precursor of mouse epidermal growth factor (EGFP) has recently been reported by two groups, both of whom noted the presence of repeated similar segments, each about 40 residues long. One of these repeat units overlaps with the sequence of epidermal growth factor itself. The sequence of epidermal growth factor has been reported to be similar to that of pancreatic secretory trypsin inhibitor (PSTI) and a somewhat better match has been found with part of the sequence of bovine factor X, one of the blood coagulating factors. We report here that there is an even stronger similarity between the sequences of some of the repeat units of epidermal growth factor precursor and certain segments in factor X. This sequence similarity is also apparent in comparisons with other blood coagulation factors. On the basis of these sequence comparisons we suggest a scheme for the evolution of the epidermal growth factor precursor. We have also identified certain structural features in the precursor sequence that bear on the way in which the mature epidermal growth factor is generated.

Amino Acid Sequence↗