PubMed Health⌕ Search

Biomedical subjects

Yi-Kuo Yu

Publications and source records attributed to Yi-Kuo Yu.

6 recordsLinked to original sources

The construction of amino acid substitution matrices for the comparison of proteins with non-standard compositions.

MOTIVATION: Amino acid substitution matrices play a central role in protein alignment methods. Standard log-odds matrices, such as those of the PAM and BLOSUM series, are constructed from large sets of protein alignments having implicit background amino acid frequencies. However, these matrices frequently are used to compare proteins with markedly different amino acid compositions, such as transmembrane proteins or proteins from organisms with strongly biased nucleotide compositions. It has been argued elsewhere that standard matrices are not ideal for such comparisons and, furthermore, a rationale has been presented for transforming a standard matrix for use in a non-standard compositional context. RESULTS: This paper presents the mathematical details underlying the compositional adjustment of amino acid or DNA substitution matrices.

Algorithms↗

Replica model for an unusual directed polymer in 1+1 dimensions and prediction of the extremal parameter of gapped sequence alignment statistics.

Sequence alignment is one of the most important bioinformatics tools for modern molecular biology. The statistical characterization of gapped alignment scores has been a long-standing problem in sequence alignment research. In this paper, we provide a self-contained exposition of sequence alignment, a short review about how this problem is related to the directed polymer problem in statistical physics, and some analytical results that can be used for predicting alignment score statistics. Basically, we present two classes of solutions for the gapped alignment statistics by explicitly calculating the evolution of the few-replica partition function in 1+1 dimensions. We have obtained the conditions under which the more important extremal parameter lambda, characterizing the alignment score statistics, becomes predictable.

Algorithms↗

Scale-free networks versus evolutionary drift.

Recent studies of properties of various biological networks revealed that many of them display scale-free characteristics. Since the theory of scale-free networks is applicable to evolving networks, one can hope that it provides not only a model of a biological network in its current state but also sheds some insight into the evolution of the network. In this work, we investigate the probability distributions and scaling properties underlying some models for biological networks and protein domain evolution. The analysis of evolutionary models for domain similarity networks indicates that models which include evolutionary drift are typically not scale free. Instead they adhere quite closely to the Yule distribution. This finding indicates that the direct applicability of scale-free models in understanding the evolution of biological network may not be as wide as it has been hoped for.

Computational Biology↗

The compositional adjustment of amino acid substitution matrices.

Amino acid substitution matrices are central to protein-comparison methods. In most commonly used matrices, the substitution scores take a log-odds form, involving the ratio of "target" to "background" frequencies derived from large, carefully curated sets of protein alignments. However, such matrices often are used to compare protein sequences with amino acid compositions that differ markedly from the background frequencies used for the construction of the matrices. Of course, the target frequencies should be adjusted in such cases, but the lack of an appropriate way to do this has been a long-standing problem. This article shows that if one demands consistency between target and background frequencies, then a log-odds substitution matrix implies a unique set of target and background frequencies as well as a unique scale. Standard substitution matrices therefore are truly appropriate only for the comparison of proteins with standard amino acid composition. Accordingly, we present and evaluate a rationale for transforming the target frequencies implicit in a standard matrix to frequencies appropriate for a nonstandard context. This rationale yields asymmetric matrices for the comparison of proteins with divergent compositions. Earlier approaches are unable to deal with this case in a fully consistent manner. Composition-specific substitution matrix adjustment is shown to be of utility for comparing compositionally biased proteins, including those of organisms with nucleotide-biased, and therefore codon-biased, genomes or isochores.

Amino Acid Sequence↗

Reentrant synclinic phase in an electric-field-temperature-phase diagram for enantiomeric mixtures of an antiferroelectric liquid crystal.

The threshold electric field E(th) for a transition from the anticlinic to the synclinic phase of enantiomeric mixtures of the liquid crystal TFMHPOBC was measured as a function of temperature T and enantiomeric excess X. For small X the phase boundary curve on a temperature-electric-field phase diagram exhibits the phase sequence synclinic-anticlinic-reentrant synclinic on decreasing the temperature. At one point along the curve the quantity dT/dE--> infinity. For large values of enantiomeric excess a reentrant phase is not observed. The results are discussed using a simple phenomenological theory that accounts for layer-layer interactions, such that the electric-field-induced transition to the synclinic phase, although completed by solitary-wave propagation, is facilitated by a percolation mechanism.

Journal Article↗

Hybrid alignment: high-performance with universal statistics.

The score statistics of a recently introduced 'hybrid alignment' algorithm is studied in detail numerically. An extensive survey across the 2216 models of protein domains contained in the Pfam v5.4 database (Bateman et al., Nucleic Acids Res., 28, 263-266, 2000) verifies the theoretical predictions: For the position-specific scoring functions used in the Pfam models, the score statistics of hybrid alignment obey the Gumbel distribution, with the key Gumbel parameter lambda taking on the asymptotic value 1 universally for all models. Thus, the use of hybrid alignment eliminates the time-consuming computer simulations normally needed to assign p-values to alignment scores, freeing the users to experiment with different scoring parameters and functions. The performance of the hybrid algorithm in detecting sequence homology is also studied. For protein sequences from the SCOP database (Murzin et al., J. Mol. Biol., 247, 536-540, 1995) using uniform scoring functions, the performance is found to be comparable to the best of the existing methods. Preliminary results using the PfamA database suggest that the hybrid algorithm achieves similar performance as existing methods for position-specific scoring systems as well. Hybrid alignment is thereby established as a high performance alignment algorithm with well-characterized, universal statistics.

Algorithms↗