PubMed HealthSearch

Biomedical subjects

T R Ioerger

Publications and source records attributed to T R Ioerger.

3 recordsLinked to original sources

The context-dependence of amino acid properties.

One of the current limitations of using sequence alignments to identify proteins with similar structures is that some proteins with similar structures do not have significant sequence similarity by identity. One way to address this "hidden-homology" problem is to match amino acids based on their chemical and physical properties. However, the amino acid properties overlap, creating orthogonal dimensions of similarity, the relative strengths of which are ambiguous. It has been observed that the role an amino acid plays (and hence the property that is important) at a site in a protein depends on its secondary and tertiary environment. To approximate and take advantage of this dependence on context for improving the sensitivity of alignments of proteins whose structures are unknown, we propose a surrogate definition of context based on the pattern of hydropathy in a small window of contiguous neighbors surrounding each amino acid. We present the results of an experiment in which a search-based program iteratively tests and selects various properties in independent contexts, and incrementally increases the ability of sequence alignments to detect relationships among distantly-related proteins. The method is shown to perform better than using the MDM78 substitution table for partial match scores.

Amino Acid Sequence

Constructive induction and protein tertiary structure prediction.

To date, the only methods that have been used successfully to predict protein structures have been based on identifying homologous proteins whose structures are known. However, such methods are limited by the fact that some proteins have similar structure but no significant sequence homology. We consider two ways of applying machine learning to facilitate protein structure prediction. We argue that a straightforward approach will not be able to improve the accuracy of classification achieved by clustering by alignment scores alone. In contrast, we present a novel constructive induction approach that learns better representations of amino acid sequences in terms of physical and chemical properties. Our learning method combines knowledge and search to shift the representation of sequences so that semantic similarity is more easily recognized by syntactic matching. Our approach promises not only to find new structural relationships among protein sequences, but also expands our understanding of the roles knowledge can play in learning via experience in this challenging domain.

Artificial Intelligence

Polymorphism at the self-incompatibility locus in Solanaceae predates speciation.

Sequences of 11 alleles of the gametophytic self-incompatibility locus (S locus) from three species of the Solanaceae family have recently been determined. Pairwise comparisons of these alleles reveal two unexpected observations: (i) amino acid sequence similarity can be as low as 40% within species and (ii) some interspecific similarities are higher than intraspecific similarities. The gene genealogy clearly illustrates this unusual pattern of relationships. The data suggest that some of the polymorphism at the S locus existed prior to the divergence of these species and has been maintained to the present. In support of this hypothesis, the number of shared polymorphic sites was found to exceed the number found in simulations with independent accumulation of mutations. Strictly neutral evolution is exceedingly unlikely to maintain the polymorphism for such a long time. The allele multiplicity and extreme age of the alleles is consistent with Wright's classic one-locus population genetic model of gametophytic self-incompatibility. Similarities between the plant S locus and the mammalian major histocompatibility complex are discussed.

Alleles