RECOMB98. Computational molecular biology: pre- and post-genomics, March 22-25, 1998, New York.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to D D Pollock.
Explore the source record for details and available documents.
The identification of protein sites undergoing correlated evolution (coevolution) is of great interest due to the possibility that these pairs will tend to be adjacent in the three-dimensional structure. Identification of such pairs should provide useful information for understanding the evolutionary process, predicting the effects of site-directed substitution, and potentially for predicting protein structure. Here, we develop and apply a maximum likelihood method with the aim of improving detection of coevolution. Unlike previous methods which have had limited success, this method allows for correlations induced by phylogenetic relationships and for variation in rate of evolution along branches, and does not rely on accurate reconstruction of ancestral nodes. In order to reduce the complexity of coevolutionary relationships and identify the primary component of pairwise coevolution between two sites, we reduce the data to a two-state system at each site, regardless of the actual number of residues observed at that site. Simulations show that this strategy is good at identifying simple correlations and at recognizing cases in which the data are insufficient to distinguish between coevolution and spurious correlations. The new method was tested by using size and charge characteristics to group the residues at each site, and then evaluating coevolution in myoglobin sequences. Grouping based on physicochemical characteristics allows categorization of coevolving sites into positive and negative coevolution, depending on the correlation between equilibrium state frequencies. We detected a striking excess of negative coevolution (corresponding to charge) at sites brought into proximity by the periodicity of the alpha-helix, and there was also a tendency for sites with significant likelihood ratios to be close in the three-dimensional structure. Sites on the surface of the protein appear to coevolve both when they are close in the structure, and when they are distant, implying a role for folding and/or avoidance of quaternary structure in the coevolution process.
Analytical molecular distance estimates can be inaccurate and biased estimates of the total number of substitutions not only when the model of evolution they are based on is incorrect, but also when the method of estimating the total is too simple. This comes about because when there are different types of substitutions occurring simultaneously, it can become extremely difficult to estimate the number of the more quickly evolving type, and the variance of this larger number can overwhelm the total estimate. In this paper, in an extension of earlier work with a simple two-parameter model of evolution, more accurate analytical distances are derived for models appropriate to a variety of known DNA types using generalized least squares principles of noise reduction. It is shown that the new estimates can be applied to achieve more accurate results for site-to-site rate variation, regions with biased nucleotide frequencies, and synonymous sites in protein-coding regions. This study also includes a methodology to obtain accurate distance estimates for large numbers of sequence regions evolving in different manners.
A symmetric stepwise mutation model with reflecting boundaries is employed to evaluate microsatellite evolution under range constraints. Methods of estimating range constraints and mutation rates under the assumptions of the model are developed. Least squares procedures are employed to improve molecular distance estimation for use in phylogenetic reconstruction in the case where range constraints and mutation rates vary across loci. The bias and accuracy of these methods are evaluated using computer simulations, and they are compared to previously existing methods which do not assume range constraints. Range constraints are seen to have a substantial impact on phylogenetic conclusions based on molecular distances, particularly for more divergent taxa. Results indicate that if range constraints are in effect, the methods developed here should be used in both the preliminary planning and final analysis of phylogenetic studies employing microsatellites. It is also seen that in order to make accurate phylogenetic inferences under range constraints, a larger number of loci are required than in their absence.
Statistical properties of the symmetric stepwise-mutation model for microsatellite evolution are studied under the assumption that the number of repeats is strictly bounded above and below. An exact analytic expression is found for the expected products of the frequencies of alleles separated by k repeats. This permits characterization of the asymptotic behavior of our distances D1 and (delta mu)2 under range constraints. Based on this characterization we develop transformations that partially restore linearity when allele size is restricted. We show that the appropriate transformation cannot be applied in the case of varying mutation rates (beta) and range constraints (R) because of statistical difficulties. In the special case of no variation in beta and R across loci, however, the transformation simplifies to a usable form and results in a distance much more linear with time than distances developed for an infinite range. Although analytically incorrect in the case of variation in beta and R, the simpler transformation is surprisingly insensitive to variation in these parameters, suggesting that it may have considerable utility in phylogenetic studies.
Various methods for detecting correlation between sites were evaluated by ascertaining their ability to discriminate positively correlated sites from background correlation at randomly evolved sites. A model for generating pairwise correlations of different degrees is also described. An assortment of physicochemical vectors and similarity and difference matrices were used to discriminate correlated change. There was little difference in effectiveness between the different matrices, but there were significant differences between the matrices and the physicochemical vectors. It is shown that all methods investigated exhibit significant inability to screen out background correlation, particularly in the presence of phylogenetic relatedness between the sequences. Methods using the matrices are unable to distinguish positively correlated from negatively correlated, or compensatory, replacements.
Since the initial work of Jukes and Cantor (1969), a number of procedures have been developed to estimate the expected number of nucleotide substitutions corresponding to a given observed level of nucleotide differentiation assuming particular evolutionary models. Unlike the proportion of different sites, the expected number of substitutions that would have occurred grows linearly with time and therefore has had great appeal as an evolutionary distance. Recently, however, a number of authors have tried to develop improved statistical approaches for generating and evaluating evolutionary distances (Schöniger and von Haeseler 1993; Goldstein and Polock 1994; Tajima and Takezaki 1994). These studies clearly show that the estimated number of nucleotide substitutions is generally not the best estimator for use in reconstruction of phylogenetic relationships. The reason for this is that there is often a large error associated with the estimation of this number. Therefore, even though its expectation is correct (i.e., on average the expected number of substitutions is proportional to time--but see Tajima 1993), it is not expected to be as useful as estimators designed to have a lower variance.
Gene duplication has produced two lactate dehydrogenase (LDH) isozymes, LDH-A and LDH-B, that are found in essentially all vertebrates. On the basis of the biochemical properties of the LDH-A and LDH-B isozymes, it has been suggested that each locus is orthologous among all vertebrates. However, phylogenetic studies have not supported a common evolutionary history among the LDH-A isozymes, particularly when those from lower vertebrates are examined. We present here the sequence of a muscle-type LDH from Fundulus heteroclitus, a teleost fish for which the LDH-B sequence has been determined and shown to be unrelated phylogenetically to tetrapod LDH-A isozymes. Although the sequence of the teleost muscle LDH shares certain features with the LDH-A of tetrapods, phylogenetic analyses do not support an orthologous relation among the LDH-A isozymes of teleost fish and tetrapod vertebrates.
Zuckerkandl and Pauling (1962, "Horizons in Biochemistry," pp. 189-225, Academic Press, New York) first noticed that the degree of sequence similarity between the proteins of different species could be used to estimate their phylogenetic relationship. Since then models have been developed to improve the accuracy of phylogenetic inferences based on amino acid or DNA sequences. Most of these models were designed to yield distance measures that are linear with time, on average. The reliability of phylogenetic reconstruction, however, depends on the variance of the distance measure in addition to its expectation. In this paper we show how the method of generalized least squares can be used to combine data types, each most informative at different points in time, into a single distance measure. This measure reconstructs phylogenies more accurately than existing non-likelihood distance measures. We illustrate the approach for a two-rate mutation model and demonstrate that its application provides more accurate phylogenetic reconstruction than do currently available analytical distance measures.
Phosphate-activated glutaminase is found in mammalian small intestine, brain, and kidney, but not in liver. The enzyme initiates the catabolism of glutamine as the principal respiratory fuel in the small intestine, may synthesize the neurotransmitter glutamate in the brain, and functions in the kidney to help maintain systemic pH homeostasis. Interleukin-9 (IL9) is a relatively new cytokine that supports the growth of helper T-cell clones, mast cells, and megakaryoblastic leukemia cells. cDNA clones have recently been obtained for each of these genes. The human loci for phosphate-activated glutaminase (GLS) and IL9 have previously been mapped to chromosomes 2 and 5, respectively, by analysis of somatic cell hybrid DNAs. By using chromosomal in situ hybridization, we have regionally mapped GLS to 2q32----q34 and IL9 to 5q31----q35.
Explore the source record for details and available documents.
Explore the source record for details and available documents.