PubMed Health⌕ Search

Biomedical subjects

Richard A Goldstein

Publications and source records attributed to Richard A Goldstein.

18 recordsLinked to original sources

Emergent robustness in competition between autocatalytic chemical networks.

The origin of auto-catalytic networks has been proposed as an initial step in pre-biotic evolution. It is possible to derive simple models where auto-catalytic networks naturally arise from simple chemical mixtures. In order for such a system to develop, there needs to be some degree of stability, what is characterised as ;robustness'. We demonstrate that competing systems generate this robustness as they create a distributed network of catalytic pathways.

Computer Simulation↗

Assessing the accuracy of ancestral protein reconstruction methods.

The phylogenetic inference of ancestral protein sequences is a powerful technique for the study of molecular evolution, but any conclusions drawn from such studies are only as good as the accuracy of the reconstruction method. Every inference method leads to errors in the ancestral protein sequence, resulting in potentially misleading estimates of the ancestral protein's properties. To assess the accuracy of ancestral protein reconstruction methods, we performed computational population evolution simulations featuring near-neutral evolution under purifying selection, speciation, and divergence using an off-lattice protein model where fitness depends on the ability to be stable in a specified target structure. We were thus able to compare the thermodynamic properties of the true ancestral sequences with the properties of "ancestral sequences" inferred by maximum parsimony, maximum likelihood, and Bayesian methods. Surprisingly, we found that methods such as maximum parsimony and maximum likelihood that reconstruct a "best guess" amino acid at each position overestimate thermostability, while a Bayesian method that sometimes chooses less-probable residues from the posterior probability distribution does not. Maximum likelihood and maximum parsimony apparently tend to eliminate variants at a position that are slightly detrimental to structural stability simply because such detrimental variants are less frequent. Other properties of ancestral proteins might be similarly overestimated. This suggests that ancestral reconstruction studies require greater care to come to credible conclusions regarding functional evolution. Inferred functional patterns that mimic reconstruction bias should be reevaluated.

Algorithms↗

Observations of amino acid gain and loss during protein evolution are explained by statistical bias.

The authors of a recent manuscript in "Nature" claim to have discovered "universal trends" of amino acid gain and loss in protein evolution. Here, we show that this universal trend can be simply explained by a bias that is unavoidable with the 3-taxon trees used in the original analysis. We demonstrate that a rigorously reversible equilibrium model, when analyzed with the same methods as the "Nature" manuscript, yields identical (and in this case, clearly erroneous) conclusions. A main source of the bias is the division of the sequence data into "informative" and "noninformative" sites, which favors the observation of certain transitions.

Algorithms↗

Divergence, recombination and retention of functionality during protein evolution.

We have only a vague idea of precisely how protein sequences evolve in the context of protein structure and function. This is primarily because structural and functional contexts are not easily predictable from the primary sequence, and evaluating patterns of evolution at individual residue positions is also difficult. As a result of increasing biodiversity in genomics studies, progress is being made in detecting context-dependent variation in substitution processes, but it remains unclear exactly what context-dependent patterns we should be looking for. To address this, we have been simulating protein evolution in the context of structure and function using lattice models of proteins and ligands (or substrates). These simulations include thermodynamic features of protein stability and population dynamics. We refer to this approach as 'ab initio evolution' to emphasise the fact that the equilibrium details of fitness distributions arise from the physical principles of the system and not from any preconceived notions or arbitrary mathematical distributions. Here, we present results on the retention of functionality in homologous recombinants following population divergence. A central result is that protein structure characteristics can strongly influence recombinant functionality. Exceptional structures with many sequence options evolve quickly and tend to retain functionality--even in highly diverged recombinants. By contrast, the more common structures with fewer sequence options evolve more slowly, but the fitness of recombinants drops off rapidly as homologous proteins diverge. These results have implications for understanding viral evolution, speciation and directed evolutionary experiments. Our analysis of the divergence process can also guide improved methods for accurately approximating folding probabilities in more complex but realistic systems.

Evolution, Molecular↗

Predicting functional sites in proteins: site-specific evolutionary models and their application to neurotransmitter transporters.

Currently there exist several computational methods for predicting the functional sites in a set of homologous proteins based on their sequences. Due to difficulties in defining the functional site in a protein, it is not trivial to compare the performance of these methods, evaluate their limitations and quantify improvements by new approaches. Here, we use extensive mutation data from two proteins, Lac repressor and subtilisin, to perform such an analysis. Along with the evaluation of existing approaches, we describe a site class model of evolution as a tool to predict functional sites in proteins. The results indicate that this model, which simulates the evolution process at the amino acid level using site-specific substitution matrices, provides the most accurate information on functional sites in a given protein family. Secondly, we present an application of this model to neurotransmitter transporters, a superfamily of proteins of which we have limited experimental knowledge. Based on this application we present testable hypotheses regarding the mechanism of action of these proteins.

Amino Acid Transport Systems↗

Performance of an iterated T-HMM for homology detection.

MOTIVATION: Much information about new protein sequences is derived from identifying homologous proteins. Such tasks are difficult when the evolutionary relationships are distant. Some modern methods achieve better results by building a model of a set of related sequences, and then identifying new proteins that fit the model. A further advance was the development of iterative methods that refine the model as more homologs are discovered. These methods are generally limited by ad hoc methods of sequence weighting, neglect of underlying evolutionary relationships and the representation of the set with a single one-size-fits-all model. These limitations are avoided through the use of a Tree hidden Markov model (T-HMM) approach. Our previous work described how a non-iterative version of the T-HMM method could identify distant homologs with superior performance compared with other non-iterated approaches, and described how this method was particularly appropriate for being implemented as an iterative algorithm. RESULTS: We describe an iterative version of the T-HMM algorithm, and evaluate its performance for the detection of distant homologs. Significant improvement over other commonly used methods is found. AVAILABILITY: The software (C++, Perl) is available from the corresponding author.

Algorithms↗

Dimerization in aminergic G-protein-coupled receptors: application of a hidden-site class model of evolution.

G-Protein-coupled receptors (GPCRs) are an important superfamily of transmembrane proteins involved in cellular communication. Recently, it has been shown that dimerization is a widely occurring phenomenon in the GPCR superfamily, with likely important physiological roles. Here we use a novel hidden-site class model of evolution as a sequence analysis tool to predict possible dimerization interfaces in GPCRs. This model aims to simulate the evolution of proteins at the amino acid level, allowing the analysis of their sequences in an explicitly evolutionary context. Applying this model to aminergic GPCR sequences, we first validate the general reasoning behind the model. We then use the model to perform a family specific analysis of GPCRs. Accounting for the family structure of these proteins, this approach detects different evolutionarily conserved and accessible patches on transmembrane (TM) helices 4-6 in different families. On the basis of these findings, we propose an experimentally testable dimerization mechanism, involving interactions among different combinations of these helices in different families of aminergic GPCRs.

Amino Acid Substitution↗

Depicting a protein's two faces: GPCR classification by phylogenetic tree-based HMMs.

Related proteins with similar biological functions generally share common features, allowing us to extract the common sequence features. These common features enable us to build statistical models that can be used to classify proteins, to predict new members, and to study the sequence-function relationship of this protein function group. Although evolution underlies the basis of multiple sequence analysis methods, most methods ignore phylogenetic relationships and the evolutionary process in building these statistical models. Previously we have shown that a phylogenetic tree-based profile hidden Markov model (T-HMM) is superior in generating a profile for a group of similar proteins. In this study we used the method to generate common features of G protein-coupled receptors (GPCRs). The profile generated by T-HMM gives high accuracy in GPCR function classification, both by ligand and by coupled G protein.

Animals↗

Probing conformational changes in neurotransmitter transporters: a structural context.

The Na+/Cl-dependent neurotransmitter transporters, a family of proteins responsible for the reuptake of neurotransmitters and other small molecules from the synaptic cleft, have been the focus of intensive research in recent years. The biogenic amine transporters, a subset of this larger family, are especially intriguing as they are the targets for many psychoactive compounds, including cocaine and amphetamines, as well as many antidepressants. In the absence of a high-resolution structure for any transporter in this family, research into the structure-function relationships of these transporters has relied on analysis of the effects of site-directed mutagenesis as well as of chemical modification of reactive residues. The aim of this review is to establish a structural context for the experimental study of these transporters through various computational approaches and to highlight what is known about the conformational changes associated with function in these transporters. We also present a novel numbering scheme to assist in the comparison of aligned positions between sequences of the neurotransmitter transporter family, a comparison that will be of increasing importance as additional experimental data is amassed.

Amino Acid Sequence↗

Detecting distant homologs using phylogenetic tree-based HMMs.

It is often desired to identify further homologs of a family of biological sequences from the ever-growing sequence databases. Profile hidden Markov models excel at capturing the common statistical features of a group of biological sequences. With these common features, we can search the biological database and find new homologous sequences. Most general profile hidden Markov model methods, however, treat the evolutionary relationships between the sequences in a homologous group in an ad-hoc manner. We hereby introduce a method to incorporate phylogenetic information directly into hidden Markov models, and demonstrate that the resulting model performs better than most of the current multiple sequence-based methods for finding distant homologs.

Algorithms↗

Optimization of a new score function for the generation of accurate alignments.

The accuracy of the alignments of protein sequences depends on the score matrix and gap penalties used in performing the alignment. Most score functions are designed to find homologs in the various databases rather than to generate accurate alignments between known homologs. We describe the optimization of a score function for the purpose of generating accurate alignments, as evaluated by using a coordinate root-mean-square deviation (RMSD)-based merit function. We show that the resulting score matrix, which we call STROMA, generates more accurate alignments than other commonly used score matrices, and this difference is not due to differences in the gap penalties. In fact, in contrast to most of the other matrices, the alignment accuracies with STROMA are relatively insensitive to the choice of gap penalty parameters.

Amino Acid Sequence↗

Performance evaluation of a new algorithm for the detection of remote homologs with sequence comparison.

A detailed analysis of the performance of hybrid, a new sequence alignment algorithm developed by Yu and coworkers that combines Smith Waterman local dynamic programming with a local version of the maximum-likelihood approach, was made to access the applicability of this algorithm to the detection of distant homologs by sequence comparison. We analyzed the statistics of hybrid with a set of nonhomologous protein sequences from the SCOP database and found that the statistics of the scores from hybrid algorithm follows an Extreme Value Distribution with lambda approximately 1, as previously shown by Yu et al. for the case of artificially generated sequences. Local dynamic programming was compared to the hybrid algorithm by using two different test data sets of distant homologs from the PFAM and COGs protein sequence databases. The studies were made with several score functions in current use including OPTIMA, a new score function originally developed to detect remote homologs with the Smith Waterman algorithm. We found OPTIMA to be the best score function for both both dynamic programming and the hybrid algorithms. The ability of dynamic programming to discriminate between homologs and nonhomologs in the two sets of distantly related sequences is slightly better than that of hybrid algorithm. The advantage of producing accurate score statistics with only a few simulations may overcome the small differences in performance and make this new algorithm suitable for detection of homologs in conjunction with a wide range of score functions and gap penalties.

Algorithms↗

Why are proteins so robust to site mutations?

There have been repeated observations that proteins are surprisingly robust to site mutations, enduring significant numbers of substitutions with little change in structure, stability, or function. These results are almost paradoxical in light of what is known about random heteropolymers and the sensitivity of their properties to seemingly trivial mutations. To address this discrepancy, the preservation of biological protein properties in the presence of mutation has been interpreted as indicating the independence of selective pressure on such properties. Such results also lead to the prediction that de novo protein design should be relatively easy, in contrast to what is observed. Here, we use a computational model with lattice proteins to demonstrate how this robustness can result from population dynamics during the evolutionary process. As a result, sequence plasticity may be a characteristic of evolutionarily derived proteins and not necessarily a property of designed proteins. This suggests that this robustness must be re-interpreted in evolutionary terms, and has consequences for our understanding of both in vivo and in vitro protein evolution.

Computer Simulation↗

Why are proteins marginally stable?

Most globular proteins are marginally stable regardless of size or activity. The most common interpretation is that proteins must be marginally stable in order to function, and so marginal stability represents the results of positive selection. We consider the issue of marginal stability directly using model proteins and the dynamical aspects of protein evolution in populations. We find that the marginal stability of proteins is an inherent property of proteins due to the high dimensionality of the sequence space, without regard to protein function. In this way, marginal stability can result from neutral, non-adaptive evolution. By allowing evolving protein sub-populations with different stability requirements for functionality to complete, we find that marginally stable populations of proteins tend to dominate. Our results show that functionalities consistent with marginal stability have a strong evolutionary advantage, and might arise because of the natural tendency of proteins towards marginal stability.

Models, Chemical↗

rtREV: an amino acid substitution matrix for inference of retrovirus and reverse transcriptase phylogeny.

Retroviral and other reverse transcriptase (RT)-containing sequences may be subject to unique evolutionary pressures, and models of molecular sequence evolution developed using other kinds of sequences may not be optimal. Here we develop and present a new substitution matrix for maximum likelihood (ML) phylogenetic analysis which has been optimized on a dataset of 33 amino acid sequences from the retroviral Pol proteins. When compared to other matrices, this model (rtREV) yields higher log-likelihood values on a range of datasets including lentiviruses, spumaviruses, betaretroviruses, gammaretroviruses, and other elements containing reverse transcriptase. We provide evidence that rtREV is a more realistic evolutionary model for analyses of the pol gene, although it is inapplicable to analyses involving the gag gene.

Amino Acid Substitution↗

Using evolutionary methods to study G-protein coupled receptors.

A novel method to analyze evolutionary change is presented and its application to the analysis of sequence data is discussed. The investigated method uses phylogenetic trees of related proteins with an evolutionary model in order to gain insight about protein structure and function. The evolutionary model, based on amino acid substitutions, contains adjustable parameters related to amino acid and sequence properties. A maximum likelihood approach is used with a phylogenetic tree to optimize these parameters. The model is applied to a set of Muscarinic receptors, members of the G-protein coupled receptor family. Here we show that the optimized parameters of the model are able to highlight the general structural features of these receptors.

Animals↗