PubMed Health⌕ Search

Biomedical subjects

John Moult

Publications and source records attributed to John Moult.

At least 19 recordsLinked to original sources

Modeling Alternative Conformational States in CASP16.

The CASP16 Ensemble Prediction experiment assessed advances in methods for modeling proteins, nucleic acids, and their complexes in multiple conformational states. Targets included systems with experimental structures determined in two or three states, evaluated by direct comparison to experimental coordinates, as well as domain-linker-domain (D-L-D) targets assessed against statistical models from NMR and SAXS data. This paper focuses on the former class of multi-state targets. Ten ensembles were released as community challenges, including ligand-induced conformational changes, protein-DNA complexes, a trimeric protein, a stem-loop RNA, and multiple oligomeric states of a single RNA. For five targets, some groups produced reasonably accurate models of both reference states (best TM-score >0.75). However, with the exception of one protein-ligand complex (T1214), where an apo structure was available as a template, predictors generally failed to capture key structural details distinguishing the states. Overall, accuracy was significantly lower than for single-state targets in other CASP experiments. The most successful approaches generated multiple AlphaFold2 models using enhanced multiple sequence alignments and sampling protocols, followed by model quality based selection. While the AlphaFold3 server performed well on several targets, individual groups outperformed it in specific cases. By contrast, predictions for one protein-DNA complex, three RNA targets, and multiple oligomeric RNA states consistently fell short (TM-score <0.75). These results highlight both progress and persistent challenges in multi-state prediction. Despite recent advances, accurate modeling of conformational ensembles, particularly RNA and large multimeric assemblies, remains a critical frontier for structural biology.

AlphaFold2↗

Detection of operons.

Operons are clusters of genes that are transcribed as a single message, and regulated by the same gene expression machinery. They are found primarily in prokaryotic genomes. Because genes in the same operon are likely to have related functions, identification of the operon structure is potentially useful for assigning gene function. We report the development and benchmarking of two different methods for detecting operons, based on an analysis of 42 fully sequenced prokaryotic organisms. The Gene Neighbor method (GNM) utilizes the relatively high conservation of gene order in operons, compared with genes in general. The Gene Gap Method (GGM) makes use of the relatively short gap between genes in operons compared with that otherwise found between adjacent genes. The methods have been benchmarked using KEGG pathway data and RegulonDB Escherichia coli operon data. With optimum parameters, the specificity of the GNM is 93% and the sensitivity is 70%. For the GGM, the specificity is 95% and the sensitivity is 68%. Together, the two methods have a sensitivity of 87.2%, while joint predictions have a sensitivity of 50% and a specificity of 98%. The methods are used to infer possible functions for some hypothetical genes in prokaryotic genomes. The methods have proven a useful addition to structure information in deriving protein function in a structural genomics project.

Computational Biology↗

Towards computing with proteins.

Can proteins be used as computational devices to address difficult computational problems? In recent years there has been much interest in biological computing, that is, building a general purpose computer from biological molecules. Most of the current efforts are based on DNA because of its ability to self-hybridize. The exquisite selectivity and specificity of complex protein-based networks motivated us to suggest that similar principles can be used to devise biological systems that will be able to directly implement any logical circuit as a parallel asynchronous computation. Such devices, powered by ATP molecules, would be able to perform, for medical applications, digital computation with natural interface to biological input conditions. We discuss how to design protein molecules that would serve as the basic computational element by functioning as a NAND logical gate, utilizing DNA tags for recognition, and phosphorylation and exonuclease reactions for information processing. A solution of these elements could carry out effective computation. Finally, the model and its robustness to errors were tested in a computer simulation.

Adenosine Triphosphate↗

Rigorous performance evaluation in protein structure modelling and implications for computational biology.

In principle, given the amino acid sequence of a protein, it is possible to compute the corresponding three-dimensional structure. Methods for modelling structure based on this premise have been under development for more than 40 years. For the past decade, a series of community wide experiments (termed Critical Assessment of Structure Prediction (CASP)) have assessed the state of the art, providing a detailed picture of what has been achieved in the field, where we are making progress, and what major problems remain. The rigorous evaluation procedures of CASP have been accompanied by substantial progress. Lessons from this area of computational biology suggest a set of principles for increasing rigor in the field as a whole.

Computational Biology↗

SNPs3D: candidate gene and SNP selection for association studies.

BACKGROUND: The relationship between disease susceptibility and genetic variation is complex, and many different types of data are relevant. We describe a web resource and database that provides and integrates as much information as possible on disease/gene relationships at the molecular level. DESCRIPTION: The resource http://www.SNPs3D.org has three primary modules. One module identifies which genes are candidates for involvement in a specified disease. A second module provides information about the relationships between sets of candidate genes. The third module analyzes the likely impact of non-synonymous SNPs on protein function. Disease/candidate gene relationships and gene-gene relationships are derived from the literature using simple but effective text profiling. SNP/protein function relationships are derived by two methods, one using principles of protein structure and stability, the other based on sequence conservation. Entries for each gene include a number of links to other data, such as expression profiles, pathway context, mouse knockout information and papers. Gene-gene interactions are presented in an interactive graphical interface, providing rapid access to the underlying information, as well as convenient navigation through the network. Use of the resource is illustrated with aspects of the inflammatory response and hypertension. CONCLUSION: The combination of SNP impact analysis, a knowledge based network of gene relationships and candidate genes, and access to a wide range of data and literature allow a user to quickly assimilate available information, and so develop models of gene-pathway-disease interaction.

Databases, Genetic↗

Identification and analysis of deleterious human SNPs.

We have developed two methods of identifying which non-synonomous single base changes have a deleterious effect on protein function in vivo. One method, described elsewhere, analyzes the effect of the resulting amino acid change on protein stability, utilizing structural information. The other method, introduced here, makes use of the conservation and type of residues observed at a base change position within a protein family. A machine learning technique, the support vector machine, is trained on single amino acid changes that cause monogenic disease, with a control set of amino acid changes fixed between species. Both methods are used to identify deleterious single nucleotide polymorphisms (SNPs) in the human population. After carefully controlling for errors, we find that approximately one quarter of known non-synonymous SNPs are deleterious by these criteria, providing a set of possible contributors to human complex disease traits.

Animals↗

Loss of protein structure stability as a major causative factor in monogenic disease.

The most common cause of monogenic disease is a single base DNA variant resulting in an amino acid substitution. In a previous study, we observed that a high fraction of these substitutions appear to result in reduction of stability of the corresponding protein structure. We have now investigated this phenomenon more fully. A set of structural effects, such as reduction in hydrophobic area, overpacking, backbone strain, and loss of electrostatic interactions, is used to represent the impact of single residue mutations on protein stability. A support vector machine (SVM) was trained on a set of mutations causative of disease, and a control set of non-disease causing mutations. In jack-knifed testing, the method identifies 74% of disease mutations, with a false positive rate of 15%. Evaluation of a set of in vitro mutagenesis data with the SVM established that the majority of disease mutations affect protein stability by 1 to 3 kcal/mol. The method's effective distinction between disease and non-disease variants, strongly supports the hypothesis that loss of protein stability is a major factor contributing to monogenic disease.

Amino Acid Substitution↗

Protein family clustering for structural genomics.

A major goal of structural genomics is the provision of a structural template for a large fraction of protein domains. The magnitude of this task depends on the number and nature of protein sequence families. With a large number of bacterial genomes now fully sequenced, it is possible to obtain improved estimates of the number and diversity of families in that kingdom. We have used an automated clustering procedure to group all sequences in a set of genomes into protein families. Bench-marking shows the clustering method is sensitive at detecting remote family members, and has a low level of false positives. This comprehensive protein family set has been used to address the following questions. (1) What is the structure coverage for currently known families? (2) How will the number of known apparent families grow as more genomes are sequenced? (3) What is a practical strategy for maximizing structure coverage in future? Our study indicates that approximately 20% of known families with three or more members currently have a representative structure. The study indicates also that the number of apparent protein families will be considerably larger than previously thought: We estimate that, by the criteria of this work, there will be about 250,000 protein families when 1000 microbial genomes have been sequenced. However, the vast majority of these families will be small, and it will be possible to obtain structural templates for 70-80% of protein domains with an achievable number of representative structures, by systematically sampling the larger families.

Genomics↗

The psychrophilic lifestyle as revealed by the genome sequence of Colwellia psychrerythraea 34H through genomic and proteomic analyses.

The completion of the 5,373,180-bp genome sequence of the marine psychrophilic bacterium Colwellia psychrerythraea 34H, a model for the study of life in permanently cold environments, reveals capabilities important to carbon and nutrient cycling, bioremediation, production of secondary metabolites, and cold-adapted enzymes. From a genomic perspective, cold adaptation is suggested in several broad categories involving changes to the cell membrane fluidity, uptake and synthesis of compounds conferring cryotolerance, and strategies to overcome temperature-dependent barriers to carbon uptake. Modeling of three-dimensional protein homology from bacteria representing a range of optimal growth temperatures suggests changes to proteome composition that may enhance enzyme effectiveness at low temperatures. Comparative genome analyses suggest that the psychrophilic lifestyle is most likely conferred not by a unique set of genes but by a collection of synergistic changes in overall genome content and amino acid composition.

Amino Acids↗

Critical assessment of methods of protein structure prediction (CASP)--round 6.

This article is an introduction to the special issue of the journal Proteins, dedicated to the sixth CASP experiment to assess the state of the art in protein structure prediction. The article describes the conduct of the experiment and the categories of prediction included, and outlines the evaluation and assessment procedures. A brief summary of progress over the decade of CASP experiments is also provided.

Algorithms↗

Progress over the first decade of CASP experiments.

CASP has now completed a decade of monitoring the state of the art in protein structure prediction. The quality of structure models produced in the latest experiment, CASP6, has been compared with that in earlier CASPs. Significant although modest progress has again been made in the fold recognition regime, and cumulatively, progress in this area is impressive. Models of previously unknown folds again appear to have modestly improved, and several mixed alpha/beta structures have been modeled in a topologically correct manner. Progress remains hard to detect in high sequence identity comparative modeling, but server performance in this area has moved forward.

Algorithms↗

A decade of CASP: progress, bottlenecks and prognosis in protein structure prediction.

For the past ten years, CASP (Critical Assessment of Structure Prediction) has monitored the state of the art in modeling protein structure from sequence. During this period, there has been substantial progress in both comparative modeling of structure (using information from an evolutionarily related structural template) and template-free modeling. The quality of comparative models depends on the closeness of the evolutionary relationship on which they are based. Template-free modeling, although still very approximate, now produces topologically near correct models for some small proteins. Current major challenges are refining comparative models so that they match experimental accuracy, obtaining accurate sequence alignments for models based on remote evolutionary relationships, and extending template-free modeling methods so that they produce more accurate models, handle parts of comparative models not available from a template and deal with larger structures.

Algorithms↗

Molecular modeling of protein function regions.

Experimental protein structures often provide extensive insight into the mode and specificity of small molecule binding, and this information is useful for understanding protein function and for the design of drugs. We have performed an analysis of the reliability with which ligand-binding information can be deduced from computer model structures, as opposed to experimentally derived ones. Models produced as part of the CASP experiments are used. The accuracy of contacts between protein model atoms and experimentally determined ligand atom positions is the main criterion. Only comparative models are included (i.e., models based on a sequence relationship between the protein of interest and a known structure). We find that, as expected, contact errors increase with decreasing sequence identity used as a basis for modeling. Analysis of the causes of errors shows that sequence alignment errors between model and experimental template have the most deleterious effect. In general, good, but not perfect, insight into ligand binding can be obtained from models based on a sequence relationship, providing there are no alignment errors in the model. The results support a structural genomics strategy based on experimental sampling of structure space so that all protein domains can be modeled on the basis of 30% or higher sequence identity.

Alpha-Globulins↗

Three-dimensional structural location and molecular functional effects of missense SNPs in the T cell receptor Vbeta domain.

The mechanisms by which human single nucleotide polymorphisms (SNPs) influence susceptibility to disease are not yet well understood. In a previous study, we developed a structure-based model that may be used to identify which missense SNPs are neutral and which are deleterious to protein function and so potentially involved in disease (Wang and Moult, Hum Mutat 2001;263-270). The model has now been applied to a set of 54 missense cSNPs in the 46 functional T-cell receptor Vbeta-genes. Most of these missense cSNPs are found to be neutral, but 10 are identified as likely deleterious to protein function. Only one was previously associated with disease. We suggest that the others may be disease related but that redundancy in the T-cell response prevents any simple, monogenic effect. Therefore, these SNPs are the most likely contributors to complex, polygenic disease traits. It has been noted that there is a surprisingly high (74%) fraction of nonsynonymous SNPs in these genes. Contrary to expectation, the analysis shows that these are not associated with an unusually high fraction of deleterious SNPs, nor do they significantly contribute to a larger range of antigen recognition or a reduced superantigen-binding repertoire.

Binding Sites↗

CAPRI: a Critical Assessment of PRedicted Interactions.

CAPRI is a communitywide experiment to assess the capacity of protein-docking methods to predict protein-protein interactions. Nineteen groups participated in rounds 1 and 2 of CAPRI and submitted blind structure predictions for seven protein-protein complexes based on the known structure of the component proteins. The predictions were compared to the unpublished X-ray structures of the complexes. We describe here the motivations for launching CAPRI, the rules that we applied to select targets and run the experiment, and some conclusions that can already be drawn. The results stress the need for new scoring functions and for methods handling the conformation changes that were observed in some of the target systems. CAPRI has already been a powerful drive for the community of computational biologists who development docking algorithms. We hope that this issue of Proteins will also be of interest to the community of structural biologists, which we call upon to provide new targets for future rounds of CAPRI, and to all molecular biologists who view protein-protein recognition as an essential process.

Algorithms↗

Assessment of progress over the CASP experiments.

The quality of structure models produced in the CASP5 experiment has been compared with that in earlier CASPs. The most significant progress is in the fold recognition regime, where the development of meta-servers has allowed more accurate consensus models to be generated. In contrast to this, there is little evidence of progress in producing more accurate comparative models, particularly those based on sequence identities > 30%. For comparative models based on low-sequence identity and for fold recognition models, accuracy depends primarily on the fraction of the target structure that is similar to an available template, and the quality of the alignment. Overall, these results indicate that there are still no effective methods of improving model quality beyond that obtained by successfully copying a template structure. For models of proteins with previously unknown folds, there appears to be a pause in the previous consistent improvement. There is some evidence that more groups are producing top-quality models, however. Although specific progress between successive experiments is sometimes difficulty to identify, over the history of all the CASPs there has been steady, if sometimes slow, progress in all modeling regimes.

Algorithms↗

Evaluation of disorder predictions in CASP5.

This paper reports an analysis of the accuracy of predictions of structural disorder received as part of the CASP5 experiment. Six groups made predictions of disorder. The predictions of the four most active groups have been compared with the experimental results, in terms of the sensitivity and specificity of the methods. All four methods succeed in detecting over half the disordered residues in the targets, with a generally low rate of over-prediction. Two of the methods perform significantly better when the structure of a related protein is available. There is a trade-off between the fraction of disordered residues detected and the extent of over-prediction, and groups have adopted different compromises in this respect. Comparison of performance at the same over-prediction rates highlights the role of related structures in some methods rather than others, with different groups achieving the highest sensitivity for different target sets. Over-all, the methods are clearly of considerable use in identifying potential disorder.

Computational Biology↗