PubMed Health⌕ Search

Biomedical subjects

Angel R Ortiz

Publications and source records attributed to Angel R Ortiz.

14 recordsLinked to original sources

Comparative binding energy analysis considering multiple receptors: a step toward 3D-QSAR models for multiple targets.

Comparative binding energy analysis, a technique to derive receptor-based three-dimensional quantitative structure-activity relationships (3D-QSAR), is herein extended to consider both affinity and selectivity in the derivation of the QSAR model. The extension is based on allowing multiple structurally related receptors to enter the X-matrix employed in the derivation of the structure-activity model. As a result, a single model common to all of them is obtained that considers both intra- and inter-receptor affinity differences for a given congeneric series. We applied the technique to a series of 88 3-amidinophenylalanines, binding to thrombin, trypsin, and factor Xa (fXa). A single predictive regression model for the three receptors involving 202 complexes, with a leave-one out (LOO) cross-validated Q(2) of 0.689, was obtained, and selectivity requirements were investigated. We find that total or partial occupancy of any of the three main pockets in the binding site (D-site, P-site, and the rim of the S1-site) leads to higher affinity across the family. However, the fact that thrombin can make stronger interactions in the P-site, as a result of its exclusive 60-loop, makes of this site a specificity pocket for this thrombin. Occupancy of the D-site leads to more active inhibitors toward fXa for the same reason, but the model does not highlight strongly the D-box because inhibitors are too short to fully occupy it. Negative charge density in the neighborhood of position 88 (a Lys insertion in thrombin) is found to be a determinant for thrombin recognition. These results were consistent with previous studies on selectivity in the thrombin/trypsin/fXa system.

Amino Acid Sequence↗

Computational approaches to model ligand selectivity in drug design.

To be effective, a designed drug must discriminate successfully the macromolecular target from alternative structures present in the organism. The last few years have witnessed the emergence of different computational tools aimed to the understanding and modeling of this process at molecular level. Although still rudimentary, these methods are shaping a coherent approach to help in the design of molecules with high affinity and specificity, both in lead discovery and in lead optimization. It is the purpose of this review to illustrate the array of computational tools available to consider selectivity in the design process, to summarize the most relevant applications, and to sketch the challenges ahead.

Amino Acid Sequence↗

Dual activation of pathways regulated by steroid receptors and peptide growth factors in primary prostate cancer revealed by Factor Analysis of microarray data.

BACKGROUND: We use an approach based on Factor Analysis to analyze datasets generated for transcriptional profiling. The method groups samples into biologically relevant categories, and enables the identification of genes and pathways most significantly associated to each phenotypic group, while allowing for the participation of a given gene in more than one cluster. Genes assigned to each cluster are used for the detection of pathways predominantly activated in that cluster by finding statistically significant associated GO terms. We tested the approach with a published dataset of microarray experiments in yeast. Upon validation with the yeast dataset, we applied the technique to a prostate cancer dataset. RESULTS: Two major pathways are shown to be activated in organ-confined, non-metastatic prostate cancer: those regulated by the androgen receptor and by receptor tyrosine kinases. A number of gene markers (HER3, IQGAP2 and POR1) highlighted by the software and related to the later pathway have been validated experimentally a posteriori on independent samples. CONCLUSION: Using a new microarray analysis tool followed by a posteriori experimental validation of the results, we have confirmed several putative markers of malignancy associated with peptide growth factor signalling in prostate cancer and revealed others, most notably ERRB3 (HER3). Our study suggest that, in primary prostate cancer, HER3, together or not with HER4, rather than in receptor complexes involving HER2, could play an important role in the biology of these tumors. These results provide new evidence for the role of receptor tyrosine kinases in the establishment and progression of prostate cancer.

Biomarkers, Tumor↗

A new progressive-iterative algorithm for multiple structure alignment.

MOTIVATION: Multiple structure alignments are becoming important tools in many aspects of structural bioinformatics. The current explosion in the number of available protein structures demands multiple structural alignment algorithms with an adequate balance of accuracy and speed, for large scale applications in structural genomics, protein structure prediction and protein classification. RESULTS: A new multiple structural alignment program, MAMMOTH-mult, is described. It is demonstrated that the alignments obtained with the new method are an improvement over previous manual or automatic alignments available in several widely used databases at all structural levels. Detailed analysis of the structural alignments for a few representative cases indicates that MAMMOTH-mult delivers biologically meaningful trees and conservation at the sequence and structural levels of functional motifs in the alignments. An important improvement over previous methods is the reduction in computational cost. Typical alignments take only a median time of 5 CPU seconds in a single R12000 processor. MAMMOTH-mult is particularly useful for large scale applications. AVAILABILITY: http://ub.cbm.uam.es/mammoth/mult.

Algorithms↗

Core deformations in protein families: a physical perspective.

An analysis is presented on how structural cores change shape within protein families, and whether or not there is a relationship between these structural changes and the vibrational modes that proteins experiment due to topological constraints. A set of 13 representative and well-populated protein families are studied. The evolutionary directions of deformation are obtained by applying a new multiple structural alignment technique to superimpose the structures and extract a conserved core, together with Principal Components Analysis (PCA) to extract the main deformation modes. A low-resolution Normal Mode Analysis (NMA) technique is used in parallel to study the properties of the mechanical core plasticity of the same proteins. We find that the evolutionary deformations span a low dimensional space. A statistically significant correspondence exists between these principal deformations and the vibrational modes accessible to a particular topology. We conclude that, to a significant extent, the structures of evolving proteins seem to respond to sequence changes by collective deformations along combinations of low-frequency modes. The findings have implications in structure prediction by homology modeling.

Chemical Phenomena↗

An analysis of core deformations in protein superfamilies.

An analysis is presented on how structural cores modify their shape across homologous proteins, and whether or not a relationship exists between these structural changes and the vibrational normal modes that proteins experience as a result of the topological constraints imposed by the fold. A set of 35 representative, well-populated protein families is studied. The evolutionary directions of deformation are obtained by using multiple structural alignments to superimpose the structures and extract a conserved core, together with principal components analysis to extract the main deformation modes from the three-dimensional superimposition. In parallel, a low-resolution normal mode analysis technique is employed to study the properties of the mechanical core plasticity of these same families. We show that the evolutionary deformations span a low dimensional space of 4-5 dimensions on average. A statistically significant correspondence exists between these principal deformations and the approximately 20 slowest vibrational modes accessible to a particular topology. We conclude that, to a significant extent, the structural response of a protein topology to sequence changes takes place by means of collective deformations along combinations of a small number of low-frequency modes. The findings have implications in structure prediction by homology modeling.

Amino Acid Sequence↗

Improvement of comparative model accuracy by free-energy optimization along principal components of natural structural variation.

Accurate high-resolution refinement of protein structure models is a formidable challenge because of the delicate balance of forces in the native state, the difficulty in sampling the very large number of alternative tightly packed conformations, and the inaccuracies in current force fields. Indeed, energy-based refinement of comparative models generally leads to degradation rather than improvement in model quality, and, hence, most current comparative modeling procedures omit physically based refinement. However, despite their inaccuracies, current force fields do contain information that is orthogonal to the evolutionary information on which comparative models are based, and, hence, refinement might be able to improve comparative models if the space that is sampled is restricted sufficiently so that false attractors are avoided. Here, we use the principal components of the variation of backbone structures within a homologous family to define a small number of evolutionarily favored sampling directions and show that model quality can be improved by energy-based optimization along these directions.

Models, Molecular↗

Virtual screening with flexible docking and COMBINE-based models. Application to a series of factor Xa inhibitors.

A two-step, fully automatic virtual screening procedure consisting of flexible docking followed by activity prediction by COMparative BINding Energy (COMBINE) analysis is presented. This novel approach has been successfully applied, as an example with medicinal chemistry interest, to a recently reported series of 133 factor Xa (fXa)(1) inhibitors whose activities encompass 4 orders of magnitude. The docking algorithm is linked to the COMBINE analysis program and used to derive independent regression models of the 133 inhibitors docked within three different fXa structures (PDB entries 1fjs, 1f0r, and 1xka), so as to explore the effect of receptor conformation on the overall results. Reliable docking conformations and predictive regression models requiring eight latent variables could be derived for two of the fXa structures, with the best model achieving a Q(2) of 0.63 and a standard deviation of errors of prediction (SDEP) of 0.51 (leave-one-out). The two-step procedure was then employed to screen a designed virtual library of 112 ligands, containing both active and inactive compounds. While docking energies alone could show a good performance for selecting hits, including structurally diverse ones, inclusion of COMBINE analysis regression models provided improved rankings for the identification of structurally related molecules in external sets. In our best case, a recognition rate of approximately 80% of known binders at approximately 15% false positives rate was achieved, corresponding to an enrichment factor of approximately 450% over random.

Algorithms↗

Gaussian mapping of chemical fragments in ligand binding sites.

We present a new approach to automatically define a quasi-optimal minimal set of pharmacophoric points mapping the interaction properties of a user-defined ligand binding site. The method is based on a fitting algorithm where a grid of sampled interaction energies of the target protein with small chemical fragments in the binding site is approximated by a linear expansion of Gaussian functions. A heuristic approximation selects from this expansion the smallest possible set of Gaussians required to describe the interaction properties of the binding site within a prespecified accuracy. We have evaluated the performance of the approach by comparing the computed Gaussians with the positions of aromatic sites found in experimental protein-ligand complexes. For a set of 53 complexes, good correspondence is found in general. At a 95% significance level, approximately 65% of the predicted interaction points have an aromatic binding site within 1.5 A. We then studied the utility of these points in docking using the program DOCK. Short docking times, with an average of approximately 0.18 s per conformer, are obtained, while retaining, both for rigid and flexible docking, the ability to sample native-like binding modes for the ligand. An average 4-5-fold speed-up in docking times and a similar success rate is estimated with respect to the standard DOCK protocol.

Algorithms↗

Solution structure of the hypothetical protein Mth677 from Methanobacterium thermoautotrophicum: a novel alpha+beta fold.

The structure of Mth677, a hypothetical protein from Methanobacterium thermoautotrophicum (Mth), has been determined by using heteronuclear nuclear magnetic resonance (NMR) methods on a double-labeled (15)N-(13)C sample. Mth677 adopts a novel alpha+beta fold, consisting of two alpha-helices (one N terminal and one C terminal) packed on the same side of a central beta-hairpin. This structure is likely shared by its three orthologs, detected in three other Archaebacteria. There are no clear features in the sequences of these proteins or in the genome organization of Mth to make a reliable functional assignment to this protein. However, the structural similarity to Escherichia coli MinE, the protein which controls that division occurs at the midcell site, lends support to the proposal that Mth677 might be, in Mth, the counterpart of the topological specificity domain of MinE in E. coli.

Amino Acid Sequence↗

CAFASP3: the third critical assessment of fully automated structure prediction methods.

We present the results of the fully automated CAFASP3 experiment, which was carried out in parallel with CASP5, using the same set of prediction targets. CAFASP participation is restricted to fully automatic structure prediction servers. The servers' performance is evaluated by using previously announced, objective, reproducible and fully automated evaluation methods. More than 60 servers participated in CAFASP3, covering all categories of structure prediction. As in the previous CAFASP2 experiment, it was possible to identify a group of 5-10 top performing independent servers. This group of top performing independent servers produced relatively accurate models for all the 32 "Homology Modeling" targets, and for up to 43% of the 30 "Fold Recognition" targets. One of the most important results of CAFASP3 was the realization of the value of all the independent servers as a group, as evidenced by the superior performance of "meta-predictors" (defined here as predictors that make use of the output of other CAFASP servers). The performance of the best automated meta-predictors was roughly 30% higher than that of the best independent server. More significantly, the performance of the best automated meta-predictors was comparable with that of the best 5-10 human CASP predictors. This result shows that significant progress has been achieved in automatic structure prediction and has important implications to the prospects of automated structure modeling in the context of structural genomics.

Computational Biology↗

Gene discovery in bladder cancer progression using cDNA microarrays.

To identify gene expression changes along progression of bladder cancer, we compared the expression profiles of early-stage and advanced bladder tumors using cDNA microarrays containing 17,842 known genes and expressed sequence tags. The application of bootstrapping techniques to hierarchical clustering segregated early-stage and invasive transitional carcinomas into two main clusters. Multidimensional analysis confirmed these clusters and more importantly, it separated carcinoma in situ from papillary superficial lesions and subgroups within early-stage and invasive tumors displaying different overall survival. Additionally, it recognized early-stage tumors showing gene profiles similar to invasive disease. Different techniques including standard t-test, single-gene logistic regression, and support vector machine algorithms were applied to identify relevant genes involved in bladder cancer progression. Cytokeratin 20, neuropilin-2, p21, and p33ING1 were selected among the top ranked molecular targets differentially expressed and validated by immunohistochemistry using tissue microarrays (n = 173). Their expression patterns were significantly associated with pathological stage, tumor grade, and altered retinoblastoma (RB) expression. Moreover, p33ING1 expression levels were significantly associated with overall survival. Analysis of the annotation of the most significant genes revealed the relevance of critical genes and pathways during bladder cancer progression, including the overexpression of oncogenic genes such as DEK in superficial tumors or immune response genes such as Cd86 antigen in invasive disease. Gene profiling successfully classified bladder tumors based on their progression and clinical outcome. The present study has identified molecular biomarkers of potential clinical significance and critical molecular targets associated with bladder cancer progression.

Aged↗

Comparative analysis of chloroplast genomes: functional annotation, genome-based phylogeny, and deduced evolutionary patterns.

All protein sequences from 19 complete chloroplast genomes (cpDNA) have been studied using a new computational method able to analyze functional correlations among series of protein sequences contained in complete proteomes. First, all open reading frames (ORFs) from the cpDNAs, comprising a total of 2266 protein sequences, were compared against the 3168 proteins from Synechocystis PCC6803 complete genome to find functionally related orthologous proteins. Additionally, all cpDNA genomes were pairwise compared to find orthologous groups not present in cyanobacteria. Annotations in the cluster of othologous proteins database and CyanoBase were used as reference for the functional assignments. Following this protocol, new functional assignments were made for ORFs of unknown function and for ycfs (hypothetical chloroplast frames), which still lack a functional assignment. Using this information, a matrix of functional relationships was derived from profiles of the presence and/or absence of orthologous proteins; the matrix included 1837 proteins in 277 orthologous clusters. A factor analysis study of this matrix, followed by cluster analysis, allowed us to obtain accurate phylogenetic reconstructions and the detection of genes probably involved in speciation as phylogenetic correlates. Finally, by grouping common evolutionary patterns, we show that it is possible to determine functionally linked protein networks. This has allowed us to suggest putative associations for some unknown ORFs.

Bacterial Proteins↗

MAMMOTH (matching molecular models obtained from theory): an automated method for model comparison.

Advances in structural genomics and protein structure prediction require the design of automatic, fast, objective, and well benchmarked methods capable of comparing and assessing the similarity of low-resolution three-dimensional structures, via experimental or theoretical approaches. Here, a new method for sequence-independent structural alignment is presented that allows comparison of an experimental protein structure with an arbitrary low-resolution protein tertiary model. The heuristic algorithm is given and then used to show that it can describe random structural alignments of proteins with different folds with good accuracy by an extreme value distribution. From this observation, a structural similarity score between two proteins or two different conformations of the same protein is derived from the likelihood of obtaining a given structural alignment by chance. The performance of the derived score is then compared with well established, consensus manual-based scores and data sets. We found that the new approach correlates better than other tools with the gold standard provided by a human evaluator. Timings indicate that the algorithm is fast enough for routine use with large databases of protein models. Overall, our results indicate that the new program (MAMMOTH) will be a good tool for protein structure comparisons in structural genomics applications. MAMMOTH is available from our web site at http://physbio.mssm.edu/~ortizg/.

Algorithms↗