PubMed Health⌕ Search

Biomedical subjects

Eldon L Ulrich

Publications and source records attributed to Eldon L Ulrich.

11 recordsLinked to original sources

RECOORD: a recalculated coordinate database of 500+ proteins from the PDB using restraints from the BioMagResBank.

State-of-the-art methods based on CNS and CYANA were used to recalculate the nuclear magnetic resonance (NMR) solution structures of 500+ proteins for which coordinates and NMR restraints are available from the Protein Data Bank. Curated restraints were obtained from the BioMagResBank FRED database. Although the original NMR structures were determined by various methods, they all were recalculated by CNS and CYANA and refined subsequently by restrained molecular dynamics (CNS) in a hydrated environment. We present an extensive analysis of the results, in terms of various quality indicators generated by PROCHECK and WHAT_CHECK. On average, the quality indicators for packing and Ramachandran appearance moved one standard deviation closer to the mean of the reference database. The structural quality of the recalculated structures is discussed in relation to various parameters, including number of restraints per residue, NOE completeness and positional root mean square deviation (RMSD). Correlations between pairs of these quality indicators were generally low; for example, there is a weak correlation between the number of restraints per residue and the Ramachandran appearance according to WHAT_CHECK (r = 0.31). The set of recalculated coordinates constitutes a unified database of protein structures in which potential user- and software-dependent biases have been kept as small as possible. The database can be used by the structural biology community for further development of calculation protocols, validation tools, structure-based statistical approaches and modeling. The RECOORD database of recalculated structures is publicly available from http://www.ebi.ac.uk/msd/recoord.

Databases, Protein↗

The CCPN data model for NMR spectroscopy: development of a software pipeline.

To address data management and data exchange problems in the nuclear magnetic resonance (NMR) community, the Collaborative Computing Project for the NMR community (CCPN) created a "Data Model" that describes all the different types of information needed in an NMR structural study, from molecular structure and NMR parameters to coordinates. This paper describes the development of a set of software applications that use the Data Model and its associated libraries, thus validating the approach. These applications are freely available and provide a pipeline for high-throughput analysis of NMR data. Three programs work directly with the Data Model: CcpNmr Analysis, an entirely new analysis and interactive display program, the CcpNmr FormatConverter, which allows transfer of data from programs commonly used in NMR to and from the Data Model, and the CLOUDS software for automated structure calculation and assignment (Carnegie Mellon University), which was rewritten to interact directly with the Data Model. The ARIA 2.0 software for structure calculation (Institut Pasteur) and the QUEEN program for validation of restraints (University of Nijmegen) were extended to provide conversion of their data to the Data Model. During these developments the Data Model has been thoroughly tested and used, demonstrating that applications can successfully exchange data via the Data Model. The software architecture developed by CCPN is now ready for new developments, such as integration with additional software applications and extensions of the Data Model into other areas of research.

Computer Graphics↗

Comparison of cell-based and cell-free protocols for producing target proteins from the Arabidopsis thaliana genome for structural studies.

We describe a comparative study of protein production from 96 Arabidopsis thaliana open reading frames (ORFs) by cell-based and cell-free protocols. Each target was carried through four pipeline protocols used by the Center for Eukaryotic Structural Genomics (CESG), one for the production of unlabeled protein to be used in crystallization trials and three for the production of 15N-labeled proteins to be analyzed by 1H-15N NMR correlation spectroscopy. Two of the protocols involved Escherichia coli cell-based and two involved wheat germ cell-free technology. The progress of each target through each of the protocols was followed with all failures and successes noted. Failures were of the following types: ORF not cloned, protein not expressed, low protein yield, no cleavage of fusion protein, insoluble protein, protein not purified, NMR sample too dilute. Those targets that reached the goal of analysis by 1H-15N NMR correlation spectroscopy were scored as HSQC+ (protein folded and suitable for NMR structural analysis), HSQC+/- (protein partially disordered or not in a single stable conformational state), HSQC- (protein unfolded, misfolded, or aggregated and thus unsuitable for NMR structural analysis). Targets were also scored as X- for failing to crystallize and X+ for successful crystallization. The results constitute a rich database for understanding differences between targets and protocols. In general, the wheat germ cell-free platform offers the advantage of greater genome coverage for NMR-based structural proteomics whereas the E. coli platform when successful yields more protein, as currently needed for crystallization trials for X-ray structure determination.

Arabidopsis↗

Addressing the intrinsic disorder bottleneck in structural proteomics.

The Center for Eukaryotic Structural Genomics (CESG), as part of the Protein Structure Initiative (PSI), has established a high-throughput structure determination pipeline focused on eukaryotic proteins. NMR spectroscopy is an integral part of this pipeline, both as a method for structure determinations and as a means for screening proteins for stable structure. Because computational approaches have estimated that many eukaryotic proteins are highly disordered, about 1 year into the project, CESG began to use an algorithm (the Predictor of Naturally Disordered Regions, PONDR to avoid proteins that were likely to be disordered. We report a retrospective analysis of the effect of this filtering on the yield of viable structure determination candidates. In addition, we have used our current database of results on 70 protein targets from Arabidopsis thaliana and 1 from Caenorhabditis elegans, which were labeled uniformly with nitrogen-15 and screened for disorder by NMR spectroscopy, to compare the original algorithm with 13 other approaches for predicting disorder from sequence. Our study indicates that the efficiency of structural proteomics of eukaryotes can be improved significantly by removing targets predicted to be disordered by an algorithm chosen to provide optimal performance.

Algorithms↗

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays.

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis↗

BioMagResBank databases DOCR and FRED containing converted and filtered sets of experimental NMR restraints and coordinates from over 500 protein PDB structures.

We present two new databases of NMR-derived distance and dihedral angle restraints: the Database Of Converted Restraints (DOCR) and the Filtered Restraints Database (FRED). These databases currently correspond to 545 proteins with NMR structures deposited in the Protein Databank (PDB). The criteria for inclusion were that these should be unique, monomeric proteins with author-provided experimental NMR data and coordinates available from the PDB capable of being parsed and prepared in a consistent manner. The Wattos program was used to parse the files, and the CcpNmr FormatConverter program was used to prepare them semi-automatically. New modules, including a new implementation of Aqua in the BioMagResBank (BMRB) software Wattos were used to analyze the sets of distance restraints (DRs) for inconsistencies, redundancies, NOE completeness, classification and violations with respect to the original coordinates. Restraints that could not be associated with a known nomenclature were flagged. The coordinates of hydrogen atoms were recalculated from the positions of heavy atoms to allow for a full restraint analysis. The DOCR database contains restraint and coordinate data that is made consistent with each other and with IUPAC conventions. The FRED database is based on the DOCR data but is filtered for use by test calculation protocols and longitudinal analyses and validations. These two databases are available from websites of the BMRB and the Macromolecular Structure Database (MSD) in various formats: NMR-STAR, CCPN XML, and in formats suitable for direct use in the software packages CNS and CYANA.

Databases, Protein↗

BioMagResBank database with sets of experimental NMR constraints corresponding to the structures of over 1400 biomolecules deposited in the Protein Data Bank.

Experimental constraints associated with NMR structures are available from the Protein Data Bank (PDB) in the form of "Magnetic Resonance" (MR) files. These files contain multiple types of data concatenated without boundary markers and are difficult to use for further research. Reported here are the results of a project initiated to annotate, archive, and disseminate these data to the research community from a searchable resource in a uniform format. The MR files from a set of 1410 NMR structures were analyzed and their original constituent data blocks annotated as to data type using a semi-automated protocol. A new software program called Wattos was then used to parse and archive the data in a relational database. From the total number of MR file blocks annotated as constraints, it proved possible to parse 84% (3337/3975). The constraint lists that were parsed correspond to three data types (2511 distance, 788 dihedral angle, and 38 residual dipolar couplings lists) from the three most popular software packages used in NMR structure determination: XPLOR/CNS (2520 lists), DISCOVER (412 lists), and DYANA/DIANA (405 lists). These constraints were then mapped to a developmental version of the BioMagResBank (BMRB) data model. A total of 31 data types originating from 16 programs have been classified, with the NOE distance constraint being the most commonly observed. The results serve as a model for the development of standards for NMR constraint deposition in computer-readable form. The constraints are updated regularly and are available from the BMRB web site (http://www.bmrb.wisc.edu).

Databases, Protein↗

Project management system for structural and functional proteomics: Sesame.

A computing infrastructure (Sesame) has been designed to manage and link individual steps in complex projects. Sesame is being developed to support a large-scale structural proteomics pilot project. When complete, the system is expected to manage all steps from target selection to data-bank deposition and report writing. We report here on the design criteria of the Sesame system and on results demonstrating successful achievement of the basic goals of its architecture. The Sesame software package, which follows the client/server paradigm, consists of a framework, which supports secure interactions among the three tiers of the system (the client, server, and database tiers), and application modules that carry out specific tasks. The framework utilizes industry standards. The client tier is written in Java2 and can be accessed anywhere through the Internet. All the development on the server tier is also carried out in Java2 so as to accommodate a wide variety of computer platforms. The database tier employs a commercial database management system. Each Sesame application module consists of a simple user interface in the client tier, corresponding objects in the server tier, and relevant data stored in the centralized database. For security, access to stored data is controlled by access privileges. The system facilitates both local and remote collaborations. Because users interact with the system using Java Web Start or through a web browser, access is limited only by the availability of an Internet connection. We describe several Sesame modules that have been developed to the point where they are being utilized routinely to support steps involved in structural and functional proteomics. This software is available to parties interested in using it and assisting to guide its further development.

Database Management Systems↗