PubMed Health⌕ Search

Biomedical subjects

Lars Malmström

Publications and source records attributed to Lars Malmström.

11 recordsLinked to original sources

2DDB - a bioinformatics solution for analysis of quantitative proteomics data.

BACKGROUND: We present 2DDB, a bioinformatics solution for storage, integration and analysis of quantitative proteomics data. As the data complexity and the rate with which it is produced increases in the proteomics field, the need for flexible analysis software increases. RESULTS: 2DDB is based on a core data model describing fundamentals such as experiment description and identified proteins. The extended data models are built on top of the core data model to capture more specific aspects of the data. A number of public databases and bioinformatical tools have been integrated giving the user access to large amounts of relevant data. A statistical and graphical package, R, is used for statistical and graphical analysis. The current implementation handles quantitative data from 2D gel electrophoresis and multidimensional liquid chromatography/mass spectrometry experiments. CONCLUSION: The software has successfully been employed in a number of projects ranging from quantitative liquid-chromatography-mass spectrometry based analysis of transforming growth factor-beta stimulated fi-broblasts to 2D gel electrophoresis/mass spectrometry analysis of biopsies from human cervix. The software is available for download at SourceForge.

Computational Biology↗

The Yeast Resource Center Public Data Repository.

The Yeast Resource Center Public Data Repository (YRC PDR) serves as a single point of access for the experimental data produced from many collaborations typically studying Saccharomyces cerevisiae (baker's yeast). The experimental data include large amounts of mass spectrometry results from protein co-purification experiments, yeast two-hybrid interaction experiments, fluorescence microscopy images and protein structure predictions. All of the data are accessible via searching by gene or protein name, and are available on the Web at http://www.yeastrc.org/pdr/.

Databases, Protein↗

Free modeling with Rosetta in CASP6.

We describe Rosetta predictions in the Sixth Community-Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction (CASP), focusing on the free modeling category. Methods developed since CASP5 are described, and their application to selected targets is discussed. Highlights include improved performance on larger proteins (100-200 residues) and the prediction of a 70-residue alpha-beta protein to near-atomic resolution.

Algorithms↗

Prediction of CASP6 structures using automated Robetta protocols.

The Robetta server and revised automatic protocols were used to predict structures for CASP6 targets. Robetta is a publicly available protein structure prediction server (http://robetta.bakerlab.org/ that uses the Rosetta de novo and homology modeling structure prediction methods. We incorporated some of the lessons learned in the CASP5 experiment into the server prior to participating in CASP6. We additionally tested new ideas that were amenable to full-automation with an eye toward improving the server. We find that the Robetta server shows the greatest promise for the more challenging targets. The most significant finding from CASP5, that automated protocols can be roughly comparable in ability with the better human-intervention predictors, is repeated here in CASP6.

Algorithms↗

Automated prediction of domain boundaries in CASP6 targets using Ginzu and RosettaDOM.

Domain boundary prediction is an important step in both experimental and computational protein structure characterization. We have developed two fully automated domain parsing methods: the first, Ginzu, which we have described previously, utilizes information from homologous sequences and structures, while the second, RosettaDOM, which has not been described previously, uses only information in the query sequence. Ginzu iteratively assigns domains by homology to structures and sequence families using successively less confident methods. RosettaDOM uses the Rosetta de novo structure prediction method to build three-dimensional models, and then applies Taylor's structure based domain assignment method to parse the models into domains. Domain boundaries observed repeatedly in the models are predicted to be domain boundaries for the protein. Interestingly, RosettaDOM produced quite good domain predictions for proteins of a size typically considered to be beyond the reach of de novo structure prediction methods. For remote fold recognition targets and new folds, both Ginzu and RosettaDOM produced promising results, and in some cases where one method failed to detect the correct domain boundary, it was correctly identified by the other method. We describe here the successes and failures using both methods, and address the possibility of incorporating both protocols into an improved hybrid method.

Algorithms↗

Nanocapillary liquid chromatography interfaced to tandem matrix-assisted laser desorption/ionization and electrospray ionization-mass spectrometry: mapping the nuclear proteome of human fibroblasts.

Miniaturized liquid chromatography nanoseparation in combination with minigel fractionation of human primary cell nuclei is presented. We obtained high-sensitivity and high-throughput identification of expressed proteins by subcellular fractionation and nanocapillary liquid chromatography interfaced to both electrospray ionization (ESI)- and matrix-assisted laser desorption/ionisation (MALDI) tandem mass spectrometry. The reversed-phase nanocapillary eluents were applied directly onto the MALDI target plate as discrete crystal spots using in-line matrix infusion. When working with primary cells, only a limited amount of sample is available. To maximize the number of identified proteins from a restricted amount of sample, miniaturized sample preparation protocols and nanoflow separation is a necessity, especially when working with low-abundant proteins. From the same isolated nuclear sample, complementary separation of intact proteins by two-dimensional (2-D) gel electrophoresis was made. In total 594 gene products from the nuclear preparations were identified out of which 261 were unique. Several proteins involved in transcriptional events were identified such as TATA-binding protein, EBNA-co-activator, and interleukin enhancer binding proteins, indicating that sufficient proteomic depth is obtained to study transcriptional controlling events. Our results suggest that by sample prefractionation and downscaled nanoflow separation along with a combined mass spectrometry strategy, it is possible to identify a large number of nuclear proteins from human primary cells. These findings are of particular importance due to the disease link of these targets cells.

Cell Nucleus↗

Automated prediction of CASP-5 structures using the Robetta server.

Robetta is a fully automated protein structure prediction server that uses the Rosetta fragment-insertion method. It combines template-based and de novo structure prediction methods in an attempt to produce high quality models that cover every residue of a submitted sequence. The first step in the procedure is the automatic detection of the locations of domains and selection of the appropriate modeling protocol for each domain. For domains matched to a homolog with an experimentally characterized structure by PSI-BLAST or Pcons2, Robetta uses a new alignment method, called K*Sync, to align the query sequence onto the parent structure. It then models the variable regions by allowing them to explore conformational space with fragments in fashion similar to the de novo protocol, but in the context of the template. When no structural homolog is available, domains are modeled with the Rosetta de novo protocol, which allows the full length of the domain to explore conformational space via fragment-insertion, producing a large decoy ensemble from which the final models are selected. The Robetta server produced quite reasonable predictions for targets in the recent CASP-5 and CAFASP-3 experiments, some of which were at the level of the best human predictions.

Algorithms↗

Assigning function to yeast proteins by integration of technologies.

Interpreting genome sequences requires the functional analysis of thousands of predicted proteins, many of which are uncharacterized and without obvious homologs. To assess whether the roles of large sets of uncharacterized genes can be assigned by targeted application of a suite of technologies, we used four complementary protein-based methods to analyze a set of 100 uncharacterized but essential open reading frames (ORFs) of the yeast Saccharomyces cerevisiae. These proteins were subjected to affinity purification and mass spectrometry analysis to identify copurifying proteins, two-hybrid analysis to identify interacting proteins, fluorescence microscopy to localize the proteins, and structure prediction methodology to predict structural domains or identify remote homologies. Integration of the data assigned function to 48 ORFs using at least two of the Gene Ontology (GO) categories of biological process, molecular function, and cellular component; 77 ORFs were annotated by at least one method. This combination of technologies, coupled with annotation using GO, is a powerful approach to classifying genes.

Computational Biology↗

De novo prediction of three-dimensional structures for major protein families.

We use the Rosetta de novo structure prediction method to produce three-dimensional structure models for all Pfam-A sequence families with average length under 150 residues and no link to any protein of known structure. To estimate the reliability of the predictions, the method was calibrated on 131 proteins of known structure. For approximately 60% of the proteins one of the top five models was correctly predicted for 50 or more residues, and for approximately 35%, the correct SCOP superfamily was identified in a structure-based search of the Protein Data Bank using one of the models. This performance is consistent with results from the fourth critical assessment of structure prediction (CASP4). Correct and incorrect predictions could be partially distinguished using a confidence function based on a combination of simulation convergence, protein length and the similarity of a given structure prediction to known protein structures. While the limited accuracy and reliability of the method precludes definitive conclusions, the Pfam models provide the only tertiary structure information available for the 12% of publicly available sequences represented by these large protein families.

Calibration↗

Proteomic 2DE database for spot selection, automated annotation, and data analysis.

We present a software solution that enables faster and more accurate data analysis of 2DE/MALDI TOF MS data. The software supports data analysis through a number of automated data selection functions and advanced graphical tools. Once protein identities are determined using MALDI TOF MS, automated data retrieval from online databases provides biological information. The software, called 2DDB, reduces analysis time to a fraction without losing any quality compared to more manual data analysis. The database contains over 100,000 data entries, and selected parts can be reached at http://2ddb.org.

Animals↗

Proteome annotations and identifications of the human pulmonary fibroblast.

We hereby report on a three year project initiative undertaken by our research team encompassing large-scale protein expression profiling and annotations of human primary lung fibroblast cells. An overview is given of proteomic studies of the fibroblast target cell involved in several diseases such as asthma, idiopatic pulmonary disease, and COPD. It has been the objective within our research team to map and identify the protein expressions occurring in both activated-, as well as resting cell states. The JGGL database www.2DDB.org has been built around these data, allowing advanced hypothesis building using the interactive query bioinformatic tools developed. Gene ontology has been applied to these annotations, classifying and correlating protein expressions to function. The localization as well as the biological processes involved for the annotations are being presented including an annotation-, and sequence-identification strategy, resulting in close to 2000 protein identities. Both gel based, high resolution 2D-gels, and liquid-phase separation (three-dimensional HPLC), as well as the combination of gel- and LC-based approaches (1D-gels and nano-capillary LC, reversed-phase) were utilized. Protein sequencing and structure identities were acquired by a combination of MALDI-, and electrospray-mass spectrometry techniques. Phenotypical and morphological characterizations were also made for this human disease target cell in both stimulated- and resting-cell states. The use of functional assays that demonstrate the key regulating role of growth factors and cytokine stimuli such as PDGF, TGF-beta, and EGF and the effect of ECM molecules such as Biglycan, are also presented and discussed.

Amino Acid Sequence↗