PubMed Health⌕ Search

Biomedical subjects

Jan Eriksson

Publications and source records attributed to Jan Eriksson.

12 recordsLinked to original sources

Optimizing search conditions for the mass fingerprint-based identification of proteins.

The two central problems in protein identification by searching a protein sequence collection with MS data are the optimal use of experimental information to allow for identification of low abundance proteins and the accurate assignment of the probability that a result is false. For comprehensive MS-based protein identification, it is necessary to choose an appropriate algorithm and optimal search conditions. We report a systematic study of the quality of PMF-based protein identifications under different sequence collection search conditions using the Probability algorithm, which assigns the statistical significance to each result. We employed 2244 PMFs from 2-DE-separated human blood plasma proteins, and performed identification under various search constraints: mass accuracy (0.01-0.3 Da), maximum number of missed cleavage sites (0-2), and size of the sequence collection searched (5.6 x 10(4)-1.8 x 10(5)). By counting the number of significant results (significance levels 0.05, 0.01, and 0.001) for each condition, we demonstrate the search condition impact on the successful outcome of proteome analysis experiments. A mass correction procedure utilizing mass deviations of albumin matching peptides was tested in an attempt to improve the statistical significance of identifications and iterative searching was employed for identification of multiple proteins from each PMF.

Blood Proteins↗

Bioinformatic and enzymatic characterization of the MAPEG superfamily.

The membrane associated proteins in eicosanoid and glutathione metabolism (MAPEG) superfamily includes structurally related membrane proteins with diverse functions of widespread origin. A total of 136 proteins belonging to the MAPEG superfamily were found in database and genome screenings. The members were found in prokaryotes and eukaryotes, but not in any archaeal organism. Multiple sequence alignments and calculations of evolutionary trees revealed a clear subdivision of the eukaryotic MAPEG members, corresponding to the six families of microsomal glutathione transferases (MGST) 1, 2 and 3, leukotriene C4 synthase (LTC4), 5-lipoxygenase activating protein (FLAP), and prostaglandin E synthase. Prokaryotes contain at least two distinct potential ancestral subfamilies, of which one is unique, whereas the other most closely resembles enzymes that belong to the MGST2/FLAP/LTC4 synthase families. The insect members are most similar to MGST1/prostaglandin E synthase. With the new data available, we observe that fish enzymes are present in all six families, showing an early origin for MAPEG family differentiation. Thus, the evolutionary origins and relationships of the MAPEG superfamily can be defined, including distinct sequence patterns characteristic for each of the subfamilies. We have further investigated and functionally characterized representative gene products from Escherichia coli, Synechocystis sp., Arabidopsis thaliana and Drosophila melanogaster, and the fish liver enzyme, purified from pike (Esox lucius). Protein overexpression and enzyme activity analysis demonstrated that all proteins catalyzed the conjugation of 1-chloro-2,4-dinitrobenzene with reduced glutathione. The E. coli protein displayed glutathione transferase activity of 0.11 micromol.min(-1).mg(-1) in the membrane fraction from bacteria overexpressing the protein. Partial purification of the Synechocystis sp. protein yielded an enzyme of the expected molecular mass and an N-terminal amino acid sequence that was at least 50% pure, with a specific activity towards 1-chloro-2,4-dinitrobenzene of 11 micromol.min(-1).mg(-1). Yeast microsomes expressing the Arabidopsis enzyme showed an activity of 0.02 micromol.min(-1).mg(-1), whereas the Drosophila enzyme expressed in E. coli was highly active at 3.6 micromol.min(-1).mg(-1). The purified pike enzyme is the most active MGST described so far with a specific activity of 285 micromol.min(-1).mg(-1). Drosophila and pike enzymes also displayed glutathione peroxidase activity towards cumene hydroperoxide (0.4 and 2.2 micromol.min(-1).mg(-1), respectively). Glutathione transferase activity can thus be regarded as a common denominator for a majority of MAPEG members throughout the kingdoms of life whereas glutathione peroxidase activity occurs in representatives from the MGST1, 2 and 3 and PGES subfamilies.

Animals↗

Cadmium in food production systems: a health risk for sensitive population groups.

This paper gives an overview of the cadmium (Cd) situation in agricultural systems and human exposure in Sweden. Cadmium levels in agricultural soils (the plow layer) increase by 0.03% to 0.05% per year. Feed can give substantial contributions of Cd to local agricultural systems. Effects on human kidney function are indicated by some measurements already at today's exposure levels. If food products reach the maximum permissible levels given by the European Union, 10% to 25% of the Swedish population will be exposed to Cd levels above the provisional tolerable weekly intake (PTWI 7 microg Cd kg(-1) body weight). Sensitive groups in the population are individuals with low iron status (mainly women) and kidney disorders. Recent studies indicate that Cd plays a role in osteoporosis and that further research is needed to clarify if Cd is neurotoxic in early developmental stages. Firm actions have to be taken in order to stop a further increase of Cd in agricultural soils. Suggestions for prevention and measures are given in this paper.

Agriculture↗

Method for differential detection and identification of components in protein mixtures analyzed by matrix-assisted laser desorption/ionization time-of-flight mass spectrometry.

We demonstrate that the semi-quantitative information in matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectra of tryptically digested protein mixtures can, via a systematic statistical approach, be utilized for the identification of a protein present in different concentrations in two samples. Multiple mass spectra were acquired from a series of tryptically digested test samples in which the concentration of one protein was varied and the concentrations of three other proteins were held constant. The mass spectra were subjected to soft independent modeling of class analogy (SIMCA) analysis assuming that spectra originating from two different samples belonged to different data classes. The SIMCA analysis yielded information on which individual m/z values discriminate between two classes. Protein identification by proteolytic peptide mass fingerprinting was performed with different numbers of mass values in the fingerprint according to the discriminatory information, beginning with the mass corresponding to the best discrimination, followed by the best together with the second best, etc. By using the Probity algorithm, which computes the statistical significance of each identification result, we demonstrate that the first protein identified at a desired significance level (0.001) is the protein that was present in a different concentration in the two samples. Differential analysis of expression is often performed by comparing 2D-gel-spot intensities followed by mass spectrometric identification of the respective protein in each spot that differs. The method presented here has the potential to allow identification of the protein component that differs in cases where a gel-spot is poorly resolved and contains several proteins.

Algorithms↗

Medical activities and statistics.

The evaluation of the medical activity is a major concern for hospitals and public health services. With the introduction of coding (IDC-10 and CHOP classifications) hospitals are now able to analyze their medical activity. A way to improve physicians' acceptance in analyzing their work is to give them valuable feedback information. Building statistics tools is costly and time consuming. Therefore introducing data warehouse tools is helpful. Nice Code is an easy-to-use software that helps medical encoding while immediately offering understandable statistics. In Switzerland, physicians demand real feedback based on the data transmitted at the "cantonal" or federal levels and more transparency in third payer's decisions. In this respect, several cantons decided to equip public health services and hospitals with this tool. The goal is to give, physicians and economists, powerful tools for analyzing the medical activity.

Hospital Information Systems↗

A model of random mass-matching and its use for automated significance testing in mass spectrometric proteome analysis.

A rapid and accurate method for testing the significance of protein identities determined by mass spectrometric analysis of protein digests and genome database searching is presented. The method is based on direct computation using a statistical model of the random matching of measured and theoretical proteolytic peptide masses. Protein identification algorithms typically rank the proteins of a genome database according to a score based on the number of matches between the masses obtained by mass spectrometry analysis and the theoretical proteolytic peptide masses of a database protein. The random matching of experimental and theoretical masses can cause false results. A result is significant only if the score characterizing the result deviates significantly from the score expected from a false result. A distribution of the score (number of matches) for random (false) results is computed directly from our model of the random matching, which allows significance testing under any experimental and database search constraints. In order to mimic protein identification data quality in large-scale proteome projects, low-to-high quality proteolytic peptide mass data were generated in silico and subsequently submitted to a database search program designed to include significance testing based on direct computation. This simulation procedure demonstrates the usefulness of direct significance testing for automatically screening for samples that must be subjected to peptide sequence analysis by e.g. tandem mass spectrometry in order to determine the protein identity.

Algorithms↗

Hardware optimization and serial implementation of a novel spiking neuron model for the POEtic tissue.

In this paper we describe the hardware implementation of a spiking neuron model, which uses a spike time dependent plasticity (STDP) rule that allows synaptic changes by discrete time steps. For this purpose an integrate-and-fire neuron is used with recurrent local connections. The connectivity of this model has been set to 24 neighbours, so there is a high degree of parallelism. After obtaining good results with the hardware implementation of the model, we proceed to simplify this hardware description, trying to keep the same behaviour. Some experiments using dynamic grading patterns have been used in order to test the learning capabilities of the model. Finally, the serial implementation has been realized.

Action Potentials↗

Dynamics of pruning in simulated large-scale spiking neural networks.

Massive synaptic pruning following over-growth is a general feature of mammalian brain maturation. This article studies the synaptic pruning that occurs in large networks of simulated spiking neurons in the absence of specific input patterns of activity. The evolution of connections between neurons were governed by an original bioinspired spike-timing-dependent synaptic plasticity (STDP) modification rule which included a slow decay term. The network reached a steady state with a bimodal distribution of the synaptic weights that were either incremented to the maximum value or decremented to the lowest value. After 1x10(6) time steps the final number of synapses that remained active was below 10% of the number of initially active synapses independently of network size. The synaptic modification rule did not introduce spurious biases in the geometrical distribution of the remaining active projections. The results show that, under certain conditions, the model is capable of generating spontaneously emergent cell assemblies.

Action Potentials↗

Event-related potentials in an auditory oddball situation in the rat.

Evoked potentials were recorded from the auditory cortex of both freely moving and anesthetized rats when deviant sounds were presented in a homogenous series of standard sounds (oddball condition). A component of the evoked response to deviant sounds, the mismatch negativity (MMN), may underlie the ability to discriminate acoustic differences, a fundamental aspect of auditory perception. Whereas most MMN studies in animals have been done using simple sounds, this study involved a more complex set of sounds (synthesized vowels). The freely moving rats had previously undergone behavioral training in which they learned to respond differentially to these sounds. Although we found little evidence in this preparation for the typical, epidurally recorded, MMN response, a significant difference between deviant and standard evoked potentials was noted for the freely moving animals in the 100-200 ms range following stimulus onset. No such difference was found in the anesthetized animals.

Animals↗

Probity: a protein identification algorithm with accurate assignment of the statistical significance of the results.

An algorithm for protein identification based on mass spectrometric proteolytic peptide mapping and genome database searching is presented. The algorithm ranks database proteins based on direct calculation of the probability of random matching and assigns the statistical significance to each result. We investigate the performance of the algorithm by simulation and show that the algorithm responds to random data in the desired manner and that the statistical significance computed indicates the risk that a particular identification result is false.

Algorithms↗

Protein identification in complex mixtures.

This paper investigates the prospects of successful mass spectrometric protein identification based on mass data from proteolytic digests of complex protein mixtures. Sets of proteolytic peptide masses representing various numbers of digested proteins in a mixture were generated in silico. In each set, different proteins were selected from a protein sequence collection and for each protein the sequence coverage was randomly selected within a particular regime (15-30% or 30-60%). We demonstrate that the Probity algorithm, which is characterized by an optimal tolerance for random interference, employed in an iterative procedure can correctly identify >95% of proteins at a desired significance level in mixtures composed of hundreds of yeast proteins under realistic mass spectrometric experimental constraints. By using a model of the distribution of protein abundance, we demonstrate that the very high efficiency of identification of protein mixtures that can be achieved by appropriate choices of informatics procedures is hampered by limitations of the mass spectrometric dynamic range. The results stress the desire to choose carefully experimental protocols for comprehensive proteome analysis, focusing on truly critical issues such as the dynamic range, which potentially limits the possibilities of identifying low abundance proteins.

Fungal Proteins↗

The statistical significance of protein identification results as a function of the number of protein sequences searched.

The potential for obtaining a true mass spectrometric protein identification result depends on the choice of algorithm as well as on experimental factors that influence the information content in the mass spectrometric data. Current methods can never prove definitively that a result is true, but an appropriate choice of algorithm can provide a measure of the statistical risk that a result is false, i.e., the statistical significance. We recently demonstrated an algorithm, Probity, which assigns the statistical significance to each result. For any choice of algorithm, the difficulty of obtaining statistically significant results depends on the number of protein sequences in the sequence collection searched. By simulations of random protein identifications and using the Probity algorithm, we here demonstrate explicitly how the statistical significance depends on the number of sequences searched. We also provide an example on how the practitioner's choice of taxonomic constraints influences the statistical significance.

Algorithms↗