PubMed Health⌕ Search

Biomedical subjects

Ingvar Eidhammer

Publications and source records attributed to Ingvar Eidhammer.

9 recordsLinked to original sources

MassSorter: a tool for administrating and analyzing data from mass spectrometry experiments on proteins with known amino acid sequences.

BACKGROUND: Proteomics is the study of the proteome, and is critical to the understanding of cellular processes. Two central and related tasks of proteomics are protein identification and protein characterization. Many small laboratories are interested in the characterization of a small number of proteins, e.g., how posttranslational modifications change under different conditions. RESULTS: We have developed a software tool called MassSorter for administrating and analyzing data from peptide mass fingerprinting experiments on proteins with known amino acid sequences. It is meant for small scale mass spectrometry laboratories that are interested in posttranslational modifications of known proteins. Several experiments can be compared simultaneously, and the matched and unmatched peak values are clearly indicated. The hits can be sorted according to m/z values (default) or according to the sequence of the protein. Filters defined by the user can mark autolytic protease peaks and other contaminating peaks (keratins, proteins co-migrating with the protein of interest, etc.). Unmatched peaks can be further analyzed for unexpected modifications by searches against a local version of the UniMod database. They can also be analyzed for unexpected cleavages, a highly useful feature for proteins that undergo maturation by proteolytic cleavage, creating new N- or C-terminals. Additional tools exist for visualization of the results, like sequence coverage, accuracy plots, different types of statistics, 3D models, etc. The program and a tutorial are freely available for academic users at http://www.bioinfo.no/software/massSorter. CONCLUSION: MassSorter has a number of useful features that can promote the analysis and administration of MS-data.

Algorithms↗

Improving the reliability and throughput of mass spectrometry-based proteomics by spectrum quality filtering.

In contemporary peptide-centric or non-gel proteome studies, vast amounts of peptide fragmentation data are generated of which only a small part leads to peptide or protein identification. This motivates the development and use of a filtering algorithm that removes spectra that contribute little to protein identification. Removal of unidentifiable spectra reduced both the amount of computational and human time spent on analyzing spectra as well as the chances of obtaining false identifications. Thorough testing on various proteome datasets from different instruments showed that the best suggested machine-learning classifier is, on average, able to recognize half of the unidentified spectra as bad spectra. Further analyses showed that several unidentified spectra classified as good were derived from peptides carrying unanticipated amino acid modifications or contained sequence tags that allowed peptide identification using homology searches. The implementation of the classifiers is available under the GNU General Public License at http://www.bioinfo.no/software/spectrumquality.

Adult↗

Analysing the outer membrane subproteome of Methylococcus capsulatus (Bath) using proteomics and novel biocomputing tools.

High-resolution two-dimensional gel electrophoresis and mass spectrometry has been used to identify the outer membrane (OM) subproteome of the Gram-negative bacterium Methylococcus capsulatus (Bath). Twenty-eight unique polypeptide sequences were identified from protein samples enriched in OMs. Only six of these polypeptides had previously been identified. The predictions from novel bioinformatic methods predicting beta-barrel outer membrane proteins (OMPs) and OM lipoproteins were compared to proteins identified experimentally. BOMP ( http://www.bioinfo.no/tools/bomp ) predicted 43 beta-barrel OMPs (1.45%) from the 2,959 annotated open reading frames. This was a lower percentage than predicted from other Gram-negative proteomes (1.8-3%). More than half of the predicted BOMPs in M. capsulatus were annotated as (conserved) hypothetical proteins with significant similarity to very few sequences in Swiss-Prot or TrEMBL. The experimental data and the computer predictions indicated that the protein composition of the M. capsulatus OM subproteome was different from that of other Gram-negative bacteria studied in a similar manner. A new program, Lipo, was developed that can analyse entire predicted proteomes and give a list of recognised lipoproteins categorised according to their lipo-box similarity to known Gram-negative lipoproteins ( http://www.bioinfo.no/tools/lipo ). This report is the first using a proteomics and bioinformatics approach to identify the OM subproteome of an obligate methanotroph.

Bacterial Outer Membrane Proteins↗

Phylogenetic reconstruction of ancestral character states for gene expression and mRNA splicing data.

BACKGROUND: As genomes evolve after speciation, gene content, coding sequence, gene expression, and splicing all diverge with time from ancestors with close relatives. A minimum evolution general method for continuous character analysis in a phylogenetic perspective is presented that allows for reconstruction of ancestral character states and for measuring along branch evolution. RESULTS: A software package for reconstruction of continuous character traits, like relative gene expression levels or alternative splice site usage data is presented and is available for download at http://www.rossnes.org/phyrex. This program was applied to a primate gene expression dataset to detect transcription factor binding sites that have undergone substitution, potentially having driven lineage-specific differences in gene expression. CONCLUSION: Systematic analysis of lineage-specific evolution is becoming the cornerstone of comparative genomics. New methods, like phyrex, extend the capabilities of comparative genomics by tracing the evolution of additional biomolecular processes.

Alternative Splicing↗

A new method for identification of protein (sub)families in a set of proteins based on hydropathy distribution in proteins.

Structural similarity among proteins is reflected in the distribution of hydropathicity along the amino acids in the protein sequence. Similarities in the hydropathy distributions are obvious for homologous proteins within a protein family. They also were observed for proteins with related structures, even when sequence similarities were undetectable. Here we present a novel method that employs the hydropathy distribution in proteins for identification of (sub)families in a set of (homologous) proteins. We represent proteins as points in a generalized hydropathy space, represented by vectors of specifically defined features. The features are derived from hydropathy of the individual amino acids. Projection of this space onto principal axes reveals groups of proteins with related hydropathy distributions. The groups identified correspond well to families of structurally and functionally related proteins. We found that this method accurately identifies protein families in a set of proteins, or subfamilies in a set of homologous proteins. Our results show that protein families can be identified by the analysis of hydropathy distribution, without the need for sequence alignment.

Algorithms↗

Genomic insights into methanotrophy: the complete genome sequence of Methylococcus capsulatus (Bath).

Methanotrophs are ubiquitous bacteria that can use the greenhouse gas methane as a sole carbon and energy source for growth, thus playing major roles in global carbon cycles, and in particular, substantially reducing emissions of biologically generated methane to the atmosphere. Despite their importance, and in contrast to organisms that play roles in other major parts of the carbon cycle such as photosynthesis, no genome-level studies have been published on the biology of methanotrophs. We report the first complete genome sequence to our knowledge from an obligate methanotroph, Methylococcus capsulatus (Bath), obtained by the shotgun sequencing approach. Analysis revealed a 3.3-Mb genome highly specialized for a methanotrophic lifestyle, including redundant pathways predicted to be involved in methanotrophy and duplicated genes for essential enzymes such as the methane monooxygenases. We used phylogenomic analysis, gene order information, and comparative analysis with the partially sequenced methylotroph Methylobacterium extorquens to detect genes of unknown function likely to be involved in methanotrophy and methylotrophy. Genome analysis suggests the ability of M. capsulatus to scavenge copper (including a previously unreported nonribosomal peptide synthetase) and to use copper in regulation of methanotrophy, but the exact regulatory mechanisms remain unclear. One of the most surprising outcomes of the project is evidence suggesting the existence of previously unsuspected metabolic flexibility in M. capsulatus, including an ability to grow on sugars, oxidize chemolithotrophic hydrogen and sulfur, and live under reduced oxygen tension, all of which have implications for methanotroph ecology. The availability of the complete genome of M. capsulatus (Bath) deepens our understanding of methanotroph biology and its relationship to global carbon cycles. We have gained evidence for greater metabolic flexibility than was previously known, and for genetic components that may have biotechnological potential.

Bacterial Proteins↗

BOMP: a program to predict integral beta-barrel outer membrane proteins encoded within genomes of Gram-negative bacteria.

This work describes the development of a program that predicts whether or not a polypeptide sequence from a Gram-negative bacterium is an integral beta-barrel outer membrane protein. The program, called the beta-barrel Outer Membrane protein Predictor (BOMP), is based on two separate components to recognize integral beta-barrel proteins. The first component is a C-terminal pattern typical of many integral beta-barrel proteins. The second component calculates an integral beta-barrel score of the sequence based on the extent to which the sequence contains stretches of amino acids typical of transmembrane beta-strands. The precision of the predictions was found to be 80% with a recall of 88% when tested on the proteins with SwissProt annotated subcellular localization in Escherichia coli K 12 (788 sequences) and Salmonella typhimurium (366 sequences). When tested on the predicted proteome of E.coli, BOMP found 103 of a total of 4346 polypeptide sequences to be possible integral beta-barrel proteins. Of these, 36 were found by BLAST to lack similarity (E-value score < 1e-10) to proteins with annotated subcellular localization in SwissProt. BOMP predicted the content of integral beta-barrels per predicted proteome of 10 different bacteria to range from 1.8 to 3%. BOMP is available at http://www.bioinfo.no/tools/bomp.

Bacterial Outer Membrane Proteins↗

A fast top-down method for constructing reliable radiation hybrid frameworks.

MOTIVATION: Radiation Hybrid Mapping (RHM) is a technique used to order a set of markers on a genome and estimating physical distances between them. RHM provides information on marker placement independent from other methods such as sequencing, and can therefore be used for example in genome sequencing to help ordering contigs. A radiation hybrid framework can be constructed by choosing a set of markers so that the chromosome coverage is good and so that the markers can be ordered with high confidence. Automatically constructing RHM frameworks is a computationally challenging problem. RESULTS: We have developed a new method for constructing radiation hybrid frameworks. Given a relatively large set of markers for a chromosome, the algorithm aims to select an ordered subset that makes up a framework, and that contains as many markers as possible. The algorithm has a time complexity that is better than any of the existing methods that we are aware of. Furthermore, we propose a method for comparing if two frameworks are consistent, giving a visual presentation as well as quantitative measures of how well the two frameworks agree. Applying our method on marker sets from 22 human chromosomes and comparing the resulting frameworks with previously published frameworks, we demonstrate that our automatic method efficiently constructs frameworks with good coverage of each chromosome and with high degree of agreement on the marker ordering.

Algorithms↗

Structure motif discovery and mining the PDB.

MOTIVATION: Many of the most interesting functional and evolutionary relationships among proteins are so ancient that they cannot be reliably detected through sequence analysis and are apparent only through a comparison of the tertiary structures. The conserved features can often be described as structural motifs consisting of a few single residues or Secondary Structure (SS) elements. Confidence in such motifs is greatly boosted when they are found in more than a pair of proteins. RESULTS: We describe an algorithm for the automatic discovery of recurring patterns in protein structures. The patterns consist of individual residues having a defined order along the protein's backbone that come close together in the structure and whose spatial conformations are similar. The residues in a pattern need not be close in the protein's sequence. The work described in this paper builds on an earlier reported algorithm for motif discovery. This paper describes a significant improvement of the algorithm which makes it very efficient. The improved efficiency allows us to use it for doing unsupervised learning of patterns occurring in small subsets in a large set of structures, a non-redundant subset of the Protein Data Bank (PDB) database of all known protein structures.

Algorithms↗