PubMed Health⌕ Search

Biomedical subjects

Marshall Bern

Publications and source records attributed to Marshall Bern.

7 recordsLinked to original sources

Data mining for proteins characteristic of clades.

A synapomorphy is a phylogenetic character that provides evidence of shared descent. Ideally a synapomorphy is ubiquitous within the clade of related organisms and nonexistent outside the clade, implying that it arose after divergence from other extant species and before the last common ancestor of the clade. With the recent proliferation of genetic sequence data, molecular synapomorphies have assumed great importance, yet there is no convenient means to search for them over entire genomes. We have developed a new program called Conserv, which can rapidly assemble orthologous sequences and rank them by various metrics, such as degree of conservation or divergence from out-group orthologs. We have used Conserv to conduct a largescale search for molecular synapomorphies for bacterial clades. The search discovered sequences unique to clades, such as Actinobacteria, Firmicutes and gamma-Proteobacteria, and shed light on several open questions, such as whether Symbiobacterium thermophilum belongs with Actinobacteria or Firmicutes. We conclude that Conserv can quickly marshall evidence relevant to evolutionary questions that would be much harder to assemble with other tools.

Amino Acid Sequence↗

Automatic determination of O-glycan structure from fragmentation spectra.

Glycosylation is one of the most important classes of post-translational protein modifications, but the identification of glycans is difficult because of their branched structures and numerous isomers. We describe an algorithm called CartoonistTwo that proposes structures for O-linked glycans by automatically analyzing fragmentation mass spectra. CartoonistTwo improves upon previous glycan identification software primarily in its scoring function, which can more successfully distinguish among a number of similar structures. CartoonistTwo was designed and tested with FTICR mass spectra, and includes automatic recalibration and peak selection especially tuned for such data, yet it can be easily adapted to fragmentation spectra (MS2 or MSn) from other instrument types. On a validated test set of 34 SORI-CID MSn FTICR spectra from Xenopus egg jelly, CartoonistTwo gave the manually determined structural assignment either the first or second highest score over 90% of the time. And for over 50% of these spectra, CartoonistTwo selected a unique highest scoring structure that agreed with the manually determined one.

Algorithms↗

De novo analysis of peptide tandem mass spectra by spectral graph partitioning.

We report on a new de novo peptide sequencing algorithm that uses spectral graph partitioning. In this approach, relationships between m/z peaks are represented by attractive and repulsive springs, and the vibrational modes of the spring system are used to infer information about the peaks (such as "likely b-ion" or "likely y-ion"). We demonstrate the effectiveness of this approach by comparison with other de novo sequencers on test sets of ion-trap and QTOF spectra, including spectra of mixtures of peptides. On all datasets, we outperform the other sequencers. Along with spectral graph theory techniques, the new de novo sequencer EigenMS incorporates another improvement of independent interest: robust statistical methods for recalibration of time-of-flight mass measurements. Robust recalibration greatly outperforms simple least-squares recalibration, achieving about three times the accuracy for one QTOF dataset.

Algorithms↗

Automatic selection of representative proteins for bacterial phylogeny.

BACKGROUND: Although there are now about 200 complete bacterial genomes in GenBank, deep bacterial phylogeny remains a difficult problem, due to confounding horizontal gene transfers and other phylogenetic "noise". Previous methods have relied primarily upon biological intuition or manual curation for choosing genomic sequences unlikely to be horizontally transferred, and have given inconsistent phylogenies with poor bootstrap confidence. RESULTS: We describe an algorithm that automatically picks "representative" protein families from entire genomes for use as phylogenetic characters. A representative protein family is one that, taken alone, gives an organismal distance matrix in good agreement with a distance matrix computed from all sufficiently conserved proteins. We then use maximum-likelihood methods to compute phylogenetic trees from a concatenation of representative sequences. We validate the use of representative proteins on a number of small phylogenetic questions with accepted answers. We then use our methodology to compute a robust and well-resolved phylogenetic tree for a diverse set of sequenced bacteria. The tree agrees closely with a recently published tree computed using manually curated proteins, and supports two proposed high-level clades: one containing Actinobacteria, Deinococcus, and Cyanobacteria ("Terrabacteria"), and another containing Planctomycetes and Chlamydiales. CONCLUSION: Representative proteins provide an effective solution to the problem of selecting phylogenetic characters.

Algorithms↗

Automatic quality assessment of peptide tandem mass spectra.

MOTIVATION: A powerful proteomics methodology couples high-performance liquid chromatography (HPLC) with tandem mass spectrometry and database-search software, such as SEQUEST. Such a set-up, however, produces a large number of spectra, many of which are of too poor quality to be useful. Hence a filter that eliminates poor spectra before the database search can significantly improve throughput and robustness. Moreover, spectra judged to be of high quality, but that cannot be identified by database search, are prime candidates for still more computationally intensive methods, such as de novo sequencing or wider database searches including post-translational modifications. RESULTS: We report on two different approaches to assessing spectral quality prior to identification: binary classification, which predicts whether or not SEQUEST will be able to make an identification, and statistical regression, which predicts a more universal quality metric involving the number of b- and y-ion peaks. The best of our binary classifiers can eliminate over 75% of the unidentifiable spectra while losing only 10% of the identifiable spectra. Statistical regression can pick out spectra of modified peptides that can be identified by a de novo program but not by SEQUEST. In a section of independent interest, we discuss intensity normalization of mass spectra.

Algorithms↗

Model-based particle picking for cryo-electron microscopy.

We describe an algorithm for finding particle images in cryo-EM micrographs. The algorithm starts from a crude 3D map of the target particle, computed from a relatively small number of manually picked images, and then projects the map in many different directions to give synthetic 2D templates. The templates are clustered and averaged and then cross-correlated with the micrographs. A probabilistic model of the imaging process then scores cross-correlation peaks to produce the final picks. We give quantitative results on two quite different target particles: keyhole limpet hemocyanin and p97 AAA ATPase. On these particles our automatic particle picker shows human performance level, as measured by the Fourier shell correlations of 3D reconstructions.

Adenosine Triphosphatases↗

Automatic particle selection: results of a comparative study.

Manual selection of single particles in images acquired using cryo-electron microscopy (cryoEM) will become a significant bottleneck when datasets of a hundred thousand or even a million particles are required for structure determination at near atomic resolution. Algorithm development of fully automated particle selection is thus an important research objective in the cryoEM field. A number of research groups are making promising new advances in this area. Evaluation of algorithms using a standard set of cryoEM images is an essential aspect of this algorithm development. With this goal in mind, a particle selection "bakeoff" was included in the program of the Multidisciplinary Workshop on Automatic Particle Selection for cryoEM. Twelve groups participated by submitting the results of testing their own algorithms on a common dataset. The dataset consisted of 82 defocus pairs of high-magnification micrographs, containing keyhole limpet hemocyanin particles, acquired using cryoEM. The results of the bakeoff are presented in this paper along with a summary of the discussion from the workshop. It was agreed that establishing benchmark particles and using bakeoffs to evaluate algorithms are useful in promoting algorithm development for fully automated particle selection, and that the infrastructure set up to support the bakeoff should be maintained and extended to include larger and more varied datasets, and more criteria for future evaluations.

Algorithms↗