PubMed Health⌕ Search

Biomedical subjects

Michael R Shortreed

Publications and source records attributed to Michael R Shortreed.

13 recordsLinked to original sources

IsoBayes: a Bayesian approach for single-isoform proteomics inference.

MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.

Proteomics↗

Improved Detection of Differentially Abundant Proteins through FDR-Control of Peptide-Identity-Propagation.

The goal of proteomics is to identify and quantify peptides and proteins within a biological sample. Almost all algorithms for the identification of peptides in LC-MS/MS data employ two steps: peptide/spectrum matching and peptide-identity-propagation (PIP), also known as match-between-runs. PIP can routinely account for up to 40% of all results, with that proportion rising as high as 75% in single-cell proteomics. Unlike peptide identities derived through peptide/spectrum matches, for which error estimation has been strictly enforced for decades, peptide identities derived through PIP have not historically been subject to statistical evaluation. As an indispensable component of label-free quantification, PIP needs a statistically rigorous method for estimating its false-discovery rate (FDR). We present a method for FDR control of PIP, called PIP-ECHO, and devise a rigorous protocol for evaluating FDR control of any PIP method. Using three different benchmark data sets, we evaluate PIP-ECHO alongside the PIP procedures implemented by FlashLFQ, IonQuant, and MaxQuant. These analyses show that only PIP-ECHO can accurately control the FDR of PIP at 1% across all data sets. When analyzing a spike-in data set, PIP-ECHO increases both the accuracy and sensitivity of differential expression analysis, yielding substantially more differentially abundant proteins than either MaxQuant or IonQuant.

Proteomics↗

Ionizable isotopic labeling reagent for relative quantification of amine metabolites by mass spectrometry.

A powerful approach to relative quantification by mass spectrometry is to employ labeling reagents that target specific functional groups in molecules of interest. A quantitative comparison of two or more samples may be readily accomplished by using a chemically identical but isotopically distinct labeling reagent for each sample. The samples may then be combined, subjected to purification steps, and mass analyzed. Comparison of the signal intensities obtained from the isotopically labeled variants of the target analyte(s) provides quantitative information on their relative concentrations in the sample. In this report, we describe the synthesis and use of heavy and light isotopic forms of methyl acetimidate for the relative quantification of amine-containing species. The principal advantages of methyl acetimidate as a labeling reagent are that the reaction product is positively charged and hydrophobicity is increased, both of which enhance electrospray ionization efficiency and increase detection sensitivity. The quantitative nature of the analysis was demonstrated in model metabolomics experiments in which heavy and light labeled Arabidopsis extracts were combined in different ratios. Finally, the labeling strategy was employed to determine differences in the amounts of amine-containing metabolites for Arabidopsis seeds germinated under two different conditions.

Amines↗

Quantitative detection of individual cleaved DNA molecules on surfaces using gold nanoparticles and scanning electron microscope imaging.

Single-nucleotide polymorphisms (SNPs) are the most frequent type of human genetic variation. Recent work has shown that it is possible to directly analyze SNPs in unamplified human genomic DNA samples using the surface-invasive cleavage reaction followed by rolling circle amplification (RCA) labeling of the cleavage products. The individual RCA amplicon molecules were counted on the surface using fluorescence microscopy. Two principal limitations of such single-molecule counting are the variability in the amplicon size, which results in a large variation in fluorescence signal intensity from the dye-labeled DNA molecules, and a high level of background fluorescence. It is shown here that an excellent alternative to RCA labeling is tagging with gold nanoparticles followed by imaging with a scanning electron microscope. Gold nanoparticles have a uniform diameter (15 +/- 0.5 nm) and provide excellent contrast against the background of the silicon substrate employed. Individual gold nanoparticles are readily counted using publicly available software. The results demonstrate that the labeling efficiency is improved by as much as approximately 15-fold, and the signal-to-noise ratio is improved by approximately 4-fold. Detection of individual cleaved DNA molecules following surface-invasive cleavage was linear and quantitative over 3 orders of magnitude in amount of target DNA (10(-18)-10(-15) mol).

DNA↗

Specific capture of mammalian cells by cell surface receptor binding to ligand immobilized on gold thin films.

Aldehyde-terminated self-assembled monolayers (SAMs) on gold surfaces were modified with proteins and employed to capture intact living cells through specific ligand-cell surface receptor interactions. In our model system, the basic fibroblast growth factor (bFGF) binding receptor was targeted on baby hamster kidney (BHK-21) cells. Negative control and target proteins were immobilized on a gold surface by coupling protein primary amines to surface aldehyde groups. Cell-binding was monitored by phase contrast microscopy or surface plasmon resonance (SPR) imaging. The specificity of the receptor-ligand interaction was confirmed by the lack of cell binding to the negative control proteins, cytochrome c and insulin, and by the disruption of cell binding by treatment with heparitinase to destroy heparan sulfate which plays an essential role in the binding of bFGF to FGF receptors. This approach can simultaneously probe a large number of receptor-ligand interactions in cell populations and has potential for targeting and isolating cells from mixtures according to the receptors expressed on their surface.

Animals↗

Scoring single-nucleotide polymorphisms at the single-molecule level by counting individual DNA cleavage events on surfaces.

Single-nucleotide polymorphisms (SNPs) are the most frequent type of human genetic variation. Recent work has shown that it is possible to directly analyze SNPs in unamplified human genomic DNA samples using the surface-invasive cleavage reaction followed by rolling circle amplification (RCA) of the cleavage products. The ability of RCA to produce single-stranded DNA tens of thousands of nucleotides in length from a single cleaved DNA molecule on the surface suggested the possibility of detecting individual cleavage events on the surface. The feasibility of this approach to SNP scoring is shown here. Individual cleavage events on the surface are detected using fluorescence microscopy to visualize the single-stranded DNA product of the RCA reaction labeled with the fluorescent dye SYBR Green I. The surface density of fluorescent features observed is dependent upon the concentration of target DNA. Future reductions of the sample volume and optimization of the reaction conditions offer the potential of being able to perform such analyses on as little as a single copy of genomic DNA target.

DNA↗

Thermodynamically based DNA strand design.

We describe a new algorithm for design of strand sets, for use in DNA computations or universal microarrays. Our algorithm can design sets that satisfy any of several thermodynamic and combinatorial constraints, which aim to maximize desired hybridizations between strands and their complements, while minimizing undesired cross-hybridizations. To heuristically search for good strand sets, our algorithm uses a conflict-driven stochastic local search approach, which is known to be effective in solving comparable search problems. The PairFold program of Andronescu et al. [M. Andronescu, Z. C. Zhang and A. Condon (2005) J. Mol. Biol., 345, 987-1001; M. Andronescu, R. Aguirre-Hernandez, A. Condon, and H. Hoos (2003) Nucleic Acids Res., 31, 3416-3422.] is used to calculate the minimum free energy of hybridization between two mismatched strands. We describe new thermodynamic measures of the quality of strand sets. With respect to these measures of quality, our algorithm consistently finds, within reasonable time, sets that are significantly better than previously published sets in the literature.

Algorithms↗

A thermodynamic approach to designing structure-free combinatorial DNA word sets.

An algorithm is presented for the generation of sets of non-interacting DNA sequences, employing existing thermodynamic models for the prediction of duplex stabilities and secondary structures. A DNA 'word' structure is employed in which individual DNA 'words' of a given length (e.g. 12mer and 16mer) may be concatenated into longer sequences (e.g. four tandem words and six tandem words). This approach, where multiple word variants are used at each tandem word position, allows very large sets of non-interacting DNA strands to be assembled from combinations of the individual words. Word sets were generated and their figures of merit are compared to sets as described previously in the literature (e.g. 4, 8, 12, 15 and 16mer). The predicted hybridization behavior was experimentally verified on selected members of the sets using standard UV hyperchromism measurements of duplex melting temperatures (T(m)s). Additional experimental validation was obtained by using the sequences in formulating and solving a small example of a DNA computing problem.

Algorithms↗

Parallel single nucleotide polymorphism genotyping by surface invasive cleavage with universal detection.

Large-scale investigations of sequence variation within the human species will provide information about the basis of heritable variation in disease susceptibility and human migration. The surface invader assay (an adaptation of the invasive cleavage reaction to an array format) is capable of exquisitely sensitive and specific detection of genetic variation. It is shown here that this genotyping technology can be multiplexed in a DNA array format, permitting the parallel analysis of a panel of single nucleotide polymorphisms (SNPs) directly from an unamplified genomic DNA target. In addition, a "universal" mode of detection was developed that makes use of a mixture of degenerate templates for DNA ligation to the surface-bound cleaved oligonucleotides and thereby makes this strategy amenable to any desired SNP site or combination of SNP sites, without regard to their particular DNA sequences. This approach was demonstrated on a proof-of-principle scale using small DNA arrays to genotype 6 SNP markers in the PTPN1 gene and 10 mutations in the cystic fibrosis transmembrane conductance regulator gene. This ability to analyze many different genetic variations in parallel, directly from unamplified human genomic DNA samples, lays the groundwork for the development of high-density arrays able to analyze hundreds of thousands or even millions of SNPs.

Base Sequence↗

Alpha-Ketoisocaproate-induced hypersecretion of insulin by islets from diabetes-susceptible mice.

Most patients at risk for developing type 2 diabetes are hyperinsulinemic. Hyperinsulinemia may be a response to insulin resistance, but another possible abnormality is insulin hypersecretion. BTBR mice are insulin resistant and hyperinsulinemic. When the leptin(ob) mutation is introgressed into BTBR mice, they develop severe diabetes. We compared the responsiveness of lean B6 and BTBR mouse islets to various insulin secretagogues. The transamination product of leucine, alpha-ketoisocaproate (KIC), elicited a dramatic insulin secretory response in BTBR islets. The KIC response was blocked by methyl-leucine or aminooxyacetate, inhibitors of branched-chain amino transferase. When dimethylglutamate was combined with KIC, the fractional insulin secretion was identical in islets from both mouse strains, predicting that the amine donor is rate-limiting for KIC-induced insulin secretion. Consistent with this prediction, glutamate levels were higher in BTBR than in B6 islets. The transamination product of glutamate, alpha-ketoglutarate, elicited insulin secretion equally from B6 and BTBR islets. Thus formation of alpha-ketoglutarate is a requisite step in the response of mouse islets to KIC. alpha-Ketoglutarate can be oxidized to succinate. However, succinate does not stimulate insulin secretion in mouse islets. Our data suggest that alpha-ketoglutarate may directly stimulate insulin secretion and that increased formation of alpha-ketoglutarate leads to hyperinsulinemia.

Animals↗

Surface amplification of invasive cleavage products.

A major focus of current efforts in genomics is to elucidate the genetic variations extent within the human population, and to study the effects of these variations upon the human system. The most common type of genetic variations are the single nucleotide polymorphisms (SNPs), which occur every 500-1000 nt in the genome. Large-scale population association studies to study the biological or medical significance of such variations may require the analysis of hundreds of thousands of SNPs on thousands of individuals. We are pursuing development of an approach to large-scale SNP analysis that combines the specificity of invasive cleavage reactions with the parallelism of high density DNA arrays. A surface-immobilized probe oligonucleotide is specifically cleaved in the presence of a complementary target sequence in unamplified human genomic DNA, yielding a 5' phosphate group. High sensitivity detection of this reaction product on the surface is achieved by the use of rolling circle amplification, with an approximate concentration detection limit of 10 fM target DNA. This combination of very specific surface cleavage and highly sensitive surface detection will make possible the rapid and parallel analysis of genetic variations across large populations.

DNA↗

Structure-specific DNA cleavage on surfaces.

The structure-specific invasive cleavage reaction is a useful means for sensitive and specific detection of single nucleotide polymorphisms, or SNPs, directly from genomic DNA without a need for prior target amplification. A new approach integrating this invasive cleavage assay and surface DNA array technology has been developed for potentially large-scale SNP scoring in a parallel format. Two surface invasive cleavage reaction strategies were designed and implemented for a model SNP system in codon 158 of the human ApoE gene. The upstream oligonucleotide, which is required for the invasive cleavage reaction, is either co-immobilized on the surface along with the probe oligonucleotide or alternatively added in solution. The ability of this approach to unambiguously discriminate a single base difference was demonstrated using PCR-amplified human genomic DNA. A theoretical model relating the surface fluorescence intensity to the progress of the invasive cleavage reaction was developed and agreed well with experimental results.

DNA↗

A surface invasive cleavage assay for highly parallel SNP analysis.

The structure-specific invasive cleavage of single-stranded DNA by 5' nucleases is a useful means for sensitive detection of single-nucleotide polymorphisms or SNPs. The solution-phase invasive cleavage reaction has sufficient sensitivity for direct detection of as few as 600 target molecules with no prior target amplification. One approach to the parallelization of SNP analysis is to adapt the invasive cleavage reaction to an addressed array format. Two surface invasive cleavage reaction strategies were designed and tested using the polymorphic site in codon 158 of the human ApoE gene as a model system, with a synthetic oligonucleotide as target. The upstream oligonucleotide, which is required for the invasive cleavage reaction, was either added in solution (strategy 1), or co-immobilized on the surface along with the probe oligonucleotide (strategy 2). Both strategies showed target-concentration and time-dependent amplification of signal. Parameters that govern the rate of the surface-invasive cleavage reactions are discussed.

Apolipoproteins E↗