PubMed Health⌕ Search

Biomedical subjects

Micah Hamady

Publications and source records attributed to Micah Hamady.

10 recordsLinked to original sources

Quantitative and qualitative beta diversity measures lead to different insights into factors that structure microbial communities.

The assessment of microbial diversity and distribution is a major concern in environmental microbiology. There are two general approaches for measuring community diversity: quantitative measures, which use the abundance of each taxon, and qualitative measures, which use only the presence/absence of data. Quantitative measures are ideally suited to revealing community differences that are due to changes in relative taxon abundance (e.g., when a particular set of taxa flourish because a limiting nutrient source becomes abundant). Qualitative measures are most informative when communities differ primarily by what can live in them (e.g., at high temperatures), in part because abundance information can obscure significant patterns of variation in which taxa are present. We illustrate these principles using two 16S rRNA-based surveys of microbial populations and two phylogenetic measures of community beta diversity: unweighted UniFrac, a qualitative measure, and weighted UniFrac, a new quantitative measure, which we have added to the UniFrac website (http://bmf.colorado.edu/unifrac). These studies considered the relative influences of mineral chemistry, temperature, and geography on microbial community composition in acidic thermal springs in Yellowstone National Park and the influences of obesity and kinship on microbial community composition in the mouse gut. We show that applying qualitative and quantitative measures to the same data set can lead to dramatically different conclusions about the main factors that structure microbial diversity and can provide insight into the nature of community differences. We also demonstrate that both weighted and unweighted UniFrac measurements are robust to the methods used to build the underlying phylogeny.

Animals↗

Using the nucleotide substitution rate matrix to detect horizontal gene transfer.

BACKGROUND: Horizontal gene transfer (HGT) has allowed bacteria to evolve many new capabilities. Because transferred genes perform many medically important functions, such as conferring antibiotic resistance, improved detection of horizontally transferred genes from sequence data would be an important advance. Existing sequence-based methods for detecting HGT focus on changes in nucleotide composition or on differences between gene and genome phylogenies; these methods have high error rates. RESULTS: First, we introduce a new class of methods for detecting HGT based on the changes in nucleotide substitution rates that occur when a gene is transferred to a new organism. Our new methods discriminate simulated HGT events with an error rate up to 10 times lower than does GC content. Use of models that are not time-reversible is crucial for detecting HGT. Second, we show that using combinations of multiple predictors of HGT offers substantial improvements over using any single predictor, yielding as much as a factor of 18 improvement in performance (a maximum reduction in error rate from 38% to about 3%). Multiple predictors were combined by using the random forests machine learning algorithm to identify optimal classifiers that separate HGT from non-HGT trees. CONCLUSION: The new class of HGT-detection methods introduced here combines advantages of phylogenetic and compositional HGT-detection techniques. These new techniques offer order-of-magnitude improvements over compositional methods because they are better able to discriminate HGT from non-HGT trees under a wide range of simulated conditions. We also found that combining multiple measures of HGT is essential for detecting a wide range of HGT events. These novel indicators of horizontal transfer will be widely useful in detecting HGT events linked to the evolution of important bacterial traits, such as antibiotic resistance and pathogenicity.

Computational Biology↗

Unraveling transcriptional control and cis-regulatory codes using the software suite GeneACT.

Deciphering gene regulatory networks requires the systematic identification of functional cis-acting regulatory elements. We present a suite of web-based bioinformatics tools, called GeneACT http://promoter.colorado.edu, that can rapidly detect evolutionarily conserved transcription factor binding sites or microRNA target sites that are either unique or over-represented in differentially expressed genes from DNA microarray data. GeneACT provides graphic visualization and extraction of common regulatory sequence elements in the promoters and 3'-untranslated regions that are conserved across multiple mammalian species.

Animals↗

UniFrac--an online tool for comparing microbial community diversity in a phylogenetic context.

BACKGROUND: Moving beyond pairwise significance tests to compare many microbial communities simultaneously is critical for understanding large-scale trends in microbial ecology and community assembly. Techniques that allow microbial communities to be compared in a phylogenetic context are rapidly gaining acceptance, but the widespread application of these techniques has been hindered by the difficulty of performing the analyses. RESULTS: We introduce UniFrac, a web application available at http://bmf.colorado.edu/unifrac, that allows several phylogenetic tests for differences among communities to be easily applied and interpreted. We demonstrate the use of UniFrac to cluster multiple environments, and to test which environments are significantly different. We show that analysis of previously published sequences from the Columbia river, its estuary, and the adjacent coastal ocean using the UniFrac interface provided insights that were not apparent from the initial data analysis, which used other commonly employed techniques to compare the communities. CONCLUSION: UniFrac provides easy access to powerful multivariate techniques for comparing microbial communities in a phylogenetic context. We thus expect that it will provide a completely new picture of many microbial interactions and processes in both environmental and medical contexts.

Bacteria↗

DivergentSet, a tool for picking non-redundant sequences from large sequence collections.

DivergentSet addresses the important but so far neglected bioinformatics task of choosing a representative set of sequences from a larger collection. We found that using a phylogenetic tree to guide the construction of divergent sets of sequences can be up to 2 orders of magnitude faster than the naive method of using a full distance matrix. By providing a user-friendly interface (available online) that integrates the tasks of finding additional sequences, building and refining the divergent set, producing random divergent sets from the same sequences, and exporting identifiers, this software facilitates a wide range of bioinformatics analyses including finding significant motifs and covariations. As an example application of DivergentSet, we demonstrate that the motifs identified by the motif-finding package MEME (Motif Elicitation by Maximum Entropy) are highly unstable with respect to the specific choice of sequences. This instability suggests that the types of sensitivity analysis enabled by DivergentSet may be widely useful for identifying the motifs of biological significance.

Amino Acid Sequence↗

Fast-Find: a novel computational approach to analyzing combinatorial motifs.

BACKGROUND: Many vital biological processes, including transcription and splicing, require a combination of short, degenerate sequence patterns, or motifs, adjacent to defined sequence features. Although these motifs occur frequently by chance, they only have biological meaning within a specific context. Identifying transcripts that contain meaningful combinations of patterns is thus an important problem, which existing tools address poorly. RESULTS: Here we present a new approach, Fast-FIND (Fast-Fully Indexed Nucleotide Database), that uses a relational database to support rapid indexed searches for arbitrary combinations of patterns defined either by sequence or composition. Fast-FIND is easy to implement, takes less than a second to search the entire Drosophila genome sequence for arbitrary patterns adjacent to sites of alternative polyadenylation, and is sufficiently fast to allow sensitivity analysis on the patterns. We have applied this approach to identify transcripts that contain combinations of sequence motifs for RNA-binding proteins that may regulate alternative polyadenylation. CONCLUSION: Fast-FIND provides an efficient way to identify transcripts that are potentially regulated via alternative polyadenylation. We have used it to generate hypotheses about interactions between specific polyadenylation factors, which we will test experimentally.

Algorithms↗

Analysis of membrane proteins from human chronic myelogenous leukemia cells: comparison of extraction methods for multidimensional LC-MS/MS.

An important strategy for "shotgun proteomics" profiling involves solution proteolysis of proteins, followed by peptide separation using multidimensional liquid chromatography and automated sequencing by mass spectrometry (LC-MS/MS). Several protocols for extracting and handling membrane proteins for shotgun proteomics experiments have been reported, but few direct comparisons of different protocols have been reported. We compare four methods for preparing membrane proteins from human cells, using acid labile surfactants (ALS), urea, and mixed organic-aqueous solvents. These methods were compared with respect to their efficiency of protein solubilization and proteolysis, peptide and protein recovery, membrane protein enrichment, and peptide coverage of transmembrane proteins. Overall, approximately 50-60% of proteins recovered were membrane-associated, identified from Gene Ontology annotations and transmembrane prediction software. Samples extracted with ALS, extracted with urea followed by dilution, or extracted with urea followed by desalting yielded comparable peptide recoveries and sequence coverage of transmembrane proteins. In contrast, suboptimal proteolysis was observed with organic solvent. Urea extraction followed by desalting may be a particularly useful approach, as it is less costly than ALS and yields satisfactory protein denaturation and proteolysis under conditions that minimize reactivity with urea-derived cyanate. Spectral counting was used to compare datasets of proteins from membrane samples with those of soluble proteins from K562 cells, and to estimate fold differences in protein abundances. Proteins most highly abundant in the membrane samples showed enrichment of integral membrane protein identifications, consistent with their isolation by differential centrifugation.

Cell Extracts↗

A method of mapping protein sumoylation sites by mass spectrometry using a modified small ubiquitin-like modifier 1 (SUMO-1) and a computational program.

Post-translational modification by small ubiquitin-like modifier 1 (SUMO-1) is a highly conserved process from yeast to humans and plays important regulatory roles in many cellular processes. Sumoylation occurs at certain internal lysine residues of target proteins via an isopeptide bond linkage. Unlike ubiquitin whose carboxyl-terminal sequence is RGG, the tripeptide at the carboxyl terminus of SUMO is TGG. The presence of the arginine residue at the carboxyl terminus of ubiquitin allows tryptic digestion of ubiquitin conjugates to yield a signature peptide containing a diglycine remnant attached to the target lysine residue and rapid identification of the ubiquitination site by mass spectrometry. The absence of lysine or arginine residues in the carboxyl terminus of mammalian SUMO makes it difficult to apply this approach to mapping sumoylation sites. We performed Arg scanning mutagenesis by systematically substituting amino acid residues surrounding the diglycine motif and found that a SUMO variant terminated with RGG can be conjugated efficiently to its target protein under normal sumoylation conditions. We developed a Programmed Data Acquisition (PDA) mass spectrometric approach to map target sumoylation sites using this SUMO variant. A web-based computational program designed for efficient identification of the modified peptides is described.

Amino Acid Sequence↗