PubMed Health⌕ Search

Biomedical subjects

David States

Publications and source records attributed to David States.

5 recordsLinked to original sources

Michigan Molecular Interactions (MiMI): putting the jigsaw puzzle together.

Protein interaction data exists in a number of repositories. Each repository has its own data format, molecule identifier and supplementary information. Michigan Molecular Interactions (MiMI) assists scientists searching through this overwhelming amount of protein interaction data. MiMI gathers data from well-known protein interaction databases and deep-merges the information. Utilizing an identity function, molecules that may have different identifiers but represent the same real-world object are merged. Thus, MiMI allows the users to retrieve information from many different databases at once, highlighting complementary and contradictory information. To help scientists judge the usefulness of a piece of data, MiMI tracks the provenance of all data. Finally, a simple yet powerful user interface aids users in their queries, and frees them from the onerous task of knowing the data format or learning a query language. MiMI allows scientists to query all data, whether corroborative or contradictory, and specify which sources to utilize. MiMI is part of the National Center for Integrative Biomedical Informatics (NCIBI) and is publicly available at: http://mimi.ncibi.org.

Databases, Protein↗

A dominant function of IKK/NF-kappaB signaling in global lipopolysaccharide-induced gene expression.

Porphyromonas gingivalis is an etiologic pathogen of periodontitis that is one of the most common inflammatory diseases. Recently, we found that P. gingivalis LPS activated the transcription factor nuclear factor-kappaB (NF-kappaB) through the IkappaB kinase complex (IKK). NF-kappaB is a transcription factor that controls inflammation and host responses. In this study, we examined the role of IKK/NF-kappaBin P. gingivalis LPS-induced gene expression on a genome-wide basis using a combination of microarray and biochemical approaches. A total of 88 early response genes were found to be induced by P. gingivalis LPS in a human THP.1 monocytic cell lines. Interestingly, the induction of most of these genes was abolished or attenuated under the inactivation of IKK/NF-kappaB. Among those IKK/NF-kappaB-dependent genes, 20 genes were NF-kappaB-inducible genes reported previously, and 59 genes represented putative novel NF-kappaB target genes. Using transcription factor binding analysis, we found that most of these putative NF-kappaB target genes contained one or multiple NF-kappaB-binding sites. Also, some transcription factor-binding motifs were overrepresented in the promoter of both known and putative NF-kappaB-dependent genes, indicating that these genes may be regulated in a similar fashion. Furthermore, we found that several transcription factors associated with metabolic and inflammatory responses, including nuclear receptors, activator of protein-1, and early growth responses, were induced by P. gingivalis LPS through IKK/NF-kappaB, indicating that IKK/NF-kappaB may utilize these transcription factors to mediate secondary responses. Taken together, our results demonstrate that IKK/NF-kappaB signaling plays a dominant role in P. gingivalis LPS-induced early response gene expression, suggesting that IKK/NF-kappaB is a therapeutic target for periodontitis.

Animals↗

Computational Proteomics Analysis System (CPAS): an extensible, open-source analytic system for evaluating and publishing proteomic data and high throughput biological experiments.

The open-source Computational Proteomics Analysis System (CPAS) contains an entire data analysis and management pipeline for Liquid Chromatography Tandem Mass Spectrometry (LC-MS/MS) proteomics, including experiment annotation, protein database searching and sequence management, and mining LC-MS/MS peptide and protein identifications. CPAS architecture and features, such as a general experiment annotation component, installation software, and data security management, make it useful for collaborative projects across geographical locations and for proteomics laboratories without substantial computational support.

Computational Biology↗

PRIDE: the proteomics identifications database.

The advent of high-throughput proteomics has enabled the identification of ever increasing numbers of proteins. Correspondingly, the number of publications centered on these protein identifications has increased dramatically. With the first results of the HUPO Plasma Proteome Project being analyzed and many other large-scale proteomics projects about to disseminate their data, this trend is not likely to flatten out any time soon. However, the publication mechanism of these identified proteins has lagged behind in technical terms. Often very long lists of identifications are either published directly with the article, resulting in both a voluminous and rather tedious read, or are included on the publisher's website as supplementary information. In either case, these lists are typically only provided as portable document format documents with a custom-made layout, making it practically impossible for computer programs to interpret them, let alone efficiently query them. Here we propose the proteomics identifications (PRIDE) database (http://www.ebi.ac.uk/pride) as a means to finally turn publicly available data into publicly accessible data. PRIDE offers a web-based query interface, a user-friendly data upload facility, and a documented application programming interface for direct computational access. The complete PRIDE database, source code, data, and support tools are freely available for web access or download and local installation.

Computational Biology↗

Selecting for functional alternative splices in ESTs.

The expressed sequence tag (EST) collection in dbEST provides an extensive resource for detecting alternative splicing on a genomic scale. Using genomically aligned ESTs, a computational tool (TAP) was used to identify alternative splice patterns for 6400 known human genes from the RefSeq database. With sufficient EST coverage, one or more alternatively spliced forms could be detected for nearly all genes examined. To identify high (>95%) confidence observations of alternative splicing, splice variants were clustered on the basis of having mutually exclusive structures, and sample statistics were then applied. Through this selection, alternative splices expected at a frequency of >5% within their respective clusters were seen for only 17%-28% of genes. Although intron retention events (potentially unspliced messages) had been seen for 36% of the genes overall, the same statistical selection yielded reliable cases of intron retention for <5% of genes. For high-confidence alternative splices in the human ESTs, we also noted significantly higher rates both of cross-species conservation in mouse ESTs and of validation in the GenBank mRNA collection. We suggest quantitative analytical approaches such as these can aid in selecting useful targets for further experimental characterization and in so doing may help elucidate the mechanisms and biological implications of alternative splicing.

Alternative Splicing↗