PubMed Health⌕ Search

Biomedical subjects

Marc Weeber

Publications and source records attributed to Marc Weeber.

8 recordsLinked to original sources

Thesaurus-based disambiguation of gene symbols.

BACKGROUND: Massive text mining of the biological literature holds great promise of relating disparate information and discovering new knowledge. However, disambiguation of gene symbols is a major bottleneck. RESULTS: We developed a simple thesaurus-based disambiguation algorithm that can operate with very little training data. The thesaurus comprises the information from five human genetic databases and MeSH. The extent of the homonym problem for human gene symbols is shown to be substantial (33% of the genes in our combined thesaurus had one or more ambiguous symbols), not only because one symbol can refer to multiple genes, but also because a gene symbol can have many non-gene meanings. A test set of 52,529 Medline abstracts, containing 690 ambiguous human gene symbols taken from OMIM, was automatically generated. Overall accuracy of the disambiguation algorithm was up to 92.7% on the test set. CONCLUSION: The ambiguity of human gene symbols is substantial, not only because one symbol may denote multiple genes but particularly because many symbols have other, non-gene meanings. The proposed disambiguation approach resolves most ambiguities in our test set with high accuracy, including the important gene/not a gene decisions. The algorithm is fast and scalable, enabling gene-symbol disambiguation in massive text mining applications.

Algorithms↗

Online tools to support literature-based discovery in the life sciences.

In biomedical research, the amount of experimental data and published scientific information is overwhelming and ever increasing, which may inhibit rather than stimulate scientific progress. Not only are text-mining and information extraction tools needed to render the biomedical literature accessible but the results of these tools can also assist researchers in the formulation and evaluation of novel hypotheses. This requires an additional set of technological approaches that are defined here as literature-based discovery (LBD) tools. Recently, several LBD tools have been developed for this purpose and a few well-motivated, specific and directly testable hypotheses have been published, some of which have even been validated experimentally. This paper presents an overview of recent LBD research and discusses methodology, results and online tools that are available to the scientific community.

Abstracting and Indexing↗

Chemical and biological profiling of an annotated compound library directed to the nuclear receptor family.

Nuclear receptors form a family of ligand-activated transcription factors that regulate a wide variety of biological processes and are thus generally considered relevant targets in drug discovery. We have constructed an annotated compound library directed to nuclear receptors (NRacl) as a means for integrating the chemical and biological data being generated within this family. Special care has been put in the appropriate storage of annotations by using hierarchical classification schemes for both molecules and nuclear receptors, which takes the ability to extract knowledge from annotated compound libraries to another level. Analysis of NRacl has ultimately led to the identification of scaffolds with highly promiscuous nuclear receptor profiles and to the classification of nuclear receptor groups with similar scaffold promiscuity patterns. This information can be exploited in the design of probing libraries for deorphanization activities as well as for devising screening batteries to address selectivity issues.

Combinatorial Chemistry Techniques↗

Contextual annotation of web pages for interactive browsing.

With the information on the World Wide Web and in specialized databases exploding, researchers and physicians are in dire need to browse efficiently though the large corpus of information resources in their field of interest. The focus is not any longer to find everything related to your interest, but it shifts to zooming in, based on context and expanding again in neighboring knowledge domains. This paper describes an attempt to develop a completely new, interactive way of browsing distributed corpora of information without the need for multiple different queries in different information resources. Classical search engines generally treat search requests in isolation. The results for a given query are identical, and do not automatically take on board the context in which the user made the request. The system described here explores implicit contexts as obtained from the document that the user is reading. The new approach merges the searching and browsing into one combined "read-and-search" mode and alleviates the shift users are normally forced to between searching and reading.

Hypermedia↗

Generating hypotheses by discovering implicit associations in the literature: a case report of a search for new potential therapeutic uses for thalidomide.

The availability of scientific bibliographies through online databases provides a rich source of information for scientists to support their research. However, the risk of this pervasive availability is that an individual researcher may fail to find relevant information that is outside the direct scope of interest. Following Swanson's ABC model of disjoint but complementary structures in the biomedical literature, we have developed a discovery support tool to systematically analyze the scientific literature in order to generate novel and plausible hypotheses. In this case report, we employ the system to find potentially new target diseases for the drug thalidomide. We find solid bibliographic evidence suggesting that thalidomide might be useful for treating acute pancreatitis, chronic hepatitis C, Helicobacter pylori-induced gastritis, and myasthenia gravis. However, experimental and clinical evaluation is needed to validate these hypotheses and to assess the trade-off between therapeutic benefits and toxicities.

Databases, Bibliographic↗

Ambiguity of human gene symbols in LocusLink and MEDLINE: creating an inventory and a disambiguation test collection.

Genes are discovered almost on a daily basis and new names have to be found. Although there are guidelines for gene nomenclature, the naming process is highly creative. Human genes are often named with a gene symbol and a longer, more descriptive term; the short form is very often an abbreviation of the long form. Abbreviations in biomedical language are highly ambiguous, i.e., one gene symbol often refers to more than one gene. Using an existing abbreviation expansion algorithm,we explore MEDLINE for the use of human gene symbols derived from LocusLink. It turns out that just over 40% of these symbols occur in MEDLINE, however, many of these occurrences are not related to genes. Along the process of making an inventory, a disambiguation test collection is constructed automatically.

Algorithms↗

Using contextual queries.

Search engines generally treat search requests in isolation. The results for a given query are identical, independent of the user, or the context in which the user made the request. An approach is demonstrated that explores implicit contexts as obtained from a document the user is reading. The approach inserts into an original (web) document functionality to directly activate context driven queries that yield related articles obtained from various information sources.

Databases as Topic↗

A probabilistic similarity metric for Medline records: a model for author name disambiguation.

We present a model for automatically generating training sets and estimating the probability that a pair of Medline records sharing a last and first name initial are authored by the same individual, based on shared title words, journal name, co-authors, medical subject headings, language, and affiliation, as well as distinctive features of the name itself (i.e., presence of middle initial, suffix, and prevalence in Medline).

Authorship↗