PubMed Health⌕ Search

Biomedical subjects

Barend Mons

Publications and source records attributed to Barend Mons.

10 recordsLinked to original sources

Databases for knowledge discovery. Examples from biomedicine and health care.

Examples are given of the use of large research databases for knowledge discovery. Such databases are not only increasingly used for research in the 'hard' mathematics-based disciplines such as physics and engineering but also in more 'soft' disciplines, such as sociology, psychology and, in general, the humanities. In between the 'hard' and the 'soft' disciplines lie disciplines such as biomedicine and health care, from which we have selected our illustrations. This latter area can be subdivided into: (1) fundamental biomedical research, related to the 'hard' scientific approach; (2) clinical research, using both 'hard' and 'soft' data and (3) population-based research, which can be subdivided into prospective and retrospective research. The examples that we shall offer are representative for using computers in scientific research in general, but in medical and health informatics in particular.

Biomedical Research↗

Thesaurus-based disambiguation of gene symbols.

BACKGROUND: Massive text mining of the biological literature holds great promise of relating disparate information and discovering new knowledge. However, disambiguation of gene symbols is a major bottleneck. RESULTS: We developed a simple thesaurus-based disambiguation algorithm that can operate with very little training data. The thesaurus comprises the information from five human genetic databases and MeSH. The extent of the homonym problem for human gene symbols is shown to be substantial (33% of the genes in our combined thesaurus had one or more ambiguous symbols), not only because one symbol can refer to multiple genes, but also because a gene symbol can have many non-gene meanings. A test set of 52,529 Medline abstracts, containing 690 ambiguous human gene symbols taken from OMIM, was automatically generated. Overall accuracy of the disambiguation algorithm was up to 92.7% on the test set. CONCLUSION: The ambiguity of human gene symbols is substantial, not only because one symbol may denote multiple genes but particularly because many symbols have other, non-gene meanings. The proposed disambiguation approach resolves most ambiguities in our test set with high accuracy, including the important gene/not a gene decisions. The algorithm is fast and scalable, enabling gene-symbol disambiguation in massive text mining applications.

Algorithms↗

Which gene did you mean?

Computational Biology needs computer-readable information records. Increasingly, meta-analysed and pre-digested information is being used in the follow up of high throughput experiments and other investigations that yield massive data sets. Semantic enrichment of plain text is crucial for computer aided analysis. In general people will think about semantic tagging as just another form of text mining, and that term has quite a negative connotation in the minds of some biologists who have been disappointed by classical approaches of text mining. Efforts so far have tried to develop tools and technologies that retrospectively extract the correct information from text, which is usually full of ambiguities. Although remarkable results have been obtained in experimental circumstances, the wide spread use of information mining tools is lagging behind earlier expectations. This commentary proposes to make semantic tagging an integral process to electronic publishing.

Abstracting and Indexing↗

Word sense disambiguation in the biomedical domain: an overview.

There is a trend towards automatic analysis of large amounts of literature in the biomedical domain. However, this can be effective only if the ambiguity in natural language is resolved. In this paper, the current state of research in word sense disambiguation (WSD) is reviewed. Several methods for WSD have already been proposed, but many systems have been tested only on evaluation sets of limited size. There are currently only very few applications of WSD in the biomedical domain. The current direction of research points towards statistically based algorithms that use existing curated data and can be applied to large sets of biomedical literature. There is a need for manually tagged evaluation sets to test WSD algorithms in the biomedical domain. WSD algorithms should preferably be able to take into account both known and unknown senses of a word. Without WSD, automatic metaanalysis of large corpora of text will be error prone.

Algorithms↗

Online tools to support literature-based discovery in the life sciences.

In biomedical research, the amount of experimental data and published scientific information is overwhelming and ever increasing, which may inhibit rather than stimulate scientific progress. Not only are text-mining and information extraction tools needed to render the biomedical literature accessible but the results of these tools can also assist researchers in the formulation and evaluation of novel hypotheses. This requires an additional set of technological approaches that are defined here as literature-based discovery (LBD) tools. Recently, several LBD tools have been developed for this purpose and a few well-motivated, specific and directly testable hypotheses have been published, some of which have even been validated experimentally. This paper presents an overview of recent LBD research and discusses methodology, results and online tools that are available to the scientific community.

Abstracting and Indexing↗

Contextual annotation of web pages for interactive browsing.

With the information on the World Wide Web and in specialized databases exploding, researchers and physicians are in dire need to browse efficiently though the large corpus of information resources in their field of interest. The focus is not any longer to find everything related to your interest, but it shifts to zooming in, based on context and expanding again in neighboring knowledge domains. This paper describes an attempt to develop a completely new, interactive way of browsing distributed corpora of information without the need for multiple different queries in different information resources. Classical search engines generally treat search requests in isolation. The results for a given query are identical, and do not automatically take on board the context in which the user made the request. The system described here explores implicit contexts as obtained from the document that the user is reading. The new approach merges the searching and browsing into one combined "read-and-search" mode and alleviates the shift users are normally forced to between searching and reading.

Hypermedia↗

Ambiguity of human gene symbols in LocusLink and MEDLINE: creating an inventory and a disambiguation test collection.

Genes are discovered almost on a daily basis and new names have to be found. Although there are guidelines for gene nomenclature, the naming process is highly creative. Human genes are often named with a gene symbol and a longer, more descriptive term; the short form is very often an abbreviation of the long form. Abbreviations in biomedical language are highly ambiguous, i.e., one gene symbol often refers to more than one gene. Using an existing abbreviation expansion algorithm,we explore MEDLINE for the use of human gene symbols derived from LocusLink. It turns out that just over 40% of these symbols occur in MEDLINE, however, many of these occurrences are not related to genes. Along the process of making an inventory, a disambiguation test collection is constructed automatically.

Algorithms↗

Mining microarray datasets aided by knowledge stored in literature.

DNA microarray technology produces large amounts of data. For data mining of these datasets, background information on genes can be helpful. Unfortunately most information is stored in free text. Here, we present an approach to use this information for DNA microarray data mining.

Databases, Genetic↗

Using contextual queries.

Search engines generally treat search requests in isolation. The results for a given query are identical, independent of the user, or the context in which the user made the request. An approach is demonstrated that explores implicit contexts as obtained from a document the user is reading. The approach inserts into an original (web) document functionality to directly activate context driven queries that yield related articles obtained from various information sources.

Databases as Topic↗

Research for research: tools for knowledge discovery and visualization.

This paper describes a method to construct from a set of documents a spatial representation that can be used for information retrieval and knowledge discovery. The proposed method has been implemented in a prototype system and allows the researcher to browse interactively and in real-time a network of relationships obtained from a set of full text articles. These relationships are combined with the potential relationships between concepts as defined in the UMLS semantic network. The browser allows the user to select a seed term and find all related concepts, to find a path between concepts (hypothesis testing), and to retrieve the references to documents or database entries that support the relationship between concepts.

Biomedical Research↗