PubMed Health⌕ Search

Biomedical subjects

Kornél Markó

Publications and source records attributed to Kornél Markó.

7 recordsLinked to original sources

Automatic lexeme acquisition for a multilingual medical subword thesaurus.

PURPOSE: We present a method for the automated acquisition of a multilingual medical lexicon (for Spanish, French and Swedish) to be used within the framework of a medical cross-language text retrieval system. METHODS: For the lexical acquisition process, we incorporate seed lexicons and lists of trusted term translations derived from the UMLS Metathesaurus. The seed lexicons for Spanish, French and Swedish are automatically generated from (previously manually constructed) Portuguese, German and English sources by simple string transformations. Lexical and semantic hypotheses are then validated by processing pairs of term translations. In a last step, we use the cleaned list of "approved" translations in order to augment, step by step, the target dictionaries by processing the parallel corpora in terms of co-occurrence patterns of hypothesized translation equivalents which cannot be derived by simple character substitutions. RESULTS: An existing multilingual lexicon for the medical domain with about 60,000 entries for English, German, and Portuguese was automatically augmented by more then 17,000 new lexemes for Spanish, French, and Swedish. CONCLUSIONS: Our approach constitutes a promising method for the automated creation of new lexicon entries and their linkage to semantic identifiers.

Electronic Data Processing↗

A language classifier that automatically divides medical documents for experts and health care consumers.

We propose a pipelined system for the automatic classification of medical documents according to their language (English, Spanish and German) and their target user group (medical experts vs. health care consumers). We use a simple n-gram based categorization model and present experimental results for both classification tasks. We also demonstrate how this methodology can be integrated into a health care document retrieval system.

Germany↗

Cross-lingual alignment of biomedical acronyms and their expansions.

We propose a method that aligns biomedical acronyms and their definitions across different languages. The approach is based upon a freely available tool for the extraction of abbreviations together with their expansions, and the subsequent normalization of language-specific variants, synonyms, and translations of the extracted acronym definitions. In this step, acronym expansions are mapped onto a language-independent concept-layer on which intra- as well as interlingual comparisons are drawn.

Germany↗

Automatic lexicon acquisition for a medical cross-language information retrieval system.

We present a method for the automated acquisition of a multilingual medical lexicon (for Spanish and Swedish) to be used within the framework of a medical cross-language text retrieval system. We incorporate seed lexicons and parallel corpora derived from the UMLS Metathesaurus. The seed lexicons for Spanish and Swedish are automatically generated from (previously manually constructed) Portuguese, German and English sources. Lexical and semantic hypotheses are then validated making iterative use of co-occurrence patterns of hypothesized translation synonyms in the parallel corpora.

Humans↗

Multilingual biomedical dictionary.

We present a unique technique to create a multilingual biomedical dictionary, based on a methodology called Morpho-Semantic indexing. Our approach closes a gap caused by the absence of free available multilingual medical dictionaries and the lack of accuracy of non-medical electronic translation tools. We first explain the underlying technology followed by a description of the dictionary interface, which makes use of a multilingual subword thesaurus and of statistical information from a domain-specific, multilingual corpus.

Abstracting and Indexing↗

A CLIR Interface to a Web search engine.

Medical document retrieval presents a unique combination of challenges for the design and implementation of retrieval engines. We introduce a method to meet these challenges by implementing a multilingual retrieval interface for biomedical content in the World Wide Web. To this end we developed an automated method for interlingual query construction by which a standard Web search engine is enabled to process non-English queries from the biomedical domain in order to retrieve English documents.

Abstracting and Indexing↗

Cross-language MeSH indexing using morpho-semantic normalization.

We consider three alternative procedures for the automatic indexing of medical documents using MeSH thesaurus identifiers as target units (document descriptors). Rather than considering complete words as the starting point of the indexing procedure, we here propose morphologically plausible subwords as basic units from which MeSH terms are derived. We describe the morphological segmentation and normalization procedures, as well as the mappings from subwords to MeSH terms, and discuss results from an evaluation carried out on a German-language corpus.

Abstracting and Indexing↗