PubMed Health⌕ Search

Biomedical subjects

Marie-Christine Jaulent

Publications and source records attributed to Marie-Christine Jaulent.

At least 19 recordsLinked to original sources

Building an ontology of pulmonary diseases with natural language processing tools using textual corpora.

Pathologies and acts are classified in thesauri to help physicians to code their activity. In practice, the use of thesauri is not sufficient to reduce variability in coding and thesauri are not suitable for computer processing. We think the automation of the coding task requires a conceptual modeling of medical items: an ontology. Our task is to help lung specialists code acts and diagnoses with software that represents medical knowledge of this concerned specialty by an ontology. The objective of the reported work was to build an ontology of pulmonary diseases dedicated to the coding process. To carry out this objective, we develop a precise methodological process for the knowledge engineer in order to build various types of medical ontologies. This process is based on the need to express precisely in natural language the meaning of each concept using differential semantics principles. A differential ontology is a hierarchy of concepts and relationships organized according to their similarities and differences. Our main research hypothesis is to apply natural language processing tools to corpora to develop the resources needed to build the ontology. We consider two corpora, one composed of patient discharge summaries and the other being a teaching book. We propose to combine two approaches to enrich the ontology building: (i) a method which consists of building terminological resources through distributional analysis and (ii) a method based on the observation of corpus sequences in order to reveal semantic relationships. Our ontology currently includes 1550 concepts and the software implementing the coding process is still under development. Results show that the proposed approach is operational and indicates that the combination of these methods and the comparison of the resulting terminological structures give interesting clues to a knowledge engineer for the building of an ontology.

France↗

Sifting abstracts from Medline and evaluating their relevance to molecular biology.

The most important knowledge in the area of biology currently consists of raw text documents. Bibliographic databases of biomedical articles can be searched, but an efficient procedure should evaluate the relevance of documents to biology. In genetics, this challenge is even trickier, because of the lack of consistency in genes' naming tradition. We aim to define a good approach for collecting relevant abstracts for biology and for studied species and genes. Our approach relies on defining best queries, detecting and filtering best sources.

France↗

Integrating anatomical pathology to the healthcare enterprise.

For medical decisions, healthcare professionals need that all required information is both correct and easily available. We address the issue of integrating anatomical pathology department to the healthcare enterprise. The pathology workflow from order to report, including specimen process and image acquisition was modeled. Corresponding integration profiles were addressed by expansion of the IHE (Integrating the Healthcare Enterprise) initiative. Implementation using respectively DICOM Structured Report (SR) and DICOM Slide-Coordinate Microscopy (SM) was tested. The two main integration profiles--pathology general workflow and pathology image workflow--rely on 13 transactions based on HL7 or DICOM standard. We propose a model of the case in anatomical pathology and of other information entities (orders, image folders and reports) and real-world objects (specimen, tissue samples, slides, etc). Cases representation in XML schemas, based on DICOM specification, allows producing DICOM image files and reports to be stored into a PACS (Picture Archiving and Communication System.

Diagnostic Imaging↗

Mapping of the WHO-ART terminology on Snomed CT to improve grouping of related adverse drug reactions.

The WHO-ART and MedDRA terminologies used for coding adverse drug reactions (ADR) do not provide formal definitions of terms. In order to improve groupings, we propose to map ADR terms to equivalent Snomed CT concepts through UMLS Metathesaurus. We performed such mappings on WHO-ART terms and can automatically classify them using a description logic definition expressing their synonymies. Our gold standard was a set of 13 MedDRA special search categories restricted to ADR terms available in WHO-ART. The overlapping of the groupings within the new structure of WHO-ART on the manually built MedDRA search categories showed a 71% success rate. We plan to improve our method in order to retrieve associative relations between WHO-ART terms.

Adverse Drug Reaction Reporting Systems↗

Knowledge acquisition for computation of semantic distance between WHO-ART terms.

Computation of semantic distance between adverse drug reactions terms may be an efficient way to group related medical conditions in pharmacovigilance case reports. Previous experience with ICD-10 on a semantic distance tool highlighted a bottleneck related to manual description of formal definitions in large terminologies. We propose a method based on acquisition of formal definitions by knowledge extraction from UMLS and morphosemantic analysis. These formal definitions are expressed with SNOMED International terms. We provide formal definitions for 758 WHO-ART terms: 321 terms defined from UMLS, 320 terms defined using morphosemantic analysis and 117 terms defined after expert evaluation. Computation of semantic distance (e.g. k-nearest neighbours) was implemented in J2EE terminology services. Similar WHO-ART terms defined by automated knowledge acquisition and ICD terms defined manually show similar behaviour in the semantic distance tool. Our knowledge acquisition method can help us to generate new formal definitions of medical terms for our semantic distance terminology services.

Adverse Drug Reaction Reporting Systems↗

Building medical ontologies by terminology extraction from texts: an experiment for the intensive care units.

In many medical fields, maintenance, comparison and aggregation of unambiguous terminologies go through formal specialized clinical terminologies: ontologies. We describe a methodology to build medical ontology from textual reports using a natural language processing tool, the SYNTEX software. The methodology is illustrated in the surgical intensive care medical domain. We have tested the possibility for an expert to build a sizeable ontology in a reasonable time. The quality of the ontology has been evaluated according to its capacity to cover the ICD-10 terminology in the field. Finally, the methodology itself is discussed.

Humans↗

Computation of semantic similarity within an ontology of breast pathology to assist inter-observer consensus.

Computer-assisted consensus in medical imaging involves automatic comparison of morphological abnormalities observed by physicians in images. We built an ontology of morphological abnormalities in breast pathology to assist inter-observer consensus. Concepts of morphological abnormalities extracted from existing terminologies, published grading systems and medical reports were organized in an taxonomic hierarchy and furthermore linked by the relation "is a diagnostic criterion of" according to diagnostic meaning. We implemented position-based, content-based and mixed semantic similarity measures between concepts in this ontology and compared the results with experts' judgment. The position-based similarity measure using both taxonomic and non-taxonomic relations performed as well as the other measures and was used for automatic comparison of morphological abnormalities within the IDEM computer-assisted consensus platform.

Breast↗

Building an ontology of adverse drug reactions for automated signal generation in pharmacovigilance.

Automated signal generation in pharmacovigilance implements unsupervised statistical machine learning techniques in order to discover unknown adverse drug reactions (ADR) in spontaneous reporting systems. The impact of the terminology used for coding ADRs has not been addressed previously. The Medical Dictionary for Regulatory Activities (MedDRA) used worldwide in pharmacovigilance cases does not provide formal definitions of terms. We have built an ontology of ADRs to describe semantics of MedDRA terms. Ontological subsumption and approximate matching inferences allow a better grouping of medically related conditions. Signal generation performances are significantly improved but time consumption related to modelization remains very important.

Adverse Drug Reaction Reporting Systems↗

Implementation of automated signal generation in pharmacovigilance using a knowledge-based approach.

Automated signal generation is a growing field in pharmacovigilance that relies on data mining of huge spontaneous reporting systems for detecting unknown adverse drug reactions (ADR). Previous implementations of quantitative techniques did not take into account issues related to the medical dictionary for regulatory activities (MedDRA) terminology used for coding ADRs. MedDRA is a first generation terminology lacking formal definitions; grouping of similar medical conditions is not accurate due to taxonomic limitations. Our objective was to build a data-mining tool that improves signal detection algorithms by performing terminological reasoning on MedDRA codes described with the DAML+OIL description logic. We propose the PharmaMiner tool that implements quantitative techniques based on underlying statistical and bayesian models. It is a JAVA application displaying results in tabular format and performing terminological reasoning with the Racer inference engine. The mean frequency of drug-adverse effect associations in the French database was 2.66. Subsumption reasoning based on MedDRA taxonomical hierarchy produced a mean number of occurrence of 2.92 versus 3.63 (p < 0.001) obtained with a combined technique using subsumption and approximate matching reasoning based on the ontological structure. Semantic integration of terminological systems with data mining methods is a promising technique for improving machine learning in medical databases.

Adverse Drug Reaction Reporting Systems↗

Electronic implementation of guidelines in the EsPeR system: a knowledge specification method.

Despite initiatives to standardize methods for the development of clinical guidelines, several barriers hinder their integration in daily clinical practice: failure to fulfil quality criteria, poor effectiveness of their dissemination. Computerization of guidelines can favor their dissemination. The initial step of computerization is the knowledge specification from the text of the guideline. We describe the method of knowledge specification, which is used in EsPeR (Personalized Estimate of Risks), a web-based decision support system in preventive medicine, which allows, for a given person, to estimate risks and access recommendations, based on clinical profile. This method is based on a structured and systematic analysis of text allowing detailed specification of a decision tree. We use decision tables to validate the decision algorithm and decision trees to specify this algorithm, along with elementary messages of recommendation. Editing tools are used to facilitate the process of validation and the workflow between expert physicians and computer scientists. Applied to eleven different guidelines, the method allows a quick and valid computerization and integration in the EsPeR system. The method used for computerization could help to define a framework usable at the initial step of guideline development in order to produce guidelines ready for electronic implementation.

Decision Support Systems, Clinical↗

Appraisal of the MedDRA conceptual structure for describing and grouping adverse drug reactions.

Computerised queries in spontaneous reporting systems for pharmacovigilance require reliable and reproducible coding of adverse drug reactions (ADRs). The aim of the Medical Dictionary for Regulatory Activities (MedDRA) terminology is to provide an internationally approved classification for efficient communication of ADR data between countries. Several studies have evaluated the domain completeness of MedDRA and whether encoded terms are coherent with physicians' original verbatim descriptions of the ADR. MedDRA terms are organised into five levels: system organ class (SOC), high level group terms (HLGTs), high level terms (HLTs), preferred terms (PTs) and low level terms (LLTs). Although terms may belong to different SOCs, no PT is related to more than one HLT within the same SOC. This hierarchical property ensures that terms cannot be counted twice in statistical studies, though it does not allow appropriate semantic grouping of PTs. For this purpose, special search categories (SSCs) [collections of PTs assembled from various SOCs] have been introduced in MedDRA to group terms with similar meanings. However, only a small number of categories are currently available and the criteria used to construct these categories have not been clarified. The objective of this work is to determine whether MedDRA contains the structural and terminological properties to group semantically linked adverse events in order to improve the performance of spontaneous reporting systems. Rossi Mori classifies terminological systems in three categories: first-generation systems, which represent terms as strings; second-generation systems, which dissect terminological phrases into a set of simpler terms; and third-generation systems, which provide advanced features to automatically retrieve the position of new terms in the classification and group sets of meaning-related terms. We applied Cimino's desiderata to show that MedDRA is not compatible with the properties of third-generation systems. Consequently, no tool can help for the automated positioning of new terms inside the hierarchy and SSCs have to be entered manually rather than automatically using the MedDRA files. One solution could be to link MedDRA to a third-generation system. This would allow the current MedDRA structure to be kept to ensure that end users have a common view on the same data and the addition of new computational properties to MedDRA.

Adverse Drug Reaction Reporting Systems↗

Structuring Clinical Guidelines through the Recognition of Deontic Operators.

In this paper, we present a novel approach to structure Clinical Guidelines through the automatic recognition of syntactic expressions called deontic operators. We defined a grammar and a set of Finite-State Transition Networks (FSTN) to automatically recognize deontic operators in Clinical Guidelines. We then implemented a dedicated FSTN parser that identifies deontic operators and marks up their occurrences in the document, thus producing a structured version of the Guideline. We evaluated our approach on a corpus (not used to define the grammar) of 5 Clinical Guidelines. As a result, 95.5% of the occurrences of deontic expressions are correctly marked up. The automatic detection of deontic operators can be a useful step to support Clinical Guidelines encoding.

Humans↗

Integration of multiple ontologies in breast cancer pathology.

The diagnostic variability in pathology, widely reported in the literature, is partly due to the use of different classification systems by pathologists. The descriptions of morphological characteristics on the same image within different classification systems can be considered as different points of view of pathologists. Our aim is to represent the points of view of the experts in pathology during image interpretation and to propose a method ological and technical solution in order to implement interoperability between these points of view. According to the hybrid ontology approach, we developed a system in three stages consisting in 1) the representation of the various points of view in local ontologies 2) the realization of a shared vocabulary and the development of a mapping tool used to allow the matching of local ontologies and shared vocabulary 3) the development of a transcoding algorithm for the translation of a case description from one point of view to another. A first evaluation of the transcoding algorithm was conducted for 33 cases of breast pathology. Our results show that the pathologists generally produce descriptions of the cases which do not follow rigorously the interpretation rules corresponding to the point of view they assert to adopt. While most of the concepts of local ontologies can be transcoded from a local ontology to another one (varying from 62.5 % to 100% according to the local ontology), the transcoding of a description which is valid according to a certain point of view, often results in a description which is not rigorously in accordance with the new point of view. These results underline the differences of interpretation rules existing in the different points of view.

Algorithms↗

Building medical ontologies based on terminology extraction from texts: an experimentation in pneumology.

Pathologies and acts are classified in thesauri to help physicians to code their activity. In practice, the use of thesauri is not sufficient to reduce variability in coding and thesauri do not fit computer processing. We think the automation of the coding task requires a conceptual modelling of medical items: an ontology. Our objective is to help pneumologists code acts and diagnoses with a software that represents medical knowledge by an ontology of the concerned specialty. The main research hypothesis is to apply natural language processing tools to corpora to develop the resources needed to build the ontology. In this paper, our objective is twofold: we have to build the ontology of pneumology and we want to develop a methodology for the knowledge engineer to build various types of medical ontologies based on terminology extraction from texts.

Humans↗

An environment for document engineering of clinical guidelines.

In this paper, we present the G-DEE system, a document engineering environment aimed at clinical guidelines. This system represents an extension of current visual interfaces for guidelines encoding, in that it supports automatic text processing functions which identify linguistic markers of document structure, such as recommendations, thereby decreasing the complexity of operations required by the user. Such markers are identified by shallow parsing of free text and are automatically marked up as an early step of document structuring. From this first representation, it is possible to identify elements of guidelines contents, such as decision variables, and produce elements of GEM encoding, using rules defined as XSL style sheets. We tested our automatic structuring system on a set of sentences extracted from French clinical guidelines. As a result, 97% of the occurrences of deontic operators and their scopes were correctly marked up. G-DEE can be used for various purposes, from research into guidelines structure to assisting the encoding of guidelines into a GEM format or into decision rules.

Data Display↗

Patient data synchronization process in a continuity of care environment.

In a distributed patient record environment, we analyze the processes needed to ensure exchange and access to EHR data. We propose an adapted method and the tools for data synchronization. Our study takes into account the issue of user rights management for data access and decreasing the amount of data exchanged over the network. We describe a XML-based synchronization model that is portable and independent of specific medical data models. The implemented platform consists of several servers, of local network clients, of workstations running user's interfaces and of data exchange and synchronization tools.

Access to Information↗

Computerization of guidelines: a knowledge specification method to convert text to detailed decision tree for electronic implementation.

The initial step for the computerization of guidelines is the knowledge specification from the prose text of guidelines. We describe a method of knowledge specification based on a structured and systematic analysis of text allowing detailed specification of a decision tree. We use decision tables to validate the decision algorithm and decision trees to specify and represent this algorithm, along with elementary messages of recommendation. Edition tools are also necessary to facilitate the process of validation and workflow between expert physicians who will validate the specified knowledge and computer scientist who will encode the specified knowledge in a guide-line model. Applied to eleven different guidelines issued by an official agency, the method allows a quick and valid computerization and integration in a larger decision support system called EsPeR (Personalized Estimate of Risks). The quality of the text guidelines is however still to be developed further. The method used for computerization could help to define a framework usable at the initial step of guideline development in order to produce guidelines ready for electronic implementation.

Decision Making, Computer-Assisted↗

A knowledge based approach for automated signal generation in pharmacovigilance.

BACKGROUND: Pharmacovigilance experts detect new adverse drug reactions (ADR) by manually reviewing spontaneous reporting systems. Automated signal generation aims to focus the attention of experts on drug-adverse event associations which are disproportionally present in the database. Although adverse events are coded by means of controlled vocabularies such as the MedDRA dictionary, this semantic information is not taken into account for signal generation. OBJECTIVE: To improve the performance of current signal detection algorithms using knowledge based approach. METHOD: We developed a formal ontology of ADRs and built a data mining tool that uses description logic representations of MedDRA terms to group medically related case reports. RESULTS: This knowledge based approach increased the sensitivity of signal detection with no decrease in specificity. DISCUSSION: A knowledge based approach improved the performance of signal detection tools. However, the huge work-load involved in the knowledge engineering step limits the use of this approach for machine learning.

Adverse Drug Reaction Reporting Systems↗