PubMed Health⌕ Search

Biomedical subjects

M Picavet

Publications and source records attributed to M Picavet.

3 recordsLinked to original sources

A preprocessing method for improving data mining techniques. Application to a large medical diabetes database.

The Knowledge Discovery in Databases (KDD) methodology seems to be attractive on the analyze of large clinical databases. In the KDD process, the preprocessing step (data cleaning and handling of missing values) is paramount since it conditions the quality of the results obtained by data mining procedures and represents about 80% of the whole project time. The aims of the present study were to analyze this step and provide tools to handle inconsistent data and missing values. We have broken down the process into 3 main stages: data cleaning--explanatory study of missing values--choice of the procedure used for handling missing values. The data cleaning stage was based on a system of logical rules to correct mistakes and on cluster analysis to discard the poorly filled files. The missing-data mechanism was analyzed by means of multivariate statistical procedures. Two methods to deal with missing values were compared: imputation by the most common value (mode) and imputation using decision trees. This study was performed on a large medical diabetes database (23,601 patients) including numerous missing values. A system of logical rules allowed to correct mistakes on essential parameters (for example, the type of diabetes). Cluster analysis allowed to identify 10% of poorly filled files. After multivariate analysis, the missing-data mechanism could be considered as random. For variables with low number of missing values (< 10%) and categories (< 4), imputation using decision trees provided better results than imputation by mode.

Data Interpretation, Statistical↗

Personalising e-learning modules: targeting Rasmussen levels using XML.

The development of Internet technologies has made it possible to increase the number and the diversity of on-line resources for teachers and students. Initiatives like the French-speaking Virtual Medical University Project (UMVF) try to organise the access to these resources. But both teachers and students are working on a partly redundant subset of knowledge. From the analysis of some French courses we propose a model for knowledge organisation derived from Rasmussen's stepladder. In the context of decision-making Rasmussen has identified skill-based, rule-based and knowledge-based levels for the mental process. In the medical context of problem-solving, we apply these three levels to the definition of three students levels: beginners, intermediate-level learners, experts. Based on our model, we build a representation of the hierarchical structure of data using XML language. We use XSLT Transformation Language in order to filter relevant data according to student level and to propose an appropriate display on students' terminal. The model and the XML implementation we define help to design tools for building personalised e-learning modules.

Academic Medical Centers↗

From data collection to knowledge data discovery: a medical application of data mining.

Prison inmates are exposed to a variety of major risk factors (psychiatric disorders, suicide attempts, illicit drug use). From 1986 to 1996, the USA prison population more than doubled while in France, it increased from 35655 in 1980 to 51623 in 1995. In spite of these findings, very little information concerning the inmates population is available. At the present time, there is a desire to adopt a policy based on the prevention of recidivism, on adequate release planning and referrals to community-based services. The aim of the RAPPEL project was to build an information system for assessing the social and health status of prison inmates. The pilot project was set up at the prison of Loos and allowed the collection and analysis of nearly 15000 records. The aim of this paper is to present the extension of the project consisting in developing a regional network grouping 11 jails. Information locally available will serve as the basis for the information system of regional jails. Data mining techniques will provide solutions for the extraction of new information. Three data mining tools were experimented : association rules, classification trees and clustering. Further extension consists in a distributed approach allowing direct access to the information system by WEB tools.

Classification↗