PubMed Health⌕ Search

Biomedical subjects

Riccardo Bellazzi

Publications and source records attributed to Riccardo Bellazzi.

At least 19 recordsLinked to original sources

Predictive data mining in clinical medicine: current issues and guidelines.

BACKGROUND: The widespread availability of new computational methods and tools for data analysis and predictive modeling requires medical informatics researchers and practitioners to systematically select the most appropriate strategy to cope with clinical prediction problems. In particular, the collection of methods known as 'data mining' offers methodological and technical solutions to deal with the analysis of medical data and construction of prediction models. A large variety of these methods requires general and simple guidelines that may help practitioners in the appropriate selection of data mining tools, construction and validation of predictive models, along with the dissemination of predictive models within clinical environments. PURPOSE: The goal of this review is to discuss the extent and role of the research area of predictive data mining and to propose a framework to cope with the problems of constructing, assessing and exploiting data mining models in clinical medicine. METHODS: We review the recent relevant work published in the area of predictive data mining in clinical medicine, highlighting critical issues and summarizing the approaches in a set of learned lessons. RESULTS: The paper provides a comprehensive review of the state of the art of predictive data mining in clinical medicine and gives guidelines to carry out data mining studies in this field. CONCLUSIONS: Predictive data mining is becoming an essential instrument for researchers and clinical practitioners in medicine. Understanding the main issues underlying these methods and the application of agreed and standardized procedures is mandatory for their deployment and the dissemination of results. Thanks to the integration of molecular and clinical data taking place within genomic medicine, the area has recently not only gained a fresh impulse but also a new set of complex problems it needs to address.

Clinical Medicine↗

A hierarchical Naïve Bayes Model for handling sample heterogeneity in classification problems: an application to tissue microarrays.

BACKGROUND: Uncertainty often affects molecular biology experiments and data for different reasons. Heterogeneity of gene or protein expression within the same tumor tissue is an example of biological uncertainty which should be taken into account when molecular markers are used in decision making. Tissue Microarray (TMA) experiments allow for large scale profiling of tissue biopsies, investigating protein patterns characterizing specific disease states. TMA studies deal with multiple sampling of the same patient, and therefore with multiple measurements of same protein target, to account for possible biological heterogeneity. The aim of this paper is to provide and validate a classification model taking into consideration the uncertainty associated with measuring replicate samples. RESULTS: We propose an extension of the well-known Naïve Bayes classifier, which accounts for biological heterogeneity in a probabilistic framework, relying on Bayesian hierarchical models. The model, which can be efficiently learned from the training dataset, exploits a closed-form of classification equation, thus providing no additional computational cost with respect to the standard Naïve Bayes classifier. We validated the approach on several simulated datasets comparing its performances with the Naïve Bayes classifier. Moreover, we demonstrated that explicitly dealing with heterogeneity can improve classification accuracy on a TMA prostate cancer dataset. CONCLUSION: The proposed Hierarchical Naïve Bayes classifier can be conveniently applied in problems where within sample heterogeneity must be taken into account, such as TMA experiments and biological contexts where several measurements (replicates) are available for the same biological sample. The performance of the new approach is better than the standard Naïve Bayes model, in particular when the within sample heterogeneity is different in the different classes.

Algorithms↗

A stochastic model to assess the variability of blood glucose time series in diabetic patients self-monitoring.

Several studies have shown that patients suffering from Diabetes Mellitus can significantly delay the onset and slow down the progression of diabetes micro- and macro-angiopathic complications through intensive monitoring and treatment. In general, intensive treatments imply a careful blood glucose level (BGL) self-monitoring. The analysis of BGL measurements is one of the most important tasks in order to assess the glucose metabolic control and to revise the therapeutic protocol. Recent clinical studies have shown the correlation between the glucose variability and the long-term diabetes related complications. In this paper, we propose a stochastic model to extract the time course of such variability from the self-monitoring BGL time series. This information can be conveniently combined with other analysis to evaluate the adequacy of the therapeutic protocol and to highlight periods characterized by an increasing glucose instability. The method here proposed has been validated on two simulated data sets and tested with success in the retrospective analysis of three patients' data sets.

Algorithms↗

Inferring gene expression networks via static and dynamic data integration.

This paper presents a novel approach for the extraction of gene regulatory networks from DNA microarray data. The approach is characterized by the integration of data coming from static and dynamic experiments, exploiting also prior knowledge on the biological process under analysis. A starting network topology is built by analyzing gene expression data measured during knockout experiments. The analysis of time series expression profiles allows to derive the complete network structure and to learn a model of the gene expression dynamics: to this aim a genetic algorithm search coupled with a regression model of the gene interactions is exploited. The method has been applied to the reconstruction of a network of genes involved into the Saccharomyces Cerevisiae cell cycle. The proposed approach was able to reconstruct known relationships among genes and to provide meaningful biological results.

Artificial Intelligence↗

Case-based retrieval to support the treatment of end stage renal failure patients.

OBJECTIVE: In the present paper, we describe an application of case-based retrieval to the domain of end stage renal failure patients, treated with hemodialysis. MATERIALS AND METHODS: Defining a dialysis session as a case, retrieval of past similar cases has to operate both on static and on dynamic features, since most of the monitoring variables of a dialysis session are time series. Retrieval is then articulated as a two-step procedure: (1) classification, based on static features and (2) intra-class retrieval, in which dynamic features are considered. As regards step (2), we concentrate on a classical dimensionality reduction technique for time series allowing for efficient indexing, namely discrete Fourier transform (DFT). Thanks to specific index structures (i.e. k -d trees), range queries (on local feature similarity) can be efficiently performed on our case base, allowing the physician to examine the most similar stored dialysis sessions with respect to the current one. RESULTS: The retrieval tool has been positively tested on real patients' data, coming from the nephrology and dialysis unit of the Vigevano hospital, in Italy. CONCLUSIONS: The overall system can be seen as a means for supporting quality assessment of the hemodialysis service, providing a useful input from the knowledge management perspective.

Decision Support Systems, Clinical↗

The relationship between focal seizures and sleep: an analysis of the cyclic alternating pattern.

PURPOSE: The aim of this study was to examine the relationship between focal epileptic seizures and sleep through analysis of the cyclic alternating pattern (CAP). METHODS: We analyzed the recordings of a total of 56 nocturnal partial seizures (13 occurring as isolated events and 43 "in clusters" during a full-night ambulatory polysomnography) in 12 adult patients affected by localization-related focal epilepsy (7 males, 5 females; mean age 35+12.8 years; range 25-71). RESULTS: On exact binomial distribution analysis, seizures were more frequent in CAP than in non-CAP sleep, and in CAP phase A than in CAP phase B (p<0.001). Seizures occurring in clusters were more frequently associated with CAP sleep (p<0.05) than isolated seizures and first seizures of seizure clusters. Increase of CAP rate during the 30-min of sleep period after the occurrence of a seizure was documented in both cluster and isolated seizures. CONCLUSIONS: Our data indicate that an intra-sleep condition of highly fluctuating vigilance constitutes the real substrate for the occurrence of epileptic seizures regardless of the NREM stage in which they occur. The occurrence of partial seizures during sleep appears to induce sleep instability, which may in turn result in the facilitation of clusters of seizures. Together, these data highlight the potential of sleep fragmentation to modulate intra-sleep seizure occurrence.

Adult↗

Reduced sampling schedule for the glucose minimal model: importance of Bayesian estimation.

The minimal model (MM) of glucose kinetics during an intravenous glucose tolerance test (IVGTT) is widely used in clinical studies to measure metabolic indexes such as glucose effectiveness (S(G)) and insulin sensitivity (S(I)). The standard (frequent) IVGTT sampling schedule (FSS) for MM identification consists of 30 points over 4 h. To facilitate clinical application of the MM, reduced sampling schedules (RSS) of 13-14 samples have also been derived for normal subjects. These RSS are especially appealing in large-scale studies. However, with RSS, the precision of S(G) and S(I) estimates deteriorates and, in certain cases, becomes unacceptably poor. To overcome this difficulty, population approaches such as the iterative two-stage (ITS) approach have been recently proposed, but, besides leaving some theoretical issues open, they appear to be oversized for the problem at hand. Here, we show that a Bayesian methodology operating at the single individual level allows an accurate determination of MM parameter estimates together with a credible measure of their precision. Results of 16 subjects show that, in passing from FSS to RSS, there are no significant changes of point estimates in nearly all of the subjects and that only a limited deterioration of parameter precision occurs. In addition, in contrast with the previously proposed ITS method, credible confidence intervals (e.g., excluding negative values) are obtained. They can be crucial for a subsequent use of the estimated MM parameters, such as in classification, clustering, regression, or risk analysis.

Adult↗

Temporal data mining for the quality assessment of hemodialysis services.

OBJECTIVE: This paper describes the temporal data mining aspects of a research project that deals with the definition of methods and tools for the assessment of the clinical performance of hemodialysis (HD) services, on the basis of the time series automatically collected during hemodialysis sessions. METHODS: Intelligent data analysis and temporal data mining techniques are applied to gain insight and to discover knowledge on the causes of unsatisfactory clinical results. In particular, two new methods for association rule discovery and temporal rule discovery are applied to the time series. Such methods exploit several pre-processing techniques, comprising data reduction, multi-scale filtering and temporal abstractions. RESULTS: We have analyzed the data of more than 5800 dialysis sessions coming from 43 different patients monitored for 19 months. The qualitative rules associating the outcome parameters and the measured variables were examined by the domain experts, which were able to distinguish between rules confirming available background knowledge and unexpected but plausible rules. CONCLUSION: The new methods proposed in the paper are suitable tools for knowledge discovery in clinical time series. Their use in the context of an auditing system for dialysis management helped clinicians to improve their understanding of the patients' behavior.

Algorithms↗

TA-clustering: cluster analysis of gene expression profiles through Temporal Abstractions.

This paper describes a new technique for clustering short time series of gene expression data. The technique is a generalization of the template-based clustering and is based on a qualitative representation of profiles which are labelled using trend Temporal Abstractions (TAs); clusters are then dynamically identified on the basis of this qualitative representation. Clustering is performed in an efficient way at three different levels of aggregation of qualitative labels, each level corresponding to a distinct degree of qualitative representation. The developed TA-clustering algorithm provides an innovative way to cluster gene profiles. We show the developed method to be robust, efficient and to perform better than the standard hierarchical agglomerative clustering approach when dealing with temporal dislocations of time series. Results of the TA-clustering algorithm can be visualized as a three-level hierarchical tree of qualitative representations and as such easy to interpret. We demonstrate the utility of the proposed algorithm on a set of two simulated data sets and on a study of gene expression data from S. cerevisiae.

Algorithms↗

Random walk models for bayesian clustering of gene expression profiles.

The analysis of gene expression temporal profiles is a topic of increasing interest in functional genomics. Model-based clustering methods are particularly interesting because they are able to capture the dynamic nature of these data and to identify the optimal number of clusters. We have defined a new Bayesian method that allows us to cope with some important issues that remain unsolved in the currently available approaches: the presence of time dislocations in gene expression, the non-stationarity of the processes generating the data, and the presence of data collected on an irregular temporal grid. Our method, which is based on random walk models, requires only mild a priori assumptions about the nature of the processes generating the data and explicitly models inter-gene variability within each cluster. It has first been validated on simulated datasets and then employed for the analysis of a dataset relative to serum-stimulated fibroblasts. In all cases, the results have been promising, showing that the method can be helpful in functional genomics research.

Journal Article↗

Comparison of two temporal abstraction procedures: a case study in prediction from monitoring data.

This paper presents an empirical comparison of two temporal abstraction procedures, that were applied to derive predictive features for a prediction problem in intensive care medicine. The first procedure employs knowledge from practitioners to derive qualitative patterns of state changes; the second procedure searches through a large number of data summaries to discover those that have predictive value. The derived features were used to predict whether postsurgical patients would need mechanical ventilation longer then 24h. The data-driven temporal abstraction procedure was found to provide more informative predictors, resulting in better predictions.

Cardiac Surgical Procedures↗

Analysing Italian voluntary abortion data using a Bayesian approach to the time series decomposition.

After the approval of the law on voluntary abortion in Italy, the Italian health care system started to practice voluntary abortion before the third month of pregnancy. Since 1980, the Italian Institute of Statistics (ISTAT) has collected data on the abortion frequency per month and per administrative local areas. Although a preliminary analysis of the data showed that, after an initial increase, the number of abortions progressively lowered over years, there is no insight on the existence of periodicity in the time series and on the local effects related to the regional habits and social environments. The aim of our study is therefore to extract local trends and periodicity from the data collected by ISTAT, by combining a 'structural model' of the time series and Bayesian statistics. This paper describes both the adopted stochastic model and its Bayesian estimation through a Markov chain Monte Carlo approach on the Italian abortion data. Abortion data are analysed both at national level and in each of the 95 Italian local areas. At the national level this analysis allows extraction of a trend component that clearly shows that the voluntary abortion trend has decreased constantly since June-July 1983 until the end of the study. The periodic component shows an astonishing regularity too, suggesting that the Italian people have a seasonal preference for voluntary abortion. In particular, abortions are concentrated in the central part of the year (April-August). Finally, at the local level this analysis allows us to find similarities/differences between different areas in trends and/or in seasonal preferences.

Abortion, Legal↗

Insulin minimal model indexes and secretion: proper handling of uncertainty by a Bayesian approach.

The identification of the insulin minimal model (MM) for the estimation of insulin secretion rate (ISR) and physiological indexes (e.g. beta-cell sensitivity) requires the knowledge of C-peptide (CP) kinetics. The four parameters of the two-compartment model of CP kinetics in a given individual can be derived either from an additional bolus experiment or, more frequently, from a population model. However, in both situations, the CP kinetics is uncertain and, in MM identification, it should be treated as such. This paper shows how to handle CP kinetics uncertainty by using a Bayesian methodology. In seven subjects, MM indexes and ISR were estimated together with their confidence intervals, using either the bolus data or the population model to assess CP kinetics. The two main results that arise from the application of the new methodology are: (i) the use of the population model in place of the bolus data to determine CP kinetics does not affect, on average, the point estimates of ISR profile and MM parameters but only the confidence intervals which becomes wider (less than 50%); (ii) in both the bolus and population situation neglecting the uncertainty of CP kinetics, as done in MM literature so far, introduces no bias, on average, on point estimates of MM indexes but only an underestimation of confidence intervals.

Adult↗

Management of patients with diabetes through information technology: tools for monitoring and control of the patients' metabolic behavior.

BACKGROUND: The junction of telemedicine home monitoring with multifaceted disease management programs seems nowadays a promising direction to combine the need for an intensive approach to deal with diabetes and the pressure to contain the costs of the interventions. Several projects in the European Union and the United States are implementing information technology-based services for diabetes management using a comprehensive approach. Within these systems, the role of tools for data analysis and automatic reminder generation seems crucial to deal with the information overload that may result from large home monitoring programs. The objective of this study was to describe the automatic reminder generation system and the summary indicators used in a clinical center within the telemedicine project M2DM, funded by the European Commission, and to show their usage during a 7-month on-field testing period. METHODS: M2DM is a multi-access service for management of patients with diabetes. The basic functionality of the technical service includes a Web-based electronic medical record and messaging system, a computer telephony integration service, a smart-modem located at home, and a set of specialized software modules for automated data analysis. The information flow is regulated by a software scheduler, called the Organizer, that, on the basis of the knowledge on the health care organization, is able to automatically send e-mails and alerts notifications as well as to commit activities to software agents, such as data analysis. Thanks to this system, it was possible to define an automatic reminder system, which relies on a data analysis tool and on a number of technologies for communication. Within the M2DM system, we have also defined and implemented a number of indexes able to summarize the patients' day-by-day metabolic control. In particular, we have defined the global risk index (GRI) of developing microangiopathic complications. RESULTS: The system for generating automatic alarms and reminders coupled with the indexes for evaluating the patients' metabolic control has been used for 7 months at the Fondazione Salvatore Maugeri (FSM) in Pavia, Italy. Twenty-two patients (43 +/- 16 years old, 12 men and 10 women) have been involved; six dropped out from the study. The average number of monthly automatic messages was 29.44 +/- 9.83, i.e., about 1.8 messages per patient per month. The number of monthly alarm reminders generated by the system was 16.44 +/- 4.39, so that the number of alarms per patient was about 1. The number of messages sent by patients and physicians during the project was about 13 per month. The GRI analysis shows, during the last trimester, a slight improvement of the performance of the FSM clinic, with a decrease in the percentage of badly controlled values from 33% to 27%. Finally, we found the presence of a linear increasing correlation between the mean GRI values and the number of alarms generated by the system. CONCLUSIONS: A telemedicine system may incorporate features that make it a suitable technological backbone for implementing a disease management program. The availability of data analysis tools, automated messaging system, and summary indicators of the effectiveness of the health care program may help in defining efficient clinical interventions.

Diabetes Mellitus↗

A careflow management system for chronic patients.

The management of chronic patients is a complex process, which requires the cooperation of all primary care professionals and their interaction with specialists, laboratories and personnel of different organizations. In this paper we show how a Careflow Management System (CfMS) may represent an essential component of an innovative Health Information System (HIS) able to handle the information and communication needs underlying chronic diseases management. On the basis of a general architecture designed for chronic diseases, we describe a CfMS implementation in the area of diabetes management; such a system embeds EPR and telemedicine functionalities as end-users applications as well as a module for inter-organizational communication based on contracts and on XML messages.

Case Management↗

Design, methods, and evaluation directions of a multi-access service for the management of diabetes mellitus patients.

Recent advances in information and communication technology allow the design and testing of new models of diabetes management, which are able to provide assistance to patients regardless of their distance from the health care providers. The M2DM project, funded by the European Commission, has the specific aim to investigate the potential of novel telemedicine services in diabetes management. A multi-access system based on the integration of Web access, telephone access through interactive voice response systems, and the use of palmtops and smart modems for data downloading has been implemented. The system is based on a technological platform that allows a tight integration between the access modalities through a middle layer called the multi-access organizer. Particular attention has been devoted to the design of the evaluation scheme for the system: A randomized controlled study has been defined, with clinical, organizational, economic, usability, and users' satisfaction outcomes. The evaluation of the system started in January 2002. The system is currently used by 67 patients and seven health care providers in five medical centers across Europe. After 6 months of usage of the system no major technical problems have been encountered, and the majority of patients are using the Web and data downloading modalities with a satisfactory frequency. From a clinical viewpoint, the hemoglobin A1c (HbA1c) of both active patients and controls decreased, and the variance of HbA1c in active patients is significantly lower than the control ones. The M2DM system allows for the implementation of an easy-to-use, user-tailored telemedicine system for diabetes management. The first clinical results are encouraging and seem to substantiate the hypothesis of its clinical effectiveness.

Diabetes Mellitus↗

Telemedicine in the management of young patients with type 1 diabetes mellitus: a follow-up study.

DCCT (Diabetes Control and Complications Trial) study showed that tight metabolic control of diabetes mellitus can delay the onset and/or reduce the frequency of vascular complications. Telemedicine, i.e. telecommunications and information technologies in health care, is a useful tool to achieve the DCCT goals. Our European Community (EC) sponsored Telematic management of Insulin-Dependent Diabetes Mellitus (T-IDDM) project implements a telemedicine service through on a careful analysis of current medical practice. The system is based on two components: Patient Unit (PU) and Medical Unit (MU) connected by a Telecommunication system (TS). PU allows data collection and transmission from the patient's house to the hospital, assists self-monitoring activity and suggests insulin variations. PU communicates patient's current metabolic state the MU. MU assists the physician in periodic evaluation and suggests the prescriptions to communicate back defining a treatment protocol. TS system is based on telephone lines, relying on the Intranet technology. To test the system functionality and potential impact in type 1 diabetes clinical practice, we enrolled 6 patients (4 males and 2 females), aged 9.9-15.8 yrs, with disease duration 2.1-6.4 yrs, intensively treated. One girl run out after a 1-year follow-up HbA1c levels decreased, but not significantly. Insulin requirement reduced, significantly in 2 patients (p = 0.02 and p = 0.07). A positive correlation was between number of links and protocol changes (p = 0.01), between number of protocols changes and HbA1c decrease (p = 0.02). In pediatric patients periodical visits are necessary, but T-IDDM enables continuity of care improving access and activities. An index is represented by the high number of messages between the 2 Units, seeming weekly exchange.

Adolescent↗