PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “predictive modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Recalibration of risk prediction models in a large multicenter cohort of admissions to adult, general critical care units in the United Kingdom.

OBJECTIVE: To assess the performance of published risk prediction models in common use in adult critical care in the United Kingdom and to recalibrate these models in a large representative database of critical care admissions. DESIGN: Prospective cohort study. SETTING: A total of 163 adult general critical care units in England, Wales, and Northern Ireland, during the period of December 1995 to August 2003. PATIENTS: A total of 231,930 admissions, of which 141,106 met inclusion criteria and had sufficient data recorded for all risk prediction models. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: The published versions of the Acute Physiology and Chronic Health Evaluation (APACHE) II, APACHE II UK, APACHE III, Simplified Acute Physiology Score (SAPS) II, and Mortality Probability Models (MPM) II were evaluated for discrimination and calibration by means of a combination of appropriate statistical measures recommended by an expert steering committee. All models showed good discrimination (the c index varied from 0.803 to 0.832) but imperfect calibration. Recalibration of the models, which was performed by both the Cox method and re-estimating coefficients, led to improved discrimination and calibration, although all models still showed significant departures from perfect calibration. CONCLUSIONS: Risk prediction models developed in another country require validation and recalibration before being used to provide risk-adjusted outcomes within a new country setting. Periodic reassessment is beneficial to ensure calibration is maintained.

APACHE↗

Use of a prediction model for high-order multiple implantation after ovarian stimulation with gonadotropins.

OBJECTIVE: To determine prospectively the effectiveness in clinical practice of a prediction model for high-order multiple pregnancies (HOMP) (triplets or more). DESIGN: Prospective study. SETTING: University teaching hospital. PATIENT(S): Eight hundred forty-nine consecutive infertile patients undergoing a total of 1,542 treatment cycles. INTERVENTION(S): Gonadotropin ovarian stimulation or induction of ovulation without IVF MAIN OUTCOME MEASURE(S): Observed and predicted overall pregnancy rates and the incidence of HOMP. RESULT(S): The use of the prediction model (implying cancellation of all cycles at high risk for HOMP) would result in an 8% (95% confidence interval, 6.8%-9.2%) reduction of overall pregnancy rate but also in a 285% (95% CI, 279%-291%) reduction of HOMP. CONCLUSION(S): By using our prediction model, it was possible to maintain a low risk of HOMP with a good pregnancy rate in patients receiving gonadotropin ovarian stimulation or induction of ovulation without IVF.

Adult↗

Development of a predictive model for growth of Listeria monocytogenes in a skim milk medium and validation studies in a range of dairy products.

A predictive model based on growth of Listeria monocytogenes in milk is described. The main aim of this work was to generate a predictive model in milk acidified with lactic acid to mimic conditions found in a range of dairy products. A complete factorial design was employed to determine the effects of pH (4.5-7.5), temperature (3-35 degrees C) and salt concentration (0-8%) on growth of the organism. There were 210 design points and growth curves were individually fitted for the Gompertz function using non-linear regression. Descriptors of the curves, such as lag phase duration (LPD), exponential growth rate (EGR) and generation time (GT) were calculated and polynomial models were developed relating these to pH, temperature and salt concentration. The selected cubic polynomial model gave acceptable predictive estimates of growth and was stable, i.e. predictions were repeatable over the range of environmental variables studied. The model was further tested to determine its capacity for predicting growth of listeria in a range of dairy foods and these validation studies confirm its usefulness as a rapid means of estimating growth of the organism under specified environmental conditions.

Animals↗

Rapid classification of positive blood cultures: validation and modification of a prediction model.

OBJECTIVE: 1) To validate a previously developed prediction model to aid physicians in differentiating true positive blood cultures from contaminants when the laboratory first calls with a positive result, and 2) to determine whether it could be modified to make it more practical for clinical use without altering predictability. DESIGN: A prospective cohort study of hospitalized patients (validation set) who had blood cultures done over a two-month period. Data collected included the seven independent predictors in the rapid classification of positive blood cultures model. The model was modified by eliminating one of the predictors (which required clinical data) but maintaining the laboratory components (morphologic and Gram stain characteristics, number of bottles positive, and time to positivity). The "blood culture episode" was the unit of evaluation. A blood culture episode was defined as a 48-hour period beginning with the drawing of blood for the culture and included any blood cultures obtained during that time period. Receiver operating characteristic (ROC) curve analysis was used to compare the predictabilities of these models. SETTING: A 550-bed, university-affiliated county hospital that is a regional trauma center and has the only burn treatment unit in the region. PATIENTS: All adult (> or = 16 years old) patients who had blood cultures done during the study period were eligible. Only patients with positive blood cultures were included in the study. INTERVENTIONS: None. MAIN RESULTS: Of 559 blood culture episodes identified, 139 (25%) included the growth of one or more organisms; 62 (45%) of the 139 episodes represented true bacteremia. By ROC curve analysis, there was no significant difference in the mean areas under the curve (AUCs) (+/- SE) of the model in the derivation set (the previously developed model) (0.93 +/- 0.02) compared with the validation set (0.89 +/- 0.03; p = 0.29). In the validation set there was no significant difference in the mean AUCs when the model was modified (0.89 +/- 0.03) by removing the clinical component vs the unmodified model (0.89 +/- 0.03; p = 0.98). CONCLUSIONS: The rapid classification of blood cultures model was validated in a general hospital population. Predictability of the model was not altered significantly by eliminating one component that required clinical data. Because the modified model requires only laboratory information, this may allow reporting of the probability of true bacteremia at the time a positive blood culture is initially reported to physicians. This information may aid physicians in interpreting the positive blood culture.

Adult↗

Cross validation of USARIEM heat strain prediction models. U.S. ARMY Research Institute of Environmental Medicine.

HYPOTHESIS: This study was a cross validation of three heat strain prediction models developed at the U.S. Army Research Institute of Environmental Medicine: the ARIEM, HSDA, and ARIEM-EXP models ability to predict core temperature. METHODS: Seven heat-acclimated subjects completed twelve experimental tests, six in each of two hot climates, at three exercise intensities and two uniform configurations in each climate. RESULTS: Experimental results showed physiological responses as expected with heat strain increasing with work load and level of protective clothing, but with similar heat strain between the two environments matched for wet bulb, globe index. Neither the ARIEM or HSDA model closely predicted core temperatures over the course of the experiment, due mostly to an abrupt initial rise in core temperature in both models. A proportionality constant in the ARIEM-EXP buffered some of this abrupt rise. CONCLUSIONS: Comparisons of the core temperature and tolerance times data with the three models led to the conclusions that for healthy males: 1) the ARIEM and HSDA models provide conservative safety limits as a result of predicting rapid initial increases in core temperature; 2) the ARIEM-EXP most closely represents core temperature responses; 3) the ARIEM-EXP requires modifications with an alternate proportionality coefficient to increase accuracy for low metabolic cost exercise; 4) all of the models require additional input from existing research on tolerance to heat strain to better predict tolerance times; and 5) additional models should be examined to investigate the transient state of the body as it is affected by environment, clothing and exercise.

Acclimatization↗

A predictive model that evaluates the effect of growth conditions on the thermal resistance of Listeria monocytogenes.

A predictive model for Listeria monocytogenes was developed using cells grown in different pH and milkfat levels before subsequent thermal inactivation in identical pH and milkfat conditions. Inactivation of the cells used combinations of temperature (55, 60, 65 degrees C), pH (5.0, 6.0, 7.0), and milkfat (0%, 2.5%, 5.0%) in a complete 3 x 3 x 3 factorial design with each test done in triplicate. A modified Gompertz equation was used to model nonlinear survival curves with the following three parameter estimates: A for the shouldering region, B for the maximum death rate, and C for the tailing region. All treatment sets were analyzed together in a regression model using the modified Gompertz equation. There was good confidence in the overall model when it was used to predict values for the entire data set. The correlation of determination, R2, between the observed log surviving fraction (LSF) of cells from each of the conditions studied in the experiment, for the overall model was 0.811. For the A and B parameter estimates, temperature or milkfat alone, and the interaction of temperature and milkfat significantly (p < 0.05) affected the shouldering region and maximum death rate of a survival curve, respectively. These results were compared to a previously published predictive model, generated for cells grown under optimum conditions (pH 7.0, 0% milkfat), where pH was the only significant (p < 0.05) factor affecting the shoulder region. These results suggested that the conditions of the growth environment had an important impact on survival curve shape and the estimates of the predictive model. Specifically, there were more factor interactions involving temperature and milkfat level. These growth factors affected the shoulder region and maximum rate of death of the survival curve when cells were grown in identical medium conditions to which they were heated. Differences related to shouldering and inactivation rates for cells grown in different conditions may have important and practical importance for estimating inactivation of L. monocytogenes. This study provides some evidence on the importance of growing conditions when evaluating microbial heat resistance.

Animals↗

A predictive model of hopefulness for adolescents.

PURPOSE: To develop a predictive model of hopefulness using the variables of age, gender, and self-esteem among a sample of adolescents with cancer and a sample of healthy adolescents. METHODS: Forty-five healthy adolescents were individually matched with 45 adolescents with cancer on the basis of gender and age. Of the 90 subjects included in this study, 48 were male and 42 were female; half of the males (n = 24) and half of the females (n = 21) had cancer. Perceived level of self-esteem was measured using the Coppersmith Self-Esteem Inventory (SEI), and their degree of hopefulness was measured with the Hopefulness Scale for Adolescents (HSA). Data were analyzed using the Statistical Analysis System (SAS) Version 8. RESULTS: Adolescents' perceived level of self-esteem and hopefulness did not differ by gender or disease status. Patients with cancer had a significantly higher mean hopefulness score than healthy subjects (p = .031), and those adolescents with cancer did not have a lower perceived sense of self-esteem than healthy adolescents. The correlation coefficients between SEI and HSA were statistically significant for females with cancer, r = 0.723 (p < .001) and for healthy females, r = 0.676 (p < .001). In contrast, the correlations between SEI and HSA for males were not statistically significant. A model was constructed to predict a subject's hopefulness score that included the variables of self-esteem (p < .001), gender (p = .001), disease status (p = .005), and the interaction between self-esteem and gender (p = .002). CONCLUSIONS: The findings of this study demonstrate that hopefulness is a coping strategy used by female adolescents, both healthy and ill, that is closely related to their perceived sense of self-esteem.

Adaptation, Psychological↗

Outcome prediction models on admission in a medical intensive care unit: do they predict individual outcome?

Prospectively acquired data from 941 patients staying greater than 24 h in a medical ICU were analyzed to determine the relevance of scoring on ICU admission by the following methods of outcome prediction: Acute Physiology and Chronic Health Evaluation (APACHE II), Simplified Acute Physiology Score (SAPS), and Mortality Prediction Model (MPM). Analysis was performed separately for all patients (group A) and for a subsample (group B), obtained by excluding coronary care patients. Calculation of risk and classification of patients were carried out as recommended in the literature for MPM, APACHE II, and SAPS. In group A, sensitivities (correct prediction of hospital mortality) were 44.7%, 51.1%, and 21.2% and specificities (correct prediction of survival) were 84.5%, 85.4%, and 96.8%, respectively; overall correct classification rates were 73.3%, 75.8%, and 75.6%. In group B, sensitivities were slightly higher, but total correct classification rates did not reach group A levels. Goodness-of-fit testing showed low levels of fit for all methods in both groups. Application of APACHE II to diagnostic subgroups, using disease-adapted risk calculations, revealed marked inconsistencies between the estimated risk and the observed mortality. We conclude that the estimation of risk on admission by the three methods investigated might be helpful for global comparisons of ICU populations, although the lack of disease specificity reduces their applicability for severity grading of a given illness. The inaccuracy of these methods makes them ineffective for predicting individual outcome; thus, they provide little advantage in clinical decision-making.

Female↗

Sources of error in road safety scheme evaluation: a method to deal with outdated accident prediction models.

This paper considers the errors that arise in using outdated accident prediction models in road safety scheme evaluation. Methods to correct for regression-to-mean (RTM) effects in scheme evaluation normally rely on the use of accident prediction models. However, because accident risk tends to decline over time, such models tend to become outdated and the estimated treatment effect is then exaggerated. A new correction procedure is described which can effectively eliminate such errors.

Accidents, Traffic↗

[Investigation of fuzzy-clustering in octane number prediction model based on detailed hydrocarbon analysis data].

A method to establish octane number prediction model based on detailed hydrocarbon analysis (DHA) data is presented. The techniques of fuzzy-clustering and the Euclidian distance are employed to select the samples needed in pattern establishment. One hundred and fifty gasoline samples and an amount of 140 characteristic components in the DHA chromatogram of each sample are used for the fuzzy-clustering research. It is found that the 3 - 10 samples, which have the nearest Euclidian distance ( < 1.5) to the prediction sample in the same cluster, are enough to build the octane number prediction model. The experimental results proved that the model obtained according to the above method has more predictable accuracy, wider application range and higher data resource utility compared with the current prediction method.

Cluster Analysis↗

Chemotherapy-induced anaemia during adjuvant treatment for breast cancer: development of a prediction model.

BACKGROUND: At present, oncologists prescribe chemotherapy according to standard dose schedules, and as a result many patients develop serious, dose-limiting toxic effects such as anaemia. We aimed to develop a prediction model for anaemia in patients with breast cancer who were receiving adjuvant chemotherapy. METHODS: We reviewed medical records of 331 patients who had received adjuvant chemotherapy for breast cancer. Patients were divided randomly into a derivation sample (n=221) and internal-validation sample (n=110). An external sample of 119 patients enrolled onto the control group of a randomised trial of epoetin alfa was used to validate the model further. Multivariable logistic regression was applied to develop the initial model. We then developed a risk-scoring system, ranging from 0 (low risk) to 50 (high risk), based on the final regression variables. A receiver operating characteristic (ROC) curve analysis was done to measure the accuracy of the scoring system when applied to both validation samples. FINDINGS: The risk of anaemia increased as the pretreatment haemoglobin concentration decreased and was reduced with successive chemotherapy cycles. Risk was also predicted by a platelet count of 200x10(9) cells/L or less before chemotherapy, age 65 years or older, type of adjuvant chemotherapy, and use of prophylactic antibiotics. ROC analysis had acceptable areas under the curve of 0.88 for the internal-validation sample and 0.84 for the external validation sample. A risk score of > or = 24 to < 25 before chemotherapy was identified as the optimum cut-off for maximum sensitivity (83.5%) and specificity (92.3%) of the prediction model. INTERPRETATION: The application and continued refinement of this prediction model will help oncologists to identify patients at risk of developing anaemia during chemotherapy for breast cancer, and might enhance patient-centred care by the application of anaemia treatment in a proactive and appropriate way.

Aged↗

[Identification of subjects at high risk of coronary disease in a working population using a prediction model].

Identification of subjects at high risk of coronary morbidity is of major interest in the prevention of cardiovascular disease. This report describes the use of a multifactorial prediction model for the identification of high risk subjects in a French male population. The PCV-METRA study (Prévention Cardiovasculaire en Médecine du Travail) monitors risk factors of cardiovascular morbidity in a population of men and women employed in big companies in the Paris region. A model adapted from a prediction model conceived by K.M. Anderson et al. in the Framingham study was used. The modified model enables an estimation of individual coronary risk based on 7 factors: age, total cholesterol, HDL-cholesterol, systolic blood pressure, smoking, diabetes and presence of left ventricular hypertrophy, taking into account the relatively low prevalence of coronary heart disease in France. The population comprised 4,131 active men aged 30 to 65 years. The average risk at 5 years was estimated to be 1.6%. Subjects at high risk (over the 80th percentile of the risk distribution curve) usually had high blood pressures and cholesterol levels. However, nearly 30% of these subjects were neither hypertensive nor hypercholesteraemic. It is important to note that 3/4 of these smoked. Moreover, they also had low HDL-cholesterol levels. A risk table, derived from the Framingham model, is presented. This table allows estimation of individual risk at 5 years in men aged 30 to 65 years. In each age group, the comparison of individual risk with the percentiles of risk distribution in the PCV-METRA population allows identification of high-risk subjects. This study proposes a tool for identifying subjects at high risk of coronary morbidity in a French male population. This multifactorial model is particularly useful for detecting subjects with several borderline factors none of which overstep the usually accepted limits.

Adult↗

Deriving the expected utility of a predictive model when the utilities are uncertain.

Predictive models are often constructed from clinical databases with the goal of eventually helping make better clinical decisions. Evaluating models using decision theory is therefore natural. When constructing a model using statistical and machine learning methods, however, we are often uncertain about precisely how the model will be used. Thus, decision-independent measures of classification performance, such as the area under an ROC curve, are popular. As a complementary method of evaluation, we investigate techniques for deriving the expected utility of a model under uncertainty about the model's utilities. We demonstrate an example of the application of this approach to the evaluation of two models that diagnose coronary artery disease.

Artificial Intelligence↗

Developing optimal prediction models for cancer classification using gene expression data.

Microarrays can provide genome-wide expression patterns for various cancers, especially for tumor sub-types that may exhibit substantially different patient prognosis. Using such gene expression data, several approaches have been proposed to classify tumor sub-types accurately. These classification methods are not robust, and often dependent on a particular training sample for modelling, which raises issues in utilizing these methods to administer proper treatment for a future patient. We propose to construct an optimal, robust prediction model for classifying cancer sub-types using gene expression data. Our model is constructed in a step-wise fashion implementing cross-validated quadratic discriminant analysis. At each step, all identified models are validated by an independent sample of patients to develop a robust model for future data. We apply the proposed methods to two microarray data sets of cancer: the acute leukemia data by Golub et al. and the colon cancer data by Alon et al. We have found that the dimensionality of our optimal prediction models is relatively small for these cases and that our prediction models with one or two gene factors outperforms or has competing performance, especially for independent samples, to other methods based on 50 or more predictive gene factors. The methodology is implemented and developed by the procedures in R and Splus. The source code can be obtained at http://hesweb1.med.virginia.edu/bioinformatics.

Colonic Neoplasms↗

Predictive model of axillary lymph node involvement in women with small invasive breast carcinoma: axillary metastases in breast carcinoma.

BACKGROUND: Axillary lymph node involvement (ALNI) remains the most accurate predictive factor for recurrence risk and survival in patients with invasive breast carcinoma (IBC) and is an essential element in therapeutic decisions. However, axillary dissection (AD) is responsible for several side effects and is now discussed in small IBC. The objective of this study was to define a predictive model of ALNI by using clinical and histologic variables available before surgery. METHODS: The authors studied 795 cases of IBC (T0, T1, T2 < or = 4 cm; N0; M0) treated between 1980 and 1997 by conservative surgery and radiation therapy. All cases had axillary dissection with at least 10 lymph nodes removed. A stepwise logistic regression analysis was performed to build a predictive model of ALNI. The authors then used the jackknife resampling technique to produce unbiased estimates of the probabilities of ALNI along with their confidence intervals. RESULTS: The global ALNI rate was 25.7%. The final predictive model included clinical tumor size, location, and histologic subtype and grade as variables independently associated with ALNI. The estimated probability of ALNI varied from 6% to 45%, according to case characteristics for these variables. CONCLUSIONS: These results show that the omission of AD in surgical procedures for these tumors is debatable. Even when ALNI rates were low, the superior bounds of the confidence intervals could be high. Consequently, we do not recommend to omit AD in women whose estimated risks are higher than 25%. Women with a risk of ALNI lower than 25% could benefit from the sentinel lymph node procedure with, likewise, a limited risk of false-negative.

Adult↗

Risk factors and predictive models of giant cell arteritis in polymyalgia rheumatica.

OBJECTIVE: To identify in polymyalgia rheumatica the best set of predictors for a positive temporal artery biopsy and to define predictive models with either a high or low probability of giant cell arteritis (GCA). PATIENTS AND METHODS: Retrospective study of 227 patients, 137 with polymyalgia rheumatica unassociated with arteritis (group A) and 90 with polymyalgia associated with biopsy-proven giant cell arteritis (group B or training set). Data on demographic features, clinical and laboratory abnormalities were collected. Risk factors for arteritis were estimated by nonlinear logistic regressions. Simple predictive models were constructed with those predictors more related to arteritis by multivariable analysis. These models were then tested in group B and in 89 cases of arteritis without polymyalgia rheumatica (group C or test set). RESULTS: The best predictors of arteritis were a new headache odds ratio (OR) 13.6 (95% confidence interval [CI] 4.7 to 39.3); age at onset < 70 years OR 0.11 (CI 0.04 to 0.35); abnormal temporal arteries OR 4.2 (CI 1.3 to 13.7); raised liver enzymes OR 2.9 (CI 1.1 to 7.8), and jaw claudication OR 4.8 (CI 1.0 to 22.7). Amaurosis was only observed in patients with arteritis. Three subsets had a very high risk of arteritis: (1) Patients with recent headache, abnormal arteries, and > or = 70 years at disease onset: sensitivity 44%, positive predictive value (PPV) 93%, likelihood ratio (LR) 20.3; (2) patients with a new headache, jaw claudication, and abnormal arteries: sensitivity 34.4%, PPV 96.9%, LR 47.2; and (3) those, that in addition to the last 3 features, were > or = 70 years of age at disease onset: sensitivity 26.7%, PPV 100%. We could also identify a subset with a very low risk of arteritis constituted by patients < 70 years, without headache, and with clinically normal temporal arteries: sensitivity 1.1%, PPV 1.7%, LR 0.03. In group C or the test set, these four predictive models correctly identified 57.3%, 29.2%, 23.6, and 3.4% of patients, respectively. CONCLUSIONS: In polymyalgia rheumatica it is feasible to identify subsets with a very high likelihood of GCA. Although in some of these subsets the diagnosis of arteritis is almost certain, we suggest that even then it should be confirmed by temporal artery biopsy. By contrast, in those patients with polymyalgia < 70 years and without cranial features of giant cell arteritis, the risk of vasculitis is so low that the biopsy could be initially avoided and the patient treated with low-dose corticosteroids.

Aged↗

[The OCRA method: updating of reference values and prediction models of occurrence of work-related musculo-skeletal diseases of the upper limbs (UL-WMSDs) in working populations exposed to repetitive movements and exertions of the upper limbs].

BACKGROUND: The paper considers a database of old (already published) and new data concerning 23 groups of workers (Total number of subjects examined=5373) with different levels of exposure to repetitive movements of the upper limbs: for all these groups data were available regarding exposure indexes (OCRA index and Checklist "OCRA" score) and clinically determined UL-WMSD outcomes (PA=Prevalence of workers Affected by one or more UL- WMSDs; PC=Prevalence of single diagnosed Cases of an UL- WMSDs). OBJECTIVES: Using these data, the paper aimed at presenting and discussing the results obtained in order to estimate: new critical values of OCRA index for discriminating different exposure levels (green, yellow, red areas); new prediction models of expected PA and PC in exposed populations based on exposure indexes. METHODS: New critical values of the OCRA index (and, consequently, of the checklist score) were estimated by an original approach in which data of the effect variable PA in a reference population not exposed to the specific risks were combined with the regression function between OCRA and PA, as resulting from the 23 available groups. RESULTS: The resulting critical values and the consequent classification system of the OCRA index and of the checklist score are synthetically reported in the following table: [table: see text]. The best simple regression functions between exposure indexes (OCRA; checklist) and health outcome variables (PA; PC) were then sought, in order to obtain prediction models of effects starting from exposure. The following were the main prediction models derived from the available set of data (standard error of b in brackets): [formula: see text]. Finally, a multiple regression model was computed for estimating PA (Y) based on OCRA index and gender structure of the group (SEXRATIO=n. females x 100/n. total) with its 5 degrees and 95 degrees percentiles (in brackets); the resulting model was. Y = 2.02 (1.72-2.32) x OCRA + 0.075 (0.035-0.115) x SEXRATIO. This model showed a very high association between the two independent variables and the effect variable (PA) (R2=0.96). DISCUSSION: Discussion of the results obtained considers their intrinsic limits, as they are based on prevalence studies, and also suggests due recommendations and caution in the use of the proposed classification system and prediction models when the OCRA methods are applied for the evaluation of occupational risk associated with repetitive movements of the upper limbs.

Adolescent↗

Caries prediction model in pre-school children in Riyadh, Saudi Arabia.

OBJECTIVES: To evaluate the significance of variables such as oral hygiene, dietary habits, socio-economic status and medical history of a child in assessing the level of caries risk and to generate a caries prediction model for pre-school Saudi children. DESIGN: Cross-sectional study of pre-school children. SETTING: Clinics and schools in Riyadh, Saudi Arabia. SAMPLE AND METHODS: A sample of 446 Saudi pre-school children, 199 males and 247 females, with a mean age of 4.13 years, were selected at random from clinics and schools. Selection was limited to subjects who either had no caries (dmft = 0) or who had high caries experience (dmft > 8). Each child was examined for caries experience and oral hygiene status. Their mothers were interviewed through a standardized questionnaire for information about oral hygiene habits of the children, diet history, childhood illness and socio-economic status. RESULTS: There was a highly significant difference between the two groups in: debris index (P < 0.0001), aged child started tooth brushing, (P < 0.0001), age breastfeeding was stopped (P < 0.005), nocturnal bottle feeding with milk formula (P < 0.001), use of sweetened milk (P < 0.0001), frequency of use of soft drinks (P < 0.0005), frequency of consumption of sweets (P < 0.0001), and age at first dental visit (P < 0.0001). A caries prediction model developed through stepwise multivariate Logistic Regression (LR) analyses showed debris index, use of sweetened milk in bottle, frequency of consumption of soft drinks, frequency of intake of sweets and child's age at the first dental visit to be significant. Predictive probability of the model was 86.31% with a sensitivity of 90.1% and a specificity of 80.6%. CONCLUSIONS: Risk factors for dental caries have been identified and a caries prediction model has been developed for Saudi pre-school children. The prediction model, if verified, may provide with guidance in identifying high caries risk Saudi preschool children as targets for preventive programmes.

Bottle Feeding↗