PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data commons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Estimating missing data: an iterative regression approach.

The problem of missing data is common in all fields of science. Various methods of estimating missing values in a dataset exist, such as deletion of cases, insertion of sample mean, and linear regression. Each approach presents problems inherent in the method itself or in the nature of the pattern of missing data. We report a method that (1) is more general in application and (2) provides better estimates than traditional approaches, such as one-step regression. The model is general in that it may be applied to singular matrices, such as small datasets or those that contain dummy or index variables. The strength of the model is that it builds a regression equation iteratively, using a bootstrap method. The precision of the regressed estimates of a variable increases as regressed estimates of the predictor variables improve. We illustrate this method with a set of measurements of European Upper Paleolithic and Mesolithic human postcranial remains, as well as a set of primate anthropometric data. First, simulation tests using the primate data set involved randomly turning 20% of the values to "missing". In each case, the first iteration produced significantly better estimates than other estimating techniques. Second, we applied our method to the incomplete set of human postcranial measurements. MISDAT estimates always perform better than replacement of missing data by means and better than classical multiple regression. As with classical multiple regression, MISDAT performs when squared multiple correlation values approach the reliability of the measurement to be estimated, e.g., above about 0. 8.

Animals↗

A nonparametric method for the derivation of alpha/beta ratios from the effect of fractionated irradiations.

Multifractionation isoeffect data are commonly analysed under the assumption that cell survival determines the observed tissue or tumour response, and that it follows a linear-quadratic dose dependence. The analysis is employed to derive the alpha/beta ratios of the linear-quadratic dose dependence, and different methods have been developed for this purpose. A common method uses the so-called Fe plot. A more complex but also more rigorous method has been introduced by Lam et al. (1979). Their method, which is based on numerical optimization procedures, is generalized and somewhat simplified in the present study. Tumour-regrowth data are used to explain the nonparametric procedure which provides alpha/beta ratios without the need to postulate analytical expressions for the relationship between cell survival and regrowth delay.

Animals↗

Analysis of rodent growth data in toxicology studies.

To evaluate compound-related effects on the growth of rodents, body weight and food consumption data are commonly collected either weekly or biweekly in toxicology studies. Body weight gain, food consumption relative to body weight, and efficiency of food utilization can be derived from body weight and food consumption for each animal in an attempt to better understand the compound-related effects. These five parameters are commonly analyzed in toxicology studies for each sex using a one-factor analysis of variance (ANOVA) at each collection point. The objective of this manuscript is to present an alternative approach to the evaluation of compound-related effects on body weight and food consumption data from both subchronic and chronic rodent toxicology studies. This approach is to perform a repeated-measures ANOVA on a selected set of parameters and analysis intervals. Compared with a standard one-factor ANOVA, this approach uses a statistical analysis method that has greater power and reduces the number of false-positive claims, and consequently provides a succinct yet comprehensive summary of the compound-related effects. Data from a mouse carcinogenicity study are included to illustrate this repeated-measures ANOVA approach to analyzing growth data in contrast with the one-factor ANOVA approach.

Analysis of Variance↗

Simulation model for enumeration of Salmonella on chicken as a function of PCR detection time score and sample size: implications for risk assessment.

A data gap commonly identified in risk assessments is the lack of quantitative information on the contamination of food with pathogens. A simulation model that predicts the incidence and distribution of Salmonella contamination on chicken as a function of PCR detection time score and sample size was developed with data from challenge studies with preenrichment samples that were composed of 25 g of chicken and 225 ml of buffered peptone water inoculated with 10(0.7) to 10(6) Salmonella and incubated at 37 degrees C. At 0, 2, 4, 6, 8, 10, 12, and 24 h of incubation, subsamples were collected and tested for Salmonella by PCR, and a PCR detection time score based on the widths of the bands in the electrophoresis gel was obtained for each preenrichment sample. Standard curves relating PCR detection time score to initial density of Salmonella inoculated were developed for sterile and nonsterile preenrichment samples. Presence of other microorganisms in the preenrichment sample decreased the PCR detection time score at low (<10(2) per 25 g) but not at high (>10(2) per 25 g) initial densities of Salmonella and resulted in a nonlinear standard curve rather than the linear standard curve obtained for sterile samples. The predicted incidence and distribution of Salmonella contamination on chicken increased in a nonlinear manner as sample size increased from 25 to 500 g. The new method reduced the time and cost of Salmonella enumeration by eliminating the selective enrichment, selective plating, and confirmation steps of the traditional most-probable-number method. Results are useful for risk assessment because they consider the uncertainty of the standard curve predictions and because they provide distributions of Salmonella contamination for different size samples of chicken that can be directly used in risk assessment.

Animals↗

Effects of polyhalogenated aromatic hydrocarbons and related contaminants on common tern reproduction: integration of biological, biochemical, and chemical data.

In eight Dutch or Belgian common tern (Sterna hirundo) colonies, breeding biology and food choice were determined, and 15 second eggs were collected from three-egg clutches for artificial incubation, biochemical analysis and analysis of yolk-sac polyhalogenated hydrocarbon (PHAH) levels. Results from these analyses were combined with biological data from the eggs remaining in each clutch. In some breeding colonies severe flooding, rainy and cold weather, and extreme predation caused extensive losses of eggs and chicks. A relationship was found between yolksac mono-ortho polychlorinated biphenyl (mo-PCB) levels and main food species (fish or insects) of the adult terns before egg-laying. Colony average breeding data differed only slightly, and were difficult to relate to PHAH-levels. When the colonies were grouped after yolksac PHAH-patterns and main food species, significant differences in average egg laying date, incubation period, egg volume and chick weight could be related to differences in yolksac PHAH and retinoid levels, and hepatic ethoxyresorufin-O-deethylase (EROD) activity. The data from all colonies also were combined into one data-set and correlated with the biochemical parameters and PHAH levels. In summary higher yolksac PHAH levels or hepatic EROD-activity correlated with and later egg laying, prolonged incubation period and smaller eggs and chicks. Lower yolksac retinoid- and plasma thyroid hormone levels, and a higher ratio of plasma retinol over yolksac retinoids correlated with later egg laying, prolonged incubation periods and smaller chicks and eggs. The dynamic environment of the terns had more obvious detrimental effects on breeding success than PHAHs. However, the more subtle effects observed for PHAHs could still be of importance during specific stress circumstances. To monitor site-specific reproduction effects, tree-nesting birds feeding on relatively big and non-migrating fishes would be most suitable. The use of specific biomarkers for exposure and effect is recommended to establish a causal relationship between a certain class of pollutants and an adverse biological effect.

Animals↗

Trends in primary care antibiotic prescribing in England 1994-1998.

PURPOSE: To investigate the changing decisions to prescribe antibiotics as manifest in the patterns of prescriptions dispensed in England, and to investigate antibiotic prescribing in different types of practice. METHODS: Antibiotic prescribing data and practice characteristics collected for every practice in England for the years 1994/5-1997/8. Morbidity data for common infections was also obtained from published sources. RESULTS: Antibiotic prescribing was related to practice characteristics, with high prescribing in deprived and single-handed practices in particular. There was a fall in antibiotic prescribing in all types of practice of practice over the period of the study. Morbidity data from other sources shows a fall in diagnosed morbidity from some infectious diseases over the same period. There were no differences in choice of antibiotic in different types of practice. CONCLUSIONS: The fall in antibiotic prescribing is universal across all kinds of practices and is possibly related to the fall in diagnoses. It is uncertain whether this reflects true morbidity.

Age Factors↗

Median-based robust algorithms for tracing neurons from noisy confocal microscope images.

This paper presents a method to exploit rank statistics to improve fully automatic tracing of neurons from noisy digital confocal microscope images. Previously proposed exploratory tracing (vectorization) algorithms work by recursively following the neuronal topology, guided by responses of multiple directional correlation kernels. These algorithms were found to fail when the data was of lower quality (noisier, less contrast, weak signal, or more discontinuous structures). This type of data is commonly encountered in the study of neuronal growth on microfabricated surfaces. We show that by partitioning the correlation kernels in the tracing algorithm into multiple subkernels, and using the median of their responses as the guiding criterion improves the tracing precision from 41% to 89% for low-quality data, with a 5% improvement in recall. Improved handling was observed for artifacts such as discontinuities and/or hollowness of structures. The new algorithms require slightly higher amounts of computation, but are still acceptably fast, typically consuming less than 2 seconds on a personal computer (Pentium III, 500 MHz, 128 MB). They produce labeling for all somas present in the field, and a graph-theoretic representation of all dendritic/axonal structures that can be edited. Topological and size measurements such as area, length, and tortuosity are derived readily. The efficiency, accuracy, and fully-automated nature of the proposed method makes it attractive for large-scale applications such as high-throughput assays in the pharmaceutical industry, and study of neuron growth on nano/micro-fabricated structures. A careful quantitative validation of the proposed algorithms is provided against manually derived tracing, using a performance measure that combines the precision and recall metrics.

Algorithms↗

Chronic liver disease in Ethiopia: a clinical study with emphasis on identifying common causes.

Between July 1986 and April 1989, 334 hospitalized adult Ethiopian patients with chronic liver disease were studied according to a protocol to define their clinical features and to identify risk factors with the aim of preventive intervention. Of these, 14 had chronic hepatitis, 208 cirrhosis and 112 hepatocellular carcinoma (HCC). Both clinical and histological diagnostic criteria were employed. A detailed questionnaire was used to document demographic and clinical data. A common clinical presentation among patients with chronic hepatitis was darkening of the face and hands with or without hypertrichosis of the face and blisters over the dorsi of the hands. This overt or latent form of porphyrea cutanea tarda (PCT) responds to chloroquine. Patients with cirrhosis of the liver commonly present for the first time with ascites, splenomegaly, haematemesis and/or melena from oesophageal varices, and mental changes due to hepatic encephalopathy. Overt or latent forms of PCT are also common features. Peculiar to these cirrhotics is the rarity of spider naevi, gynaecomastia, testicular atrophy, Dupuytren's contracture, parotid gland enlargement and clubbing of the fingers. Exhaustion, loss of appetite, rapid loss of weight, right upper quadrant and/or epigastric pain (all often of less than 6 months' duration, a big, hard, tender and grossly nodular liver with bruit, signs of portal hypertension, and/or hepatic encephalopathy, in a young male with a rapid down hill course characterize the Ethiopian patient with HCC. Serum anti-nuclear factor, anti-mitochondrial anti-bodies and anti-smooth muscle anti-bodies were absent in those with chronic hepatitis and were uncommon in the cirrhotics and HCC cases. One or more hepatitis B virus markers were found in 86% of chronic hepatitis, 88% cirrhosis and 78% HCC and the HBsAg carrier state was found in 36%, 29% and 23%, respectively. Among the HBsAg carriers, HBeAg positivity was less common than anti-HBe but anti-HDV was significantly higher than in the healthy general population. Alphafetoprotein (AFP) levels greater than 500 mg/ml were present in 16 (8%) cirrhotics and 58 (52%) patients with HCC. Histologically, 3 of the chronic hepatitis patients had progressed to cirrhosis, 8 of the cirrhotic patients had chronic active hepatitis and 85% of HCC cases occurred in a background of macronodular cirrhosis. Three cirrhotics developed HCC during follow-up.(ABSTRACT TRUNCATED AT 400 WORDS)

Adult↗

Extended least squares nonlinear regression: a possible solution to the "choice of weights" problem in analysis of individual pharmacokinetic data.

It is often difficult to specify weights for weighted least squares nonlinear regression analysis of pharmacokinetic data. Improper choice of weights may lead to inaccurate and/or imprecise estimates of pharmacokinetic parameters. Extended least squares nonlinear regression provides a possible solution to this problem by allowing the incorporation of a general parametric variance model. Weighted least squares and extended least squares analyses of data from a simulated pharmacokinetic experiment were compared. Weighted least squares analysis of the simulated data, using commonly used weighting schemes, yielded estimates of pharmacokinetic parameters that were significantly biased, whereas extended least squares estimates were unbiased. Extended least squares estimates were often significantly more precise than were weighted least squares estimates. It is suggested that extended least squares regression should be further investigated for individual pharmacokinetic data analysis.

Computers↗

Critical levels of atmospheric pollution: criteria and concepts for operational modelling of mercury in forest and lake ecosystems.

Mercury (Hg) is regarded as a major environmental concern in many regions, traditionally because of high concentrations in freshwater fish, and now also because of potential toxic effects on soil microflora. The predominant source of Hg in most watersheds is atmospheric deposition, which has increased 2- to >20-fold over the past centuries. A promising approach for supporting current European efforts to limit transboundary air pollution is the development of emission-exposure-effect relationships, with the aim of determining the critical level of atmospheric pollution (CLAP, cf. critical load) causing harm or concern in sensitive elements of the environment. This requires a quantification of slow ecosystem dynamics from short-term collections of data. Aiming at an operational tool for assessing the past and future metal contamination of terrestrial and aquatic ecosystems, we present a simple and flexible modelling concept, including ways of minimizing requirements for computation and data collection, focusing on the exposure of biota in forest soils and lakes to Hg. Issues related to the complexity of Hg biogeochemistry are addressed by (1) a model design that allows independent validation of each model unit with readily available data, (2) a process- and scale-independent model formulation based on concentration ratios and transfer factors without requiring loads and mass balance, and (3) an equilibration concept that accounts for relevant dynamics in ecosystems without long-term data collection or advanced calculations. Based on data accumulated in Sweden over the past decades, we present a model to determine the CLAP-Hg from standardized values of region- or site-specific synoptic concentrations in four key matrices of boreal watersheds: precipitation (atmospheric source), large lacustrine fish (aquatic receptor and vector), organic soil layers (terrestrial receptor proxy and temporary reservoir), as well as new and old lake sediments (archives of response dynamics). Key dynamics in watersheds are accounted for by quantifying current states of equilibration in both soils and lakes based on comparison of contamination factors in sediment cores. Future steady-state concentrations in soils and fish in single watersheds or entire regions are then determined by corresponding projection of survey data. A regional-scale application to southern Sweden suggests that the response of environmental Hg levels to changes in atmospheric Hg pollution is delayed by centuries and initially not proportional among receptors (atmosphere >> soils not equal sediments>fish; clearwater lakes >> humic lakes). This has implications for the interpretation of common survey data as well as for the implementation of pollution control strategies. Near Hg emission sources, the pollution of organic soils and clearwater lakes deserves attention. Critical receptors, however, even in remote areas, are humic waters, in which biotic Hg levels are naturally high, most likely to increase further, and at high long-term risk of exceeding the current levels of concern: </=0.5 mg (kg fw)(-1) in freshwater fish, and 0.5 mg (kg dw)(-1) in soil organic matter. If environmental Hg concentrations are to be reduced and kept below these critical limits, virtually no man-made atmospheric Hg emissions can be permitted.

Air Pollutants↗

The case of the missing data: methods of dealing with dropouts and other research vagaries.

Missing data are common in most studies, especially when subjects are followed over time. This can jeopardize the validity of a study because of reduced power to detect differences, and especially because subjects who are lost to follow-up rarely represent the group as a whole. There are several approaches to handling missing data, but some may result in biased estimates of the treatment effect, and others may overestimate the significance of the statistical tests. When cross-sectional data (for example, demographic and background information and a single outcome measurement time) are missing, replacement with the group mean leads to an underestimate of the standard deviation (SD) and inflation of the Type I error rate. Using regression estimates, especially with error built into the imputed value, lessens but does not eliminate this problem. Multiple imputation preserves the estimates of both the mean and the SD, even when a significant proportion of the data are missing. With longitudinal studies, the last observation carried forward (LOCF) approach preserves the sample size, but may make unwarranted assumptions about the missing data, resulting in either underestimating or overestimating the treatment effects. Growth curve analysis makes maximal use of the existing data and makes fewer assumptions.

Clinical Trials as Topic↗

Haplotype motifs: an algorithmic approach to locating evolutionarily conserved patterns in haploid sequences.

The promise of plentiful data on common human genetic variations has given hope that we will be able to uncover genetic factors behind common diseases that have proven difficult to locate by prior methods. Much recent interest in this problem has focused on using haplotypes (contiguous regions of correlated genetic variations), instead of the isolated variations, in order to reduce the size of the statistical analysis problem. In order to most effectively use such variation data, we will need a better understanding of haplotype structure, including both the general principles underlying haplotype structure in the human population and the specific structures found in particular genetic regions or sub-populations. This paper presents a probabilistic model for analyzing haplotype structure in a population using conserved motifs found in statistically significant sub-populations. It describes the model and computational methods for deriving the predicted motif set and haplotype structure for a population. It further presents results on simulated data, in order to validate the method, and on two real datasets from the literature, in order to illustrate its practical application.

Algorithms↗

Internet based multicenter study for thoracolumbar injuries: a new concept and preliminary results.

This article reports about the internet based, second multicenter study (MCS II) of the spine study group (AG WS) of the German trauma association (DGU). It represents a continuation of the first study conducted between the years 1994 and 1996 (MCS I). For the purpose of one common, centralised data capture methodology, a newly developed internet-based data collection system ( http://www.memdoc.org ) of the Institute for Evaluative Research in Orthopaedic Surgery of the University of Bern was used. The aim of this first publication on the MCS II was to describe in detail the new method of data collection and the structure of the developed data base system, via internet. The goal of the study was the assessment of the current state of treatment for fresh traumatic injuries of the thoracolumbar spine in the German speaking part of Europe. For that reason, we intended to collect large number of cases and representative, valid information about the radiographic, clinical and subjective treatment outcomes. Thanks to the new study design of MCS II, not only the common surgical treatment concepts, but also the new and constantly broadening spectrum of spine surgery, i.e. vertebro-/kyphoplasty, computer assisted surgery and navigation, minimal-invasive, and endoscopic techniques, documented and evaluated. We present a first statistical overview and preliminary analysis of 18 centers from Germany and Austria that participated in MCS II. A real time data capture at source was made possible by the constant availability of the data collection system via internet access. Following the principle of an application service provider, software, questionnaires and validation routines are located on a central server, which is accessed from the periphery (hospitals) by means of standard Internet browsers. By that, costly and time consuming software installation and maintenance of local data repositories are avoided and, more importantly, cumbersome migration of data into one integrated database becomes obsolete. Finally, this set-up also replaces traditional systems wherein paper questionnaires were mailed to the central study office and entered by hand whereby incomplete or incorrect forms always represent a resource consuming problem and source of error. With the new study concept and the expanded inclusion criteria of MCS II 1, 251 case histories with admission and surgical data were collected. This remarkable number of interventions documented during 24 months represents an increase of 183% compared to the previously conducted MCS I. The concept and technical feasibility of the MEMdoc data collection system was proven, as the participants of the MCS II succeeded in collecting data ever published on the largest series of patients with spinal injuries treated within a 2 year period.

Adolescent↗

Accuracy of identification of patients with immune thrombocytopenic purpura through administrative records: a data validation study.

Administrative data are commonly used to estimate the prevalence of a disease, but the validity of the coding system needs to be evaluated before its use. We assessed the validity of the International Classification of Disease, 9(th) version, Clinical Modification (ICD-9-CM) code of 287.3 for identifying patients with immune thrombocytopenic purpura (ITP). Administrative data from inpatients and outpatients seen were retrieved if the patient or insurer was billed with one of three ICD-9-CM codes for thrombocytopenic disorders, 287.3, 287.4, and 287.5, as a primary or secondary diagnosis; or was physician-identified as having ITP. The electronic medical records for these patients were systematically reviewed to identify patients with ITP and with non-ITP diagnoses. Sensitivity, specificity, positive and negative predictive values, and kappa scores were calculated separately for inpatients and outpatients. Four-hundred eighteen records were reviewed. Among inpatients, the sensitivity of code 287.3 for indicating a diagnosis of ITP was 100% [95% confidence interval 94-100%]. The specificity was 89% [95% confidence interval 84-94%]. The percent agreement was 92%, and the kappa statistic was 0.80. For outpatients, the sensitivity of the billing code 287.3 was 84% [95% confidence interval 76-91%], a conservative estimate because of how the patients with other diagnoses were selected. The specificity for outpatients was 66% [95% confidence interval 56-76%]. ICD-9-CM code 287.3 in administrative billing data is likely to be sufficiently sensitive and specific, particularly when inpatient data are used, for the estimation of the prevalence of ITP.

Accounts Payable and Receivable↗

Carotid endarterectomy with homologous vein patch angioplasty: a review of 1006 cases.

PURPOSE: Because homologous vein is rarely used in vascular reconstructions, we evaluated the homologous vein as a patch for the reconstruction of the carotid bifurcation after endarterectomy. METHODS: Excess vein harvested during open heart operations was either refrigerated in saline solution or cryopreserved in a solution of 10% dimethyl sulfoxide. Donors were tested for transmissible infections, and the veins were cultured for common pathogens. Data were analyzed from 837 consecutive patients (1006 cases) who underwent carotid endarterectomy with homologous vein patch angioplasty between 1981 and 1993. RESULTS: The perioperative mortality rate was 0.8% (eight patients). Two deaths (0.2%) were attributed to ipsilateral strokes. Ischemic strokes occurred in 12 patients (1.2%; 10 ipsilateral), and ipsilateral transient ischemic attacks occurred in three patients (0.3%). Follow-up data were obtained for 482 patients (56%; mean follow-up time, 61 months; range, 1 to 132 months). Ipsilateral recurrent symptoms occurred in eight patients (1.7%; seven strokes, one transient ischemic attack). Of the 63 late deaths (13%), the majority (25 patients; 40%) were caused by complications of coronary artery disease. The 10-year overall survival rate was 76% +/- 3.2%, and the 10-year rate of freedom from late ipsilateral morbidity was 96% +/- 1.4%. The 10-year rate of freedom from late stenosis (a reduction in diameter of > or = 20%) in the 220 arteries (22%) that were studied by duplex scan was 84% +/- 2.3%. CONCLUSIONS: The postoperative mortality and neurologic morbidity rates of carotid endarterectomy with homologous vein patch angioplasty are similar to those in the best series with all types of closure. The existing long-term follow-up data indicate that the homologous vein is a durable patch that behaves like other patches used in the same location.

Aged↗

The accuracy of seven mathematical functions in modeling dairy cattle lactation curves based on test-day records from varying sample schemes.

Daily milk yield over the course of the lactation follows a curvilinear pattern, so a suitable function is required to model this curve. In this study, 7 functions (Wood, Wilmink, Ali and Schaeffer, cubic splines, and 3 Legendre polynomials) were used to model the lactation curve at the phenotypic level, using both daily observations and data from commonly used recording schemes. The number of observations per lactation varied from 4 to 11. Several criteria based on the analysis of the real error were used to compare models. The performance of models showed few discrepancies in the comparison criteria when daily or 4-weekly (with first test at days in milk 8) data by lactation were used. The performance of the Wood, Wilmink, and Ali and Schaeffer models were highly affected by the reduction of the sample dimension. The results of this work support the idea that the performance of these models depends on the sample properties but also shows considerable variation within the sampling groups.

Animals↗

[Process indicators and standards for the evaluation of breast cancer screening programmes].

In order to obtain the maximum benefit from breast cancer screening it is essential for every programme to reach high levels of sensitivity and specificity. This can only be achieved if skill and a comprehensive quality assurance system is applied to the entire process, involving each individual part of the programme. Monitoring of outcomes and continuous evaluation of the entire screening process are key operational objectives for a successful population screening programme. The aim of this document, born in the framework of the Italian Group for Mammography Screening (GISMa), is to propose a unique methodology for collecting and reporting screening data using commonly agreed terminology, definitions and classifications. The indicators considered are those referred to the entire screening process and its sequelae, such as organizational, logistic and performance indicators. The indicators are provided under form of a synthetic and easy to use card. Every card is structured in short sections: definition, aim of the indicator, the data necessary to build it, the summarizing formula, possible problems of interpretation, the acceptable and desirable standards (derived both from the experience of national and European breast cancer screening programmes).

Breast Neoplasms↗

Interpretable gene expression classifier with an accurate and compact fuzzy rule base for microarray data analysis.

An accurate classifier with linguistic interpretability using a small number of relevant genes is beneficial to microarray data analysis and development of inexpensive diagnostic tests. Several frequently used techniques for designing classifiers of microarray data, such as support vector machine, neural networks, k-nearest neighbor, and logistic regression model, suffer from low interpretabilities. This paper proposes an interpretable gene expression classifier (named iGEC) with an accurate and compact fuzzy rule base for microarray data analysis. The design of iGEC has three objectives to be simultaneously optimized: maximal classification accuracy, minimal number of rules, and minimal number of used genes. An "intelligent" genetic algorithm IGA is used to efficiently solve the design problem with a large number of tuning parameters. The performance of iGEC is evaluated using eight commonly-used data sets. It is shown that iGEC has an accurate, concise, and interpretable rule base (1.1 rules per class) on average in terms of test classification accuracy (87.9%), rule number (3.9), and used gene number (5.0). Moreover, iGEC not only has better performance than the existing fuzzy rule-based classifier in terms of the above-mentioned objectives, but also is more accurate than some existing non-rule-based classifiers.

Algorithms↗