PubMed HealthSearch

PubMed · 8664590

Incomplete data sets: coping with inadequate databases.

Abstract

Three problems arise in handling numerical values in databases: bad data, missing data, and sloppy data. The effects of bad data are mitigated by using statistical subterfuges such as robust statistics or outlier removal. Missing data are replaced by creating a substitute through interpolation or by using statistics appropriate to unbalanced designs. Sloppy, semiquantitative data are relegated to innocuous positions by using nonparametric, rank, or attribute statistics. These techniques are illustrated by the telephone directory, a database of carcinogenicity test results, and a database of precision parameters derived from method performance (collaborative) studies.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

R H Albert, W Horwitz. Incomplete data sets: coping with inadequate databases.. https://pubmed.ncbi.nlm.nih.gov/8664590/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Commentary on the application of (Q)SAR to the toxicological evaluation of existing chemicals.

For ethical and financial reasons it is impossible to perform thorough toxicological testing for all of the more than 100,000 substances registered in the European Inventory of Existing Substances. It was therefore investigated whether the application of (quantitative) structure-activity relationships (QSAR) with commercially available computer programs could predict the toxicological profile and help identify those substances requiring priority toxicological testing. Whereas predictions with respect to complex endpoints such as carcinogenicity, chronic toxicity and teratogenicity are still disappointing, more reliable predictions should be forthcoming in the immediate future for sensitisation, mutagenicity and genotoxicity endpoints.

Carcinogenicity Tests