PubMed Health⌕ Search

Biomedical subjects

Peter Christen

Publications and source records attributed to Peter Christen.

4 recordsLinked to original sources

Some methods for blindfolded record linkage.

BACKGROUND: The linkage of records which refer to the same entity in separate data collections is a common requirement in public health and biomedical research. Traditionally, record linkage techniques have required that all the identifying data in which links are sought be revealed to at least one party, often a third party. This necessarily invades personal privacy and requires complete trust in the intentions of that party and their ability to maintain security and confidentiality. Dusserre, Quantin, Bouzelat and colleagues have demonstrated that it is possible to use secure one-way hash transformations to carry out follow-up epidemiological studies without any party having to reveal identifying information about any of the subjects - a technique which we refer to as "blindfolded record linkage". A limitation of their method is that only exact comparisons of values are possible, although phonetic encoding of names and other strings can be used to allow for some types of typographical variation and data errors. METHODS: A method is described which permits the calculation of a general similarity measure, the n-gram score, without having to reveal the data being compared, albeit at some cost in computation and data communication. This method can be combined with public key cryptography and automatic estimation of linkage model parameters to create an overall system for blindfolded record linkage. RESULTS: The system described offers good protection against misdeeds or security failures by any one party, but remains vulnerable to collusion between or simultaneous compromise of two or more parties involved in the linkage operation. In order to reduce the likelihood of this, the use of last-minute allocation of tasks to substitutable servers is proposed. Proof-of-concept computer programmes written in the Python programming language are provided to illustrate the similarity comparison protocol. CONCLUSION: Although the protocols described in this paper are not unconditionally secure, they do suggest the feasibility, with the aid of modern cryptographic techniques and high speed communication networks, of a general purpose probabilistic record linkage system which permits record linkage studies to be carried out with negligible risk of invasion of personal privacy.

Computer Communication Networks↗

Characterization of cytotoxic and genotoxic effects of different compounds in CHO K5 cells with the comet assay (single-cell gel electrophoresis assay).

Different variants of the comet assay were used to study the genotoxic and cytotoxic properties of the following eight compounds: chloral hydrate, colchicine, hydroquinone, DL-menthol, mitomycin C, sodium iodoacetate, thimerosal and valinomycin. Colchicine, mitomycin C, sodium iodoacetate and thimerosal induced genotoxic effects. The other compounds were found to be inactive. The compounds were tested in the standard comet assay as well as in the all cell comet assay (recovery of floating cells after treatment), designed in our laboratory for adherently-growing cells. This latter procedure proved to be more adequate for the assessment of the cytotoxicity for some of the compounds tested (hydroquinone, DL-menthol, thimerosal, valinomycin). Colchicine was positive in the standard comet assay (3h treatment) and in the all cell comet assay (24h treatment). Sodium iodoacetate and thimerosal were positive in the standard and/or the all cell comet assay. Chloral hydrate, hydroquinone, sodium iodoacetate, mitomycin C and thimerosal were also tested in the modified comet assay using lysed cells. Mitomycin C and thimerosal showed effects in this assay, whereas sodium iodoacetate was inactive. This indicates that it does not induce direct DNA damage. Compounds that are known or suspected to form DNA-DNA cross-links or DNA-protein cross-links (chloral hydrate, hydroquinone, mitomycin C and thimerosal) were checked for their ability to reduce ethyl methanesulfonate (EMS)-induced DNA damage. This mode of action could be demonstrated for mitomycin C only.

Animals↗

Preparation of name and address data for record linkage using hidden Markov models.

BACKGROUND: Record linkage refers to the process of joining records that relate to the same entity or event in one or more data collections. In the absence of a shared, unique key, record linkage involves the comparison of ensembles of partially-identifying, non-unique data items between pairs of records. Data items with variable formats, such as names and addresses, need to be transformed and normalised in order to validly carry out these comparisons. Traditionally, deterministic rule-based data processing systems have been used to carry out this pre-processing, which is commonly referred to as "standardisation". This paper describes an alternative approach to standardisation, using a combination of lexicon-based tokenisation and probabilistic hidden Markov models (HMMs). METHODS: HMMs were trained to standardise typical Australian name and address data drawn from a range of health data collections. The accuracy of the results was compared to that produced by rule-based systems. RESULTS: Training of HMMs was found to be quick and did not require any specialised skills. For addresses, HMMs produced equal or better standardisation accuracy than a widely-used rule-based system. However, accuracy was worse when used with simpler name data. Possible reasons for this poorer performance are discussed. CONCLUSION: Lexicon-based tokenisation and HMMs provide a viable and effort-effective alternative to rule-based systems for pre-processing more complex variably formatted data such as addresses. Further work is required to improve the performance of this approach with simpler data such as names. Software which implements the methods described in this paper is freely available under an open source license for other researchers to use and improve.

Data Collection↗