PubMed Health⌕ Search

PubMed · 9145539

Development of computerized storage facilities for twin data: a relational database system for a twin register.

Abstract

Many twin registers hold information on flat file systems such as those provided by statistical packages or spreadsheets. Demographic details may be maintained separately from data collected in multiple different studies, leading to considerable problems with data consistency, redundancy, and integration. Ad hoc requests may be difficult. Implementation of a relational database system permits storage and maintenance of all records, simple data entry and validation procedures, linking of information from different projects with security of access, and the flexibility to provide rapid answers to ad hoc enquiries using standard Structured Query Language (SQL). Twin data provide a challenge for relational database design which rests on the technique of normalization and the use of unique identifiers to access associated groups of variables; for twins, "uniqueness" must preserve identification of both the pair and the individual twin subjects in the data structure to enable flexible access to and analysis of the data. An application on the Institute of Psychiatry Volunteer Twin Register (IOPVTR) database is described, through reference to one study of a sample of the twins, with simulated data. We show how a balance of adherence to database design principles and attention to ongoing clerical and research procedures has been used to produce an integrated, flexible, and open-ended system.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A M Macdonald, S A Hamer. 1997. Development of computerized storage facilities for twin data: a relational database system for a twin register.. https://doi.org/10.1023/a%3A1025655023496

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Probability estimation when some observations are grouped.

This paper considers the use of additional questions for decreasing survey non-response rates and an approach for estimating a probability based on the results obtained. In a survey, the respondents are asked to answer an original question and follow-up questions, where the answers for the follow-up questions are grouped answers for the original question. For example, respondents are asked to provide an exact number of incidents, but in cases of 'Do not know' or 'Refuse' responses, they are subsequently asked to pick an answer from a less specific categorical scale. The new estimator obtains smaller variance asymptotically and does not depend on a distribution family. This method is applied to income questions in a survey regarding injury prevention and behaviours. Another application is survey data on intimate partner violence, where some amendments were applied for incorporating post-stratification weights and for using non-random grouping. For additional illustration, an example of parameter estimation on artificially generated data is presented.

Data Collection↗