PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data commons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Creating a general practice national minimum data set: present possibility or future plan?

AIM: To assess the feasibility of implementing the recommendations of the New Zealand National Minimum Data Set working party in computerised general practices. METHOD: Doctors from 12 computerised general practices belonging to the Royal New Zealand College of General Practitioners' Dunedin Research Unit Computer Network participated in the study (five Dunedin practices, four in rural Otago and Southland, and three in Christchurch). A three-month sample of data was extracted from practice computers and evaluated for completeness and compliance to the national minimum data set structure. Rates of recording practice identifier, provider, patient identifiers, sex, ethnicity, government subsidy eligibility, consultation identifier and date, prescriptions and Read codes were calculated for each practice. RESULTS: Apart from data recorded automatically by computers, there was a wide range in the extent of missing data. Of the data requiring manual computer entry, patient demography and subsidy eligibility were most comprehensively recorded (date of birth 99.9%, sex 99.6%, eligibility to subsidies 98.5%). Data with little immediate clinical or management relevance were poorly recorded (Read codes 32.4% and ethnicity 5.0%). CONCLUSIONS: It is possible to derive a common minimum data set from different computerised general practices. However some data elements will be missing unless suitable education and support are provided for the doctors and other staff members who record patient information.

Data Collection↗

Paired individual and mean postsynaptic currents recorded in four-cell networks of Aplysia.

1. Presynaptic neurons B4 and B5 of Aplysia buccal ganglia produce similar inhibitory postsynaptic currents (PSCs) in several postsynaptic follower cells. Two previous papers have characterized the variability of synaptic current amplitude and decay time both for individual PSCs and also for mean values characterizing synapses and have compared PSC amplitude and time course at different synapses sharing a common presynaptic or postsynaptic neuron. 2. To distinguish similarity in synaptic current amplitude or decay introduced by a common pre- or postsynaptic neuron from similarity because of factors common to the particular ganglion or animal, paired synapses were analyzed in four-cell networks in which each of two identified presynaptic neurons produces similar PSCs in each of two postsynaptic cells. Pairing the same synaptic data by common presynaptic or postsynaptic neuron tests if the presynaptic or postsynaptic element partially specifies a parameter; cross-pairing controls for more global factors. Paired values of peak conductance gpeak and decay time constant tau were compared for both individual sequential PSCs and for averages characterizing synapses. Analyses of individual PSCs examine processes affecting synaptic plasticity on a time scale of seconds to minutes, while average values compare more slowly varying factors. 3. Peak amplitudes were compared between individual PSCs in each of 24 paired sets. Correlations of gpeak fluctuations were significantly larger for PSCs produced by the same presynaptic neuron than for postsynaptic or cross pairings (P less than 0.05), consistent with partially correlated fluctuations in transmitter release at different presynaptic terminals. 4. Firing rates of individual presynaptic neurons were modulated to induce variability of test PSCs. These manipulations altered synaptic peak amplitudes in paired postsynaptic neurons, although not to the same degree. Manipulation of a single presynaptic neuron modulated input from that neuron alone to common postsynaptic cells without any effect on input from the paired presynaptic neuron. When fluctuations in the amplitude of gpeak were examined in runs incorporating presynaptic modulation, correlations were strong for sets of PSCs sharing a common presynaptic neuron (R = 0.87), significantly greater (P less than 0.001) than for other pairings. 5. In contrast to the partial presynaptic specification of fluctuations of individual PSCs, values of synaptic amplitude and time course averaged over 21-132 PSCs at a given synapse reflect postsynaptic determinants. Mean values of gpeak characterizing synapses paired by common postsynaptic cell are highly similar (P = 0.0001), in contrast to the lack of similarity seen when the same data are presynaptically (P = 0.11) or cross (P = 0.36) paired.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Use of medical insurance claims data for occupational health research.

OBJECTIVE: The objective of this study was to demonstrate that health claims data, widely available due to the unique nature of the U.S. healthcare system, can be linked to other relevant databases such as personnel files and exposure data maintained by large employers. These data offer great potential for occupational health research. METHODS: In this article, we describe the process for linking claims data to industrial hygiene exposure data and personnel files of a single large employer to conduct epidemiologic research. RESULTS: Our results demonstrate the ability to replicate previously published findings using commonly maintained data sets and illustrate methodological issues that may arise as newer hypotheses are tested in this way. CONCLUSIONS: Health claims files offer potential for epidemiologic research in the United States, although the full extent and guidelines for successful application await further clarification through empiric research.

Adult↗

Supporting continuity of information in the patient transfer process should there be a minimum data set across care settings?

Existing clinical information systems do not facilitate the easy transfer of relevant clinical information; hence there are significant disparities in both the type and timeliness of information that accompanies a patient transfer between care settings. Patient data is commonly fragmented between various health care settings and providers, further contributing to a reduction in the quality of information that is shared. Delineating a transfer minimum data set would provide the basis for communicating consistent information between care settings.

Continuity of Patient Care↗

Using linear and non-linear regression to fit biochemical data.

For biochemists or chemists the most common form of data analysis is likely to be regression analysis. This is a technique to find the 'best' values for various experimental parameters; defined as those values which, when used in an appropriate equation, result in the minimum deviation of the calculated results from the experimental data. Despite the widespread application of regression analysis, the basis of the technique and the underlying assumptions are often poorly understood or appreciated. This article describes the basics of linear and non-linear regression, the role of 'weighting' and the potential pitfalls of such analyses.

Biochemical Phenomena↗

Analysis of consecutive pseudo-first-order reactions. I: An evaluation of available methods to calculate the rate constants from co-product or co-reactant data.

Co-product or co-reactant data is commonly used to obtain the hydrolysis rate constants for two-step consecutive pseudo-first-order reactions. Different methods to calculate the rate constants from experimental data were evaluated using data simulated with and without a +/- 2% random error. The results of this analysis indicate that the ability to determine the value and accuracy of the rate constants obtained by the different methods depends on the k2/k1 ratio. In certain ranges of k2/k1 ratios, erroneous results are obtained which have occasionally led to incorrect conclusions by authors in the literature.

Chemical Phenomena↗

The bootstrap: a technique for data-driven statistics. Using computer-intensive analyses to explore experimental data.

BACKGROUND: The concept of resampling data--more commonly referred to as bootstrapping--has been in use for more than three decades. Bootstrapping has considerable theoretical advantages when it is applied to non-Gaussian data. Most of the published literature is concerned with the mathematical aspects of the bootstrap but increasingly this technique is being utilized in medical and other fields. METHODS: I reviewed the published literature following a 1994 publication assessing the transfer of technology, including the bootstrap, to the biomedical literature. RESULTS: In the ten-year period following that 1994 paper there were 1679 published references to the technique in Medline. In that same time period the following citations were found in the four major medical journals-British Medical Journal (48), JAMA (51), Lancet (52) and the New England Journal of Medicine (45). CONTENT: I introduce the basic theory of the bootstrap, the jackknife, and permutation tests. The bootstrap is used to estimate the accuracy of an estimator such as the standard error, a confidence interval, or the bias of an estimator. The technique may be useful for analysing smallish expensive-to-collect data sets where prior information is sparse, distributional assumptions are unclear, and where further data may be difficult to acquire. Some of the elementary uses of bootstrapping are illustrated by considering the calculation of confidence intervals such as for reference ranges or for experimental data findings, hypothesis testing such as comparing experimental findings, linear regression, and correlation when studying association and prediction of variables, non-linear regression such as used in immunoassay techniques, and ROC curve processing. CONCLUSIONS: These techniques can supplement current nonparametric statistical methods and should be included, where appropriate, in the armamentarium of data processing methodologies.

Computers↗

Whose data set is it anyway? Sharing raw data from randomized trials.

BACKGROUND: Sharing of raw research data is common in many areas of medical research, genomics being perhaps the most well-known example. In the clinical trial community investigators routinely refuse to share raw data from a randomized trial without giving a reason. DISCUSSION: Data sharing benefits numerous research-related activities: reproducing analyses; testing secondary hypotheses; developing and evaluating novel statistical methods; teaching; aiding design of future trials; meta-analysis; and, possibly, preventing error, fraud and selective reporting. Clinical trialists, however, sometimes appear overly concerned with being scooped and with misrepresentation of their work. Both possibilities can be avoided with simple measures such as inclusion of the original trialists as co-authors on any publication resulting from data sharing. Moreover, if we treat any data set as belonging to the patients who comprise it, rather than the investigators, such concerns fall away. CONCLUSION: Technological developments, particularly the Internet, have made data sharing generally a trivial logistical problem. Data sharing should come to be seen as an inherent part of conducting a randomized trial, similar to the way in which we consider ethical review and publication of study results. Journals and funding bodies should insist that trialists make raw data available, for example, by publishing data on the Web. If the clinical trial community continues to fail with respect to data sharing, we will only strengthen the public perception that we do clinical trials to benefit ourselves, not our patients.

Editorial↗

Common nursing terminology for clinical information systems.

UNLABELLED: The lack of professional agreement upon chosen terminology in nursing detracts from the role of Clinical Information Systems (CIS) as central repositories of patient health records. The purposes of this paper are: (1) Identification of common terminology for clinical nursing information in CHS according to the following stages: patient history of health and illnesses; nursing assessment; nursing interventions and outcomes. (2) Implementation of the common terminology into computerized applications in several nursing settings. The sample included 224 nurses divided into four groups. Each group was asked to identify the common initial data for patient history and nursing interventions, based on professional experience, expertise, clinical standards and organizational / legal policy. The identification of nursing assessments and outcomes was done according to evidenced-based Clinical Guide-Lines (CGL) for each nursing setting. The CGL were chosen as a source for assessment and outcome classification for two main reasons. First, the CGL include criteria of the clinical state by the degree of severity base, which are acceptable and comprehensible to other disciplines within the healthcare system. Second, the lack of evidence-based researches related to clinical nursing outcomes. RESULTS: Standard patient history of health and illnesses (admission and discharge) was developed for all departments in the hospital with flexibility to add any specific clinical data upon requirement. A total of 62 nursing assessments / outcomes were identified from the CGL in the four chosen nursing settings. 43 (70%) nursing assessments / outcomes were common both for nursing practice in hospitals and community clinics. 30 (40%) were implemented in the community clinics CIS application, 19 (31%) in the oncology CIS application, and 16 (26%) in the delivery CIS application. The groups identified a total of 70 nursing interventions. 49 (70%) nursing interventions were common both for nursing practice in hospitals and community clinics. 59 (84%) were implemented in the community clinics CIS application, 18 (26%) in the oncology CIS application, and 29 (41%) in the delivery CIS application. For summary, the definition process, including computerization, spread across four years. The community CIS application serves about 1500 clinics in CHS Israel (which employs about 2500 nurses). The admission and discharge CIS application serves 7 general hospitals, and is currently implemented in the internal and surgical departments (about 30 departments, 35 average beds each). The oncology CIS application is implemented in two oncology centers, and the delivery CIS application will soon be implemented in 8 hospitals.

Delivery of Health Care↗

Instrument monitoring, data sharing, and archiving using Common Instrument Middleware Architecture (CIMA).

The Common Instrument Middleware Architecture (CIMA) aims at Grid-enabling a wide range of scientific instruments and sensors to enable easy access to and sharing and storage of data produced by these instruments and sensors. This paper describes the implementation of CIMA applied to the field of single-crystal X-ray crystallography. To allow the researchers to easily view the current and past data streams from the instruments or sensors in a laboratory, a crystallography portal and associated portlets were developed for this application. The CIMA-based crystallography system provides an opportunity for anyone with Web access to observe and use crystallographic and other data from laboratories that previously had only limited access.

Journal Article↗

Application of computer technology to the collection, analysis and use of veterinary data.

The value of a common pool of veterinary data, using clinical records from general practices, welfare organisations, research bodies and veterinary schools is described. Developments in computer technology are outlined and the computer's application to integrated data collection, storage, querying and dissemination is indicated. Proposals for a computerised integrated veterinary clinical data base, using a standard coded case record, are presented.

Computers↗

Descriptive analytical data and consequences for calculation of common reference intervals in the Nordic Reference Interval Project 2000.

In the Nordic Reference Interval Project (NORIP), data from 102 Nordic clinical chemical laboratories were obtained. Each laboratory reported analytical data on up to 25 of the most commonly used clinical biochemical properties, including results from each of a minimum of 25 reference individuals. A reference material consisting of a liquid frozen pool of serum with values traceable to reference methods (used as the project "calibrator" for non-enzymes to correct reference values) was measured together with other serum pool controls in each laboratory in the same analytical series as the project samples. The data on the controls were used to evaluate the analytical quality of the routine methods. For reference interval calculations, only such reference values on enzymes were accepted that were obtained by applying the International Federation of Clinical Chemistry (IFCC) compatible methods (37 degrees C), while "calibrator"-corrected reference values were used in the cases of non-enzymes. For each property, gender- and age-specific reference intervals were estimated, based on simple non-parametric calculations and using objective criteria to perform partitioning into subgroups. It is concluded that the same reference intervals are applicable in all five Nordic countries. The following descriptive data for the considered properties are presented in the tables: number of measurement values from each country and measurement system, certified/indicative target values for controls, differences between methods and measurement systems together with coefficients of variation, effects of control correction on the measurement values, differences between subgroups as determined by age, gender, country and material, and comparison of the new reference intervals with those presented in standard textbooks. The 25 components involved in this project were (listed in alphabetical order): Alanine transaminase, albumin, alkaline phosphatase, amylase, amylase pancreatic type, aspartate transaminase, bilirubin, calcium, carbamide, cholesterol, creatine kinase, creatininium, gamma-glutamyltransferase, glucose, HDL-cholesterol, iron, iron-binding capacity, lactate dehydrogenase, magnesium, phosphate, potassium, protein, sodium, triglyceride and urate.

Blood Chemical Analysis↗

Cell and tumor classification using gene expression data: construction of forests.

The advent of gene chips has led to a promising technology for cell, tumor, and cancer classification. We exploit and expand the methodology of recursive partitioning trees for tumor and cell classification from microarray gene expression data. To improve classification and prediction accuracy, we introduce a deterministic procedure to form forests of classification trees and compare their performance with extant alternatives. When two published and commonly used data sets are used, we find that the deterministic forests perform similarly to the random forests in terms of the error rate obtained from the leave-one-out procedure, and all of the forests are far better than the single trees. In addition, we provide graphical presentations to facilitate interpretation of complex forests and compare our findings with the current biological literature. In addition to numerical improvement, the main advantage of deterministic forests is reproducibility and scientific interpretability of all steps in tree construction.

Cells↗

Haplotype frequency estimation error analysis in the presence of missing genotype data.

BACKGROUND: Increasingly researchers are turning to the use of haplotype analysis as a tool in population studies, the investigation of linkage disequilibrium, and candidate gene analysis. When the phase of the data is unknown, computational methods, in particular those employing the Expectation-Maximisation (EM) algorithm, are frequently used for estimating the phase and frequency of the underlying haplotypes. These methods have proved very successful, predicting the phase-known frequencies from data for which the phase is unknown with a high degree of accuracy. Recently there has been much speculation as to the effect of unknown, or missing allelic data - a common phenomenon even with modern automated DNA analysis techniques - on the performance of EM-based methods. To this end an EM-based program, modified to accommodate missing data, has been developed, incorporating non-parametric bootstrapping for the calculation of accurate confidence intervals. RESULTS: Here we present the results of the analyses of various data sets in which randomly selected known alleles have been relabelled as missing. Remarkably, we find that the absence of up to 30% of the data in both biallelic and multiallelic data sets with moderate to strong levels of linkage disequilibrium can be tolerated. Additionally, the frequencies of haplotypes which predominate in the complete data analysis remain essentially the same after the addition of the random noise caused by missing data. CONCLUSIONS: These findings have important implications for the area of data gathering. It may be concluded that small levels of drop out in the data do not affect the overall accuracy of haplotype analysis perceptibly, and that, given recent findings on the effect of inaccurate data, ambiguous data points are best treated as unknown.

Alleles↗

Perspective of an emergency physician group as a data provider for syndromic surveillance.

The need for enhanced biologic surveillance has led to the search for new sources of data. Beginning in September 2001, Emergency Medical Associates (EMA) of New Jersey, an emergency physician group practice, undertook a series of surveillance projects in collaboration with state and federal agencies. This paper examines EMA's motivations and concerns and discusses the collaborative opportunities available to data suppliers for syndromic surveillance. Motivations for supplying data included altruism and public service, previous involvement in terrorism and disaster preparedness, academic research interests, and the opportunity to find added value in the group's existing information systems. Concerns and barriers included cost, maintaining patient confidentiality, and challenges in interacting with the public health community. The extensive and carefully maintained electronic medical record enabled EMA to conduct multiple studies in collaboration with state and federal agencies. The electronic medical record provides useful data that might be more sensitive and specific in detecting outbreaks than the patient-chief-complaint data more commonly used for surveillance. EMA's experience also indicates that opportunities exist for the public health community to work with emergency physicians and emergency physician groups as suppliers of data. Such collaborations not only are useful for syndromic surveillance systems but also can help build relations that might facilitate a response to an actual biologic attack.

Bioterrorism↗

Quantifying the pathology of neurodegenerative disorders: quantitative measurements, sampling strategies and data analysis.

The use of quantitative methods has become increasingly important in the study of neurodegenerative disease. Disorders such as Alzheimer's disease (AD) are characterized by the formation of discrete, microscopic, pathological lesions which play an important role in pathological diagnosis. This article reviews the advantages and limitations of the different methods of quantifying the abundance of pathological lesions in histological sections, including estimates of density, frequency, coverage, and the use of semiquantitative scores. The major sampling methods by which these quantitative measures can be obtained from histological sections, including plot or quadrat sampling, transect sampling, and point-quarter sampling, are also described. In addition, the data analysis methods commonly used to analyse quantitative data in neuropathology, including analyses of variance (anova) and principal components analysis (PCA), are discussed. These methods are illustrated with reference to particular problems in the pathological diagnosis of AD and dementia with Lewy bodies (DLB).

Alzheimer Disease↗

Analyzing qualitative data.

Data analysis in qualitative research is a creative process. As the instrument of data analysis, the researcher explores and reflects on the meaning of the data. In most qualitative traditions, the data analysis phase overlaps the data collection phase. As data analysis proceeds, the researcher moves back and forth between data analysis and data collection in order to create and explain the findings. Using data from the authors' research, common techniques of data analysis in qualitative research are presented.

Data Interpretation, Statistical↗