PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Dataset”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Statistical comparison of two ROC-curve estimates obtained from partially-paired datasets.

The authors propose a new generalized method for ROC-curve fitting and statistical testing that allows researchers to utilize all of the data collected in an experimental comparison of two diagnostic modalities, even if some patients have not been studied with both modalities. Their new algorithm, ROCKIT, subsumes previous algorithms as special cases. It conducts all analyses available from previous ROC software and provides 95% confidence intervals for all estimates. ROCKIT was tested on more than half a million computer-simulated datasets of various sizes and configurations representing a range of population ROC curves. The algorithm successfully converged for more than 99.8% of all datasets studied. The type I error rates of the new algorithm's statistical test for differences in Az estimates were excellent for datasets typically encountered in practice, but diverged from alpha for datasets arising from some extreme situations.

Algorithms↗

A comparison of univariate and multivariate gene selection techniques for classification of cancer datasets.

BACKGROUND: Gene selection is an important step when building predictors of disease state based on gene expression data. Gene selection generally improves performance and identifies a relevant subset of genes. Many univariate and multivariate gene selection approaches have been proposed. Frequently the claim is made that genes are co-regulated (due to pathway dependencies) and that multivariate approaches are therefore per definition more desirable than univariate selection approaches. Based on the published performances of all these approaches a fair comparison of the available results can not be made. This mainly stems from two factors. First, the results are often biased, since the validation set is in one way or another involved in training the predictor, resulting in optimistically biased performance estimates. Second, the published results are often based on a small number of relatively simple datasets. Consequently no generally applicable conclusions can be drawn. RESULTS: In this study we adopted an unbiased protocol to perform a fair comparison of frequently used multivariate and univariate gene selection techniques, in combination with a ränge of classifiers. Our conclusions are based on seven gene expression datasets, across several cancer types. CONCLUSION: Our experiments illustrate that, contrary to several previous studies, in five of the seven datasets univariate selection approaches yield consistently better results than multivariate approaches. The simplest multivariate selection approach, the Top Scoring method, achieves the best results on the remaining two datasets. We conclude that the correlation structures, if present, are difficult to extract due to the small number of samples, and that consequently, overly-complex gene selection algorithms that attempt to extract these structures are prone to overtraining.

Algorithms↗

Candidate gene analysis of the Price Foundation anorexia nervosa affected relative pair dataset.

The eating disorders are severe psychiatric illnesses with significant morbidity and mortality that exhibit statistically significant familial risk and heritability, providing support for a molecular genetic approach toward defining etiological factors. An emerging candidate gene literature has concentrated on serotinergic and dopaminergic candidates. With the financial support of the Price Foundation, a group of investigators initiated an international multi-center collaboration (Price Foundation Collaborative Group) in 1995 to study the genetics of anorexia and bulimia nervosa by collecting and analyzing phenotypes and genotypes of individuals and their relatives affected with eating disorders. The first sample of families collected by this collaborative group, known as the Price Foundation Anorexia Nervosa Affected Relative Pair (AN-ARP) dataset, was ascertained on an proband affected with Anorexia Nervosa (AN), with relative pairs affected with the eating disorders AN, Bulimia Nervosa or Eating Disorders Not Otherwise Specified [1]. Biognosis U.S., Inc. was founded to identify and characterize candidate susceptibility genes for anorexia and bulimia nervosa phenotypes in the Price Foundation eating disorder datasets. During 2000-2001, Biognosis U.S., Inc. developed and implemented a research program with a focus on the analysis of candidate genes nominated by neurochemical characteristics of eating disorder patients [2], serotonergic and dopaminergic candidate gene polymorphisms [3], neuroendocrine regulation of appetite [4], and by a positional hypothesis from a linkage analysis of the AN-ARP dataset [5]. This report reviews the anorexia nervosa candidate gene literature through 2001, the candidate gene research program implemented at Biognosis U.S., Inc. and selected candidate gene findings in the AN-ARP dataset derived from that research program.

Animals↗

Participatory development of a minimum dataset for the Khayelitsha district.

BACKGROUND: Traditional 'data-led' information systems have created excessive amounts of poor-quality and poorly utilised data. The Health Information Systems Pilot Project (HISPP), a Western Cape project that started in 1996, initiated a process in one of its three pilot sites to model an alternative approach to developing a district health information system. OBJECTIVE: To develop a minimum dataset for Khayelitsha as part of an action-led district health and management information system in a participatory 'bottom-up' process. METHOD: The HISPP, in conjunction with health workers in the proposed Khayelitsha district, developed a minimum dataset through a process of defining local goals, targets and indicators. This dataset was integrated with data requirements at regional and provincial levels. RESULTS: A minimum dataset was produced that defined all the data needed according to the frequency of reporting and the level at which it was required. CONCLUSION: The HISPP has demonstrated an alternative model for defining health information needs at district level. This participatory process has enabled health workers to appraise their own information needs critically and has encouraged local use of information for planning and action.

Adult↗

Characterization of the Caucasian haplogroups present in the SWGDAM forensic mtDNA dataset for 1771 human control region sequences. Scientific Working Group on DNA Analysis Methods.

Currently, the Scientific Working Group on DNA Analysis Methods (SWGDAM) mtDNA dataset is used to infer the relative rarity of mtDNA profiles (i.e., haplotypes) obtained from evidence samples and for identification of missing persons. The Caucasian haplogroup patterns in this forensic dataset have been characterized using phylogenetic methods. The assessment reveals that the dataset is relevant and representative of U.S. and European Caucasians. The comparisons carried out were both the observation of variable sites within the control region (CR) and the selection of a subset of these sites, which partition the variation within human mtDNA control region sequences into clusters (i.e., haplogroups). The aligned sequence matrix was analyzed to determine both single nucleotide polymorphisms (SNPs) in a phylogenetic context, as well as to check and standardize haplogroup designations with a focus on determining the characters that define these groups. To evaluate the dataset for forensic utility, the haplogroup identifications and frequencies were compared with those reported from other published studies.

DNA, Mitochondrial↗

Developing metadata to organize public health datasets.

The Centers for Disease Control and Prevention (CDC) has available a large number of datasets from previous and current surveillance and research.1 Until now, these datasets have not been catalogued. Metadata would organize these datasets and enhance CDC's ability to efficiently use this data to quickly gain the broader view of the nation's health status to effectively carry out public health activities. This project was to develop metadata for cataloguing CDC datasets and a system that would allow researchers to search at least 95% of databases within CDC based on the most relevant criteria for research. It also explored the need to involve stakeholders and users in the project. The resulting metadata and system are available only to CDC researchers on the CDC intranet.

Cataloging↗

Revision of the Belgian Nursing Minimum Dataset: From data to information.

The Ministry of Public Health commissioned a research project to the Catholic University of Leuven and the University Hospital of Liège to revise the Belgian Nursing Minimum Dataset (B-NMDS). The study started in 2000 and will end with the implementation of the revised B-NMDS in January 2007. The study entailed four major phases. The first phase involved the development of a conceptual framework based on a literature review and secondary data analysis. The second phase focused on language development and development of a data collection tool. The third phase focused on data collection and validation of the new tool. In the fourth phase the validity and reliability of the dataset was tested. The new dataset is without avail if it is not leading to new information. Four applications of the dataset has been defined from the beginning: evaluation of the appropriateness of stay (AEP) in the hospital, nurse staffing, hospital financing and quality management. The aim of this paper is to describe how the B-NMDS can contribute to each of these applications.

Belgium↗

MelanoDB: A dataset of clinical and molecular features of patients with advanced melanoma treated with MAPK inhibitors.

MAPK inhibitors (MAPKi) have revolutionized the treatment of patients with advanced melanoma. However, primary and acquired resistance mechanisms limit their efficacy. Predicting MAPKi response from the tumor baseline features remains challenging due to the limited size of patient cohorts. Therefore, we collected data from nine different patient cohorts (total n = 417 patients with advanced melanoma treated with MAPKi) to identify clinical and molecular features. Our curated dataset, named MelanoDB, includes whole or partial exome sequencing data for 191 patients, copy number alteration information for 66 patients, and gene expression data for 132 patients. We provide a web application to explore the integrated dataset and data distribution across the collected studies, and we share this dataset with the scientific community according to the Findable, Accessible, Interoperable, Reusable (FAIR) principles.

Humans↗

Genetic analysis of IDDM: the GAW5 multiplex family dataset.

In a collaborative effort by 12 centers from Europe and North America, data were assembled from 94 multiplex families with insulin-dependent diabetes mellitus (IDDM) for analysis of genetic and other factors of possible etiological importance. The dataset contains information on the following genetic markers: HLA-DR beta and -DQ beta restriction fragment length polymorphisms (RFLPs), three RFLPs detected with two probes that map 5' to the insulin gene, the serologically defined HLA loci, and the immunoglobulin allotypes. Data also were included for auto-antibodies to insulin and pancreatic islet cells as possible indicators of pathogenesis and for antibodies to certain viruses that have been implicated as "triggering" agents in IDDM. Medical history of family members was obtained by means of a uniform questionnaire. Identical copies of the dataset were distributed to anyone wishing to participate in the analysis for the IDDM component of GAW5. The multiplex IDDM family dataset is now available on request for further analysis.

Adolescent↗

Combining National Health Interview Survey Datasets: issues and approaches.

This paper identifies issues, special problems, and approaches in preparing estimates when combining National Health Interview Survey (NHIS) datasets. Such datasets can be used to produce estimates jointly based on individual NHIS survey components. This paper illustrates several issues associated with the analysis of multiple related datasets.

Analysis of Variance↗

Relative quality of different systematic datasets for cetartiodactyl mammals: assessments within a combined analysis framework.

High congruence, support, stability, resolution and decisiveness are seen as positive attributes by many systematists. Within a cladistic context, the consistency index, the retention index, branch support, data decisiveness, the number of nodes resolved in a strict consensus tree and the incongruence length difference are direct measures of these qualities. Phylogenetic analyses of 29 datasets for cetartiodactyl mammals show that for a particular character partition, these indices can vary radically in separate versus combined analysis of datasets. The quality of any single dataset is of little importance in comparison to a thorough sampling of the available character space.

Animals↗

MR neurography with multiplanar reconstruction of 3D MRI datasets: an anatomical study and clinical applications.

INTRODUCTION: Extracranial MR neurography has so far mainly been used with 2D datasets. We investigated the use of 3D datasets for peripheral neurography of the sciatic nerve. METHODS: A total of 40 thighs (20 healthy volunteers) were examined with a coronally oriented magnetization-prepared rapid acquisition gradient echo sequence with isotropic voxels of 1 x 1 x 1 mm and a field of view of 500 mm. Anatomical landmarks were palpated and marked with MRI markers. After MR scanning, the sciatic nerve was identified by two readers independently in the resulting 3D dataset. RESULTS: In every volunteer, the sciatic nerve could be identified bilaterally over the whole length of the thigh, even in areas of close contact to isointense muscles. The landmark of the greater trochanter was falsely palpated by 2.2 cm, and the knee joint by 1 cm. The mean distance between the bifurcation of the sciatic nerve and the knee-joint gap was 6 cm (+/-1.8 cm). The mean results of the two readers differed by 1-6%. CONCLUSION: With the described method of MR neurography, the sciatic nerve was depicted reliably and objectively in great anatomical detail over the whole length of the thigh. Important anatomical information can be obtained. The clinical applications of MR neurography for the brachial plexus and lumbosacral plexus/sciatic nerve are discussed.

Adolescent↗

Guidelines for managing datasets, programs and printouts in scientific research.

This paper presents guidelines and naming conventions to assist in managing datasets, programs and printouts for small to medium sized research projects. First, the limits and nature of the type of project are considered. A number of definitions are included to help further specify the problem. Four main areas are addressed. First, a consistent set of rules for naming programs and datasets is presented. Next, techniques for storing and archiving the programs and datasets are presented, using the naming conventions just developed. Third, rules for printout management are given, based on the naming rules. Finally, the documentation of all components of the project is discussed. Although trivial to implement, adherence to the naming rules provides a sound basis for a high quality, easy-to-use documentation system.

Computers↗

Metallic and organic contaminants in sediments of Sydney Harbour, Australia and vicinity-- a chemical dataset for evaluating sediment quality guidelines.

An internally consistent dataset comprising 103 surficial estuarine sediment samples were collected from Sydney Harbour, Australia and locations south of Sydney. This paper describes the chemical characteristics of the dataset and evaluates its suitability for use in evaluating biological effects-based sediment quality guidelines (SQGs). The sediments contained mixtures of chemicals, the most prevalent chemical classes being metals and polycyclic aromatic hydrocarbons, whereas sediments from coastal lakes/estuaries south of Sydney had low concentrations of contaminants. Maximum concentrations of the prevalent contaminants zinc, lead, copper and pyrene were 11,300, 1,420, 1,060 mg kg(-1) and 23,300 microg kg(-1), respectively. For the majority of samples, concentrations of individual chemicals exceeded most effects-based SQGs that have been adopted for use in Australia, implying occasional or frequent adverse biological effects are expected. Comparing mixtures of contaminants to ranges in numbers of SQGs exceeded and mean SQG quotients showed that most samples (57% to 68%) had contamination characteristics associated with moderate probabilities (30% to 52%) of acute toxicity, based on North American data. A smaller proportion of samples (15% to 17%) had contamination characteristics associated with high probabilities (74% to 85%) of toxicity. The wide range of chemicals and concentrations, associated with low, medium and high probabilities of toxicity, indicated that the dataset was suitable for future use in evaluating predictive abilities of SQGs. This is relevant, given the recent introduction of North American-derived SQGs for Australia.

Australia↗

Revising the Belgian Nursing Minimum Dataset: from concept to implementation.

The process of revising the Belgian Nursing Minimum Dataset (B-NMDS) started in 2000 and entailed four major phases. The first phase (June-October 2002) involved the development of a conceptual framework based on a literature review and secondary data analysis. The Nursing Interventions Classification (NIC) was selected as a framework for the revision of the original B-NMDS. The second phase (November 2002-September 2003) focused on language development for six care programs evaluated by panels of clinical experts (N=75). These panels identified the following items as priorities for the revised B-NMDS: hospital financing, nurse staffing allocation, assessment of the appropriateness of hospitalisation, and quality management. During this period, we developed a draft instrument with 92 variables using the NIC. This led to an alpha version of a revised B-NMDS. The third phase (October 2003-December 2004) focused on data collection and validation of the new tool. The revised B-NMDS (alpha version) was tested in 158 nursing wards in 66 Belgian hospitals from December 2003 until March 2004. This test generated data for some 95,000 in-patient days. The interrater reliability of the revised B-NMDS was assessed. The criterion-related validity of the revised B-NMDS was compared to that of the original B-NMDS. The discriminative power of the revised B-NMDS was also assessed to select the most relevant variables for data collection. This resulted in a beta version of the revised B-NMDS in December 2004. The records of the revised B-NMDS were linked to the Hospital Discharge Dataset and other mandatory datasets to integrate the revised B-NMDS into the overall healthcare management system. The fourth phase (January 2005-December 2005) is presently focusing on information management. Nationwide implementation is foreseen by January 2007.

Belgium↗

Immune cell transcriptome datasets reveal novel leukocyte subset-specific genes and genes associated with allergic processes.

BACKGROUND: The precise function of various resting and activated leukocyte subsets remains unclear. For instance, mast cells, basophils, and eosinophils play important roles in allergic inflammation but also participate in other immunologic responses. One strategy to understand leukocyte subset function is to define the expression and function of subset-restricted molecules. OBJECTIVE: To use a microarray dataset and bioinformatics strategies to identify novel leukocyte markers as well as genes associated with allergic or innate responses. METHODS: By using Affymetrix microarrays, we generated an immune transcriptome dataset composed of gene profiles from all of the major leukocyte subsets, including rare enigmatic subsets such as mast cells, basophils, and plasma cells. We also assessed whether analysis of genes expressed commonly by certain groups of leukocytes, such as allergic leukocytes, might identify genes associated with particular responses. RESULTS: Transcripts highly restricted to a single leukocyte subset were readily identified (>2000 subset-specific transcripts), many of which have not been associated previously with leukocyte functions. Transcripts expressed exclusively by allergy-related leukocytes revealed well known as well as novel molecules, many of which presumably contribute to allergic responses. Likewise, Nearest Neighbor Analysis of genes coexpressed with Toll-like receptors identified genes of potential relevance for innate immunity. CONCLUSION: Gene profiles from all of the major human leukocyte subsets provide a powerful means to identify genes associated with single leukocyte subsets, or different types of immune response. CLINICAL IMPLICATIONS: A comprehensive dataset of gene expression profiles of human leukocytes should provide new targets or biomarkers for human inflammatory diseases.

Gene Expression Profiling↗

Detection of signal synchronizations in resting-state fMRI datasets.

In this paper, we propose a generic framework for the analysis of steady-state fMRI datasets, applied here to resting-state datasets. Our approach avoids the introduction of user-defined seed regions for the study of spontaneous activity. Unlike existing techniques, it yields a sparse representation of resting-state activity networks which can be characterized and investigated fairly easily in a semi-interactive fashion. We proceed in several steps, based on the idea that spectral coherence of the fMRI time courses in the low frequency band carries the information of interest. In particular, we address the question of building adapted representations of the data from the spectral coherence matrix. We analyze nine datasets taken from three subjects and show resting-state networks validated by EEG-fMRI simultaneous acquisition literature, with low intra-subject variability; we also discuss the merits of different (rapid/slow) fMRI acquisition schemes.

Algorithms↗

Diagnostic variability for schizophrenia and major depression in a large public mental health care system dataset.

Administrative datasets can provide information about mental health treatment in real world settings; however, an important limitation in using these datasets is the uncertainty regarding psychiatric diagnosis. To better understand the psychiatric diagnoses, we investigated the diagnostic variability of schizophrenia and major depression in a large public mental health system. Using schizophrenia and major depression as the two comparison diagnoses, we compared the variability of diagnoses assigned to patients with one recorded diagnosis of schizophrenia or major depression. In addition, for both of these diagnoses, the diagnostic variability was compared across seven types of treatment settings. Statistical analyses were conducted using t tests for continuous data and chi-square tests for categorical data. We found that schizophrenia had greater diagnostic variability than major depression (31% vs. 43%). For both schizophrenia and major depression, variability was significantly higher in jail and the emergency psychiatric unit than in inpatient or outpatient settings. These findings demonstrate that the variability of psychiatric diagnoses recorded in the administrative dataset of a large public mental health system varies by diagnosis and by treatment setting. Further research is needed to clarify the relationship between psychiatric diagnosis, diagnostic variability and treatment setting.

Adult↗