PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data commons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

Genetic diversity and function in the human cytosolic sulfotransferases.

Amino-acid substitutions, which result from common nonsynonymous (NS) polymorphisms, may dramatically alter the function of the encoded protein. Gaining insight into how these substitutions alter function is a step toward acquiring predictability. In this study, we incorporated gene resequencing, functional genomics, amino-acid characterization and crystal structure analysis for the cytosolic sulfotransferases (SULTs) to attempt to gain predictability regarding the function of variant allozymes. Previously, four SULT genes were resequenced in 118 DNA samples. With additional resequencing of the remaining eight SULT family members in the same DNA samples, a total of 217 polymorphisms were revealed. Of 64 polymorphisms identified within 8785 bp of coding regions from SULT genes examined, 25 were synonymous and 39 were NS. Overall, the proportion of synonymous changes was greater than expected from a random distribution of mutations, suggesting the presence of a selective pressure against amino-acid substitutions. Functional data for common variants of five SULT genes have been previously published. These data, together with the SULT1A1 variant allozyme data presented in this paper, showed that the major mechanism by which amino acid changes altered function in a transient expression system was through decreases in immunoreactive protein rather than changes in enzyme kinetics. Additional insight with regard to mechanisms by which NS single nucleotide polymorphisms alter function was sought by analysis of evolutionary conservation, physicochemical properties of the amino-acid substitutions and crystal structure analysis. Neither individual amino-acid characteristics nor structural models were able to accurately and reliably predict the function of variant allozymes. These results suggest that common amino-acid substitutions may not dramatically alter the protein structure, but affect interactions with the cellular environment that are currently not well understood.

Amino Acid Substitution↗

The Mouse Phenome Project.

The laboratory mouse is the organism of choice for many studies in biology and medicine. Reliable phenotypic data are essential for the full utility of genotypic information emerging from efforts to sequence human and mouse genomes. The Mouse Phenome Project has been organized to help accomplish this task by establishing a collection of baseline phenotypic data on commonly used and genetically diverse inbred mouse strains and making this information publicly available through a web-accessible database. The Mouse Phenome Database (MPD) is being developed to manage these data and to provide researchers with tools for exploring both raw phenotypic data and comparative summary analyses. The MPD serves as a repository for detailed protocols and raw data. This resource enables investigators to identify appropriate strains for (1) physiological testing, (2) drug discovery, (3) toxicology studies, (4) mutagenesis, (5) modeling human diseases, (6) QTL analyses and identification of new genes and (7) unraveling the influence of environment on genotype.

Animals↗

Polyhazard models for lifetime data.

We propose a polyhazard model to deal with lifetime data associated with latent competing risks. The causes of failure are assumed unobserved and affecting individuals independently. The general framework allows a broad class of hazard models that includes the most common hazard-based models. The model accommodates bathtub and multimodal hazards, keeping enough flexibility for common lifetime data that cannot be accommodated by usual hazard-based models. Maximum likelihood estimation is discussed, and parametric simulation is used for hypothesis testing.

Animals↗

[Time series analysis in environmental epidemiology: short-term effects of air pollution on mortality and morbidity].

This work gives an overview of design and analysis of temporal studies using aggregated data in air pollution epidemiology. In the last years time series are often used to study the short-term association between ambient air pollution levels and aggregated health data. Health endpoints are usually daily mortality and/or daily hospital admission data from routine health registers. Air quality data are commonly obtained from one (or a few) fixed site monitoring stations. To detect the temporal association between the time-pattern in air pollution and the time-pattern in health data particular attention needs to be given to the autocorrelation structure, to the seasonality and long term trend in the data, and to the weather variables. Poisson regression with autocorrelated residuals is the suitable statistical method to analyze time studies. Furthermore, the pollutant variable can be analyzed at different lag-times to account for short latency periods in the manifestation of diseases. Studies with temporal aggregated data show the same disadvantages of the ecologic studies, although, in this case, confounding is less of a problem. Temporal studies usually are based on a large database, so that sufficient power can be achieved to detect even weak associations. Finally, the exposure information on subjects is often better characterized by short-term fluctuations in ambient air quality than is the case in geographic aggregations.

Air Pollutants↗

An introductory practical guide to secondary data analysis in pediatric urology.

INTRODUCTION: Secondary data analysis (SDA) has become an increasingly important approach in pediatric urology, enabling the study of long-term outcomes, care variation, and disparities in populations with chronic or congenital urologic conditions. With the growing availability of large datasets, a structured approach to designing and conducting SDA studies is increasingly relevant. OBJECTIVES: To provide an introductory, practical guide to SDA in pediatric urology by (1) summarizing commonly used data sources with representative studies, (2) outlining a stepwise approach to designing and executing SDA studies, and (3) highlighting key methodological considerations, limitations, and opportunities for future work. STUDY DESIGN: Narrative review of existing literature and commonly used datasets relevant to pediatric urology, including administrative claims, hospital encounter databases, clinical registries, electronic health record networks, and population-based surveys. RESULTS: Data sources differ in scope, clinical granularity, longitudinal follow-up, and representativeness, and each is suited to specific research questions. We present a practical workflow for SDA, including dataset selection, cohort definition, and analytic planning. Linkage across datasets can provide a more comprehensive view of care patterns and outcomes, although feasibility is influenced by legal, technical, and data-quality constraints. DISCUSSION: SDA enables population-level analyses and the study of rare conditions that are challenging to evaluate through single-center or prospective designs. However, careful cohort definition, feasibility assessment, and awareness of data limitations are essential to ensure validity and interpretability. CONCLUSION: SDA provides a scalable, cost-efficient framework for generating meaningful evidence in pediatric urology. Continued efforts to harmonize data elements, improve linkage infrastructure, and support cross-institution collaboration will enhance the quality and impact of future research. This article provides a practical framework and examples to support the design and execution of SDA studies.

Humans↗

Near-infrared spectroscopic monitoring of a series of industrial batch processes using a bilinear grey model.

A good process understanding is the foundation for process optimization, process monitoring, end-point detection, and estimation of the end-product quality. Performing good process measurements and the construction of process models will contribute to a better process understanding. To improve the process knowledge it is common to build process models. These models are often based on first principles such as kinetic rates or mass balances. These types of models are also known as hard or white models. White models are characterized by being generally applicable but often having only a reasonable fit to real process data. Other commonly used types of models are empirical or black-box models such as regression and neural nets. Black-box models are characterized by having a good data fit but they lack a chemically meaningful model interpretation. Alternative models are grey models, which are combinations of white models and black models. The aim of a grey model is to combine the advantages of both black-box models and white models. In a qualitative case study of monitoring industrial batches using near-infrared (NIR) spectroscopy, it is shown that grey models are a good tool for detecting batch-to-batch variations and an excellent tool for process diagnosis compared to common spectroscopic monitoring tools.

Chemical Industry↗

Structure and stability of common sesquiterpenes.

We present data on the electronic structure, polarity and relative stability of 14 common sesquiterpenes. The data were obtained by a combination of spectroscopic and high-level theoretical analysis. The discussion also includes comments on possible implication of molecular properties for physiological behaviour.

Chromatography, Gas↗

Detection of consistently task-related activations in fMRI data with hybrid independent component analysis.

fMRI data are commonly analyzed by testing the time course from each voxel against specific hypothesized waveforms, despite the fact that many components of fMRI signals are difficult to specify explicitly. In contrast, purely data-driven techniques, by focusing on the intrinsic structure of the data, lack a direct means to test hypotheses of interest to the examiner. Between these two extremes, there is a role for hybrid methods that use powerful data-driven techniques to fully characterize the data, but also use some a priori hypotheses to guide the analysis. Here we describe such a hybrid technique, HYBICA, which uses the initial characterization of the fMRI data from Independent Component Analysis and allows the experimenter to sequentially combine assumed task-related components so that one can gracefully navigate from a fully data-derived approach to a fully hypothesis-driven approach. We describe the results of testing the method with two artificial and two real data sets. A metric based on the diagnostic Predicted Sum of Squares statistic was used to select the best number of spatially independent components to combine and utilize in a standard regressional framework. The proposed metric provided an objective method to determine whether a more data-driven or a more hypothesis-driven approach was appropriate, depending on the degree of mismatch between the hypothesized reference function and the features in the data. HYBICA provides a robust way to combine the data-derived independent components into a data-derived activation waveform and suitable confounds so that standard statistical analysis can be performed.

Algorithms↗

Cough and the common cold: ACCP evidence-based clinical practice guidelines.

OBJECTIVE: To review the literature on cough and the common cold. METHODS: MEDLINE was searched through May 2004 for studies published in the English language since 1980 on human subjects using the medical subject heading terms "cough" and "common cold." Selected case series and prospective descriptive clinical trials were reviewed. Additional references from these studies that were pertinent to the topic were also reviewed. RESULTS: Based on extrapolation from epidemiologic data, the common cold is believed to be the single most common cause of acute cough. The most likely mechanism is the direct irritation of upper airway structures. It is also clear that viral infections of the upper respiratory tract that produce the common cold syndrome frequently produce a rhinosinusitis. In the setting of a cold, the presence of abnormalities seen on sinus roentgenograms or sinus CT scans are frequently due to the viral infection and are not diagnostic of bacterial sinus infection. CONCLUSION: Cough due to the common cold is probably the most common cause of acute cough. In a significant subset of patients with "postinfectious" cough, the etiology is probably an inflammatory response triggered by a viral upper respiratory infection (ie, the common cold). The resultant subacute or chronic cough can be considered to be due to an upper airway cough syndrome, previously referred to as postnasal drip syndrome. This process can be self-perpetuating unless interrupted with active treatment.

Acute Disease↗

A geometric approach to the analysis of physiological flow data.

Physiological flow data are common in various medical fields. Examples include urinary, blood and expiratory flows. They are widely used in assessing functions in the urinary, circulatory, or pulmonary systems, respectively. Current statistical methods for analysing these flow data in clinical trials are either univariate analyses, which do not utilize all the information together, or some conventional multivariate methods (such as regression analyses) which yield results that do not render clear medical interpretations. This paper presents a new approach to analysing the flow data, using urinary flow as the primary focus. The basic idea and technical steps are applicable to other flow data as well. The proposed method aims to transform the flow measurements back to the shape of the flow graphs. Since the whole geometric pattern of the flow graph provides more information about the patient's flow condition than any individual flow parameter alone, the method is a meaningful way of combining and analysing the flow data in both statistical and clinical senses. The method is a three-stage procedure. Patients are classified into three classes in the first stage and then ranked in sequence in the second stage, according to the geometry of the shape pattern and some clinical criteria. The classification procedure is shown to be very reliable when compared with the clinician's visual evaluation, and hence can be implemented by computer programming to aid clinical trials involving many patients. The whole ranking score is then readily analysed at the third stage for comparing treatment effects by the analysis of covariance method based on ranks, with the post-treatment score as the response variable and the baseline score as the covariate. An example of a urinary flow data set is provided to illustrate the use of the procedure.

Analysis of Variance↗

Vehicular trauma triage by mechanism: avoidance of the unproductive evaluation.

An instrument was developed using routinely available field data to identify the sizable subgroup of stable vehicular trauma victims initially triaged to the trauma center by mechanism indicators alone who are in reality at minimal risk for serious injury. The six most common vehicular mechanism indicators seen at a level I trauma center were evaluated: rollover, head-on greater than 30 mph, intrusion, prolonged extrication, other death in same vehicle, and ejection. Review of 1235 consecutive trauma team activations yielded 349 victims with a qualifying vehicular mechanism. Outcome indicators were used to classify patients into two groups: Minor Injury (MI) and Severe Injury (SI). Nineteen common field data elements routinely reported on arrival by the regional Emergency Medical Service (EMS) personnel were then reviewed. Data patterns associated only with the MI group were sought. A checklist was developed for Mechanism vehicular trauma utilizing physiologic, anatomic, and neurologic elements. A single positive element would define trauma team activations. Retrospectively, use of this instrument would have excluded 56% of the MI group from unproductive trauma team referral, but nearly none of the SI group. We conclude that an identifiable subset of trauma patients referred by vehicular mechanism criteria alone could be safely evaluated on arrival in the emergency department as a form of secondary triage rather than by referral to the trauma team. The use of an appropriate exclusionary instrument can still preserve the sensitivity of trauma team activation for severely injured victims.

Accidents, Traffic↗

Mechanism of the process formation; podocytes vs. neurons.

In this review article we discuss the common mechanism for cellular process formation. Besides the podocyte, the mechanism of process formation, including cytoskeletal organization and signal transduction, etc., has been studied using neurons and glias as model systems. There has been an accumulation of data showing common cell biological features of the podocyte and the neuron: 1) Both cells possess long and short cell processes equipped with highly organized cytoskeletal systems; 2) Both show cytoskeletal segregation; microtubules (MTs) and intermediate filaments (IFs) in podocyte primary processes and in neurites, while actin filaments (AFs) are abundant in podocyte foot processes in neuronal synaptic regions; 3) In both cells, process formation is mechanically dependent on MTs, whose assembly is regulated by various microtubule- associated proteins (MAPs); 4) In both cells, process formation is positively regulated by PP2A, a Ser/Thr protein phosphatase; 5) In both cells, process formation is accelerated by laminin, an extracellular matrix protein. In addition, recent data from our and other laboratories have shown that podocyte processes share many features with neuronal dendrites: 1) Podocyte processes and neuronal dendrites possess MTs with mixed polarity, namely, plus-end-distal and minus-end-distal MTs coexist in these processes; 2) To establish the mixed polarity of MTs, both express CHO1/MKLP1, a kinesin-related motor protein, and when its expression is inhibited formation of both podocyte processes and neuronal dendrites is abolished; 3) The elongation of both podocyte processes and neuronal dendrites is supported by rab8-regulated basolateral-type membrane transport; 4) Both podocyte processes and neuronal dendrites express synaptopodin, an actin-associated protein, in a development-dependent manner; interestingly, in both cells, synaptopodin is localized not in the main shaft of processes but in thin short projections from the main shaft. We propose that the podocyte process and the neuronal dendrite share many features, while the neuronal axon should be thought of as an exceptionally differentiated cellular process.

Animals↗

Method of moments and treatment of nonrandom error.

If one has a convoluted fluorescence decay and wishes to analyze it for a sum of exponential, then one can begin by asking either of two questions: (1) What sum of exponentials best fits the data? (2) What physical decay parameters gave rise to the data? At first these two questions may sound equivalent; in fact, they represent different philosophical approaches to data analysis. In resolving the first question, one adjusts the decay parameters until a calculated curve agrees within arbitrarily chosen limits to the original data. This is what we did in the fourth section of Table II. The fit obtained was decent, but the resulting parameters were wrong. A more difficult approach is to design a method of data analysis which is intrinsically insensitive to the presence of anticipated errors, aiming directly at recovering the decay parameters without regard to the fit. This is what we have done with the method of moments with MD. If particular errors do not have much effect on the recovered parameters, then such a method of data analysis is said to be robust with respect to those errors. Robust methods are widely used in engineering but have not seen much introduction yet to biophysics. Least-squares, the basis of the commonly used data fitting methods for pulse fluorometry, is nonrobust with respect to underlying noise distributions. Isenberg has shown that least-squares is nonrobust with respect to the nonrandom light scatter, time origin shift, and lamp width errors as well. As shown in Isenberg's paper, as well as here, the method of moments with MD is quite robust with respect to these nonrandom errors. Perhaps question (1) could be modified to include all of the errors that might be present in the data; but then, how would one decide which errors to include and whether an error is present? What fitting criterion would tell one this? Why choose a method which depends so strongly on this information when robust alternatives exist? As a rule, fitting should not be used as a criterion for correct decay parameters, unless all of the significant nonrandom errors have been included in the fit. If one fits the data but has not incorporated an important error, then the best fit will necessarily give the wrong answer. The method of moments provides clear criteria for accepting or rejecting an analysis.(ABSTRACT TRUNCATED AT 400 WORDS)

Data Interpretation, Statistical↗

Fatal firearm-related injury surveillance in Maryland.

CONTEXT: Maryland began a statewide firearm-related injury surveillance system in 1995. The system now focuses on firearm-related deaths; a system to monitor nonfatal injuries is being developed. The system is passive; it accesses, integrates, and analyzes data collected by Maryland's Office of the Chief Medical Examiner, Maryland State Police, and Division of Health Statistics. OBJECTIVE: To evaluate the surveillance system's ability to ascertain cases in the absence of a standard for the true number of cases. DESIGN: Link records of the same firearm-related death captured by the surveillance system's multiple data sources, comparing the rate of false positives and false negatives, and assessing errors in linkage variables. SETTING: Maryland, 1991-1994. PARTICIPANTS: All deaths occurring in the state of Maryland as a result of a firearm-related injury. MAIN OUTCOME MEASURES: Sensitivity and positive predictive value. RESULTS: The system is extremely sensitive, detecting 99.61% of cases, and it has a very high positive predictive value, with 99.87% of the cases identified from medical examiner's office data being confirmed as actual cases. CONCLUSIONS: Maryland's database of information from the medical examiner's office is highly accurate for ascertaining firearm-related deaths that occur in the state. A unique identifier common across data sources would ease record linkage efforts, and improve the system's ability to monitor firearm-related deaths.

Adolescent↗

A caveat concerning independence estimating equations with multivariate binary data.

Clustered binary data occur commonly in both the biomedical and health sciences. In this paper, we consider logistic regression models for multivariate binary responses, where the association between the responses is largely regarded as a nuisance characteristic of the data. In particular, we consider the estimator based on independence estimating equations (IEE), which assumes that the responses are independent. This estimator has been shown to be nearly efficient when compared with maximum likelihood (ML) and generalized estimating equations (GEE) in a variety of settings. The purpose of this paper is to highlight a circumstance where assuming independence can lead to quite substantial losses of efficiency. In particular, when the covariate design includes within-cluster covariates, assuming independence can lead to a considerable loss of efficiency in estimating the regression parameters associated with those covariates.

Biometry↗

A model for quality achievement--the NORDKEM protein project.

UNLABELLED: The Nordic protein project demonstrates a model for the process of achieving analytical quality. GOAL: based on use of common reference intervals leading to the quality specifications. Creation of quality: through common high quality calibrator (with IFCC-values) (external factor) and individual trouble-shooting and guidelines (internal factor). Control of quality: with specially designed set of control samples and problem-related evaluation of control data. Establishing common reference intervals: through associated projects.

Blood Proteins↗

Avoiding model selection bias in small-sample genomic datasets.

MOTIVATION: Genomic datasets generated by high-throughput technologies are typically characterized by a moderate number of samples and a large number of measurements per sample. As a consequence, classification models are commonly compared based on resampling techniques. This investigation discusses the conceptual difficulties involved in comparative classification studies. Conclusions derived from such studies are often optimistically biased, because the apparent differences in performance are usually not controlled in a statistically stringent framework taking into account the adopted sampling strategy. We investigate this problem by means of a comparison of various classifiers in the context of multiclass microarray data. RESULTS: Commonly used accuracy-based performance values, with or without confidence intervals, are inadequate for comparing classifiers for small-sample data. We present a statistical methodology that avoids bias in cross-validated model selection in the context of small-sample scenarios. This methodology is valid for both k-fold cross-validation and repeated random sampling.

Algorithms↗

Use of administrative data to estimate mass vaccination campaign coverage, Burkina Faso, 1999.

Administrative coverage data are commonly used to assess coverage of mass vaccination campaigns. These estimates are obtained by dividing the number of doses administered by the number of children of eligible age, usually at the health district level. This study used data from a cluster survey conducted in each of the 53 Burkina Faso health districts immediately after 1999 the National Immunization Days to assess whether administrative estimates correlated with those obtained through survey and whether the former identified districts that achieved suboptimal coverage as measured by cluster survey. During the first round of the campaign there was no significant correlation between data obtained by either method. The correlation was only marginally better during the second round. Although useful to help plan the logistics of a campaign, administrative coverage data should be used with other evaluation techniques in order to determine the number of eligible children vaccinated during a mass campaign.

Burkina Faso↗