PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Dataset”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Registration of 3D CT and ultrasound datasets of the spine using bone structures.

OBJECTIVE: In navigated orthopedic surgery, accurate registration of bones is of major interest. Usually, this registration is performed using landmarks positioned directly on the bone surface. These landmarks must be exposed during surgery. Our goal is to avoid the exposure of bone surface for the sole purpose of registration by using an intraoperative ultrasound device that can localize the bone through tissue. METHOD: We propose an algorithm for the registration of CT and ultrasound datasets that takes into account the fact that ultrasound produces very noisy images (speckle) and shows only parts of the bone surface. This part is made from the CT dataset. Next, a surface volume registration is performed by searching for a position of the estimated surface that maximizes the average gray value of the voxels in the ultrasound dataset covered by the surface. RESULTS: The algorithm was implemented and validated using an ex vivo preparation of a human lumbar spine with surrounding muscle tissue. On the basis of this data, the method has a large radius of convergence and a repeatability of 0.5 mm for displacement and 0.5 degrees for rotation. CONCLUSIONS: A robust algorithm for the registration of 3D CT and ultrasound datasets is presented. The computation time seems sufficiently short to permit intraoperative use.

Algorithms↗

A dataset of protein-protein interfaces generated with a sequence-order-independent comparison technique.

While there are a number of structurally non-redundant datasets of protein monomers, there is none of protein-protein interfaces. Yet, the availability of such a dataset is expected to provide an added insight into a number of investigations. First and foremost among these is analyzing the interfaces to obtain their prevailing architectures, the forces that account for the protein-protein associations and their packing considerations. Their comparisons with those of the monomers are likely to shed additional light on protein-protein recognition on the one hand and on the folding of the polypeptide chain on the other. Docking simulations are also expected to benefit from the existence of such a dataset. A major stumbling block to the generation of a dataset of interfaces has been that the interface is composed of at least two chains. Furthermore, in the interfaces, each of the chains might be represented by non-contiguous pieces. Their order in the interfaces being compared might be different as well. This discontinuity stems from the definition of an interface. An interface consists of interacting residues between the chains, and those that are in their vicinity in the supporting scaffold, within a certain distance threshold. This necessarily yields unordered fragments, as well as isolated residues. Our novel, efficient, sequence-order-independent structural comparison technique is ideally suited to handle the task of the generation of a library of structurally non-redundant protein-protein interfaces. As it is computer-vision based, it views atoms as collections of points in space, disregarding their chain connectivity. In this work, 351 interface-families are created. Comparisons of the interfaces, and separately, of the chains which contribute to them, yield some interesting cases. In one of the cases, while two interfaces are similar, the structure of only one of the two chains is similar between the two complexes. The structure of the second chain of the first complex differs from that of the second chain of the second complex. Here the structure of the cleft in the first chain dictates the specific binding interactions. In another case, while the interfaces in the two complexes are similar, both chains composing them differ between the complexes. Lastly, the chains composing the complexes are similar, but the interfaces are dissimilar, providing a set of data for investigations of the favorable orientations of protein-protein associations.

Algorithms↗

The hospital cost of vertebral fractures in the EU: estimates using national datasets.

The purpose of this study was to estimate the hospital cost of vertebral fractures in the EU using national datasets to explore some of the methodologic limitations associated with such an approach. Hospital costs for vertebral fractures across the EU were compared with the hospital costs associated with hip fractures. Additionally, these costs were placed into the health care context by making comparisons with national health care expenditure. All EU Ministries of Health were contacted to identify national datasets to estimate the average length of stay, cost per diem and the number of patients discharged with vertebral fractures. Where national information was not available expert opinion and data from the relevant literature were used instead. Countries show a marked difference in the length of stay between men and women, with differences ranging from 0.32 days in Austria to 20.2 days in Spain. The average hospitalization rate was found to be 8% across the EU, with higher rates found for men than for women. Interestingly a positive correlation between health expenditure per capita and hospitalization rates was found. The total cost of vertebral fractures in the EU was estimated at euro 377 million per year. Across the EU the hospital cost of a vertebral fracture was on average 63% that of a hip fracture. National datasets allow us to estimate the cost of vertebral fractures in the EU but show limitations. In the absence of large scale prospective studies, national datasets need to be further refined to ensure more accurate estimations of the cost of vertebral fractures in the EU.

Cost of Illness↗

Spatial distribution of forest fires and controlling factors in Andhra Pradesh, India using SPOT satellite datasets.

Fires are one of the major causes of forest disturbance and destruction in several dry deciduous forests of southern India. In this study, we use remote sensing data sets in conjunction with topographic, vegetation, climate and socioeconomic factors for determining the potential causes of forest fires in Andhra Pradesh, India. Spatial patterns in fire characteristics were analyzed using SPOT satellite remote sensing datasets. We then used nineteen different metrics in concurrence with fire count datasets in a robust statistical framework to arrive at a predictive model that best explained the variation in fire counts across diverse geographical and climatic gradients. Results suggested that, of all the states in India, fires in Andhra Pradesh constituted nearly 13.53% of total fires. District wise estimates of fire counts for Andhra Pradesh suggested that, Adilabad, Cuddapah, Kurnool, Prakasham and Mehbubnagar had relatively highest number of fires compared to others. Results from statistical analysis suggested that of the nineteen parameters, population density, demand of metabolic energy (DME), compound topographic index, slope, aspect, average temperature of the warmest quarter (ATWQ) along with literacy rate explained 61.1% of total variation in fire datasets. Among these, DME and literacy rate were found to be negative predictors of forest fires. In overall, this study represents the first statewide effort that evaluated the causative factors of fire at district level using biophysical and socioeconomic datasets. Results from this study identify important biophysical and socioeconomic factors for assessing 'forest fire danger' in the study area. Our results also identify potential 'hotspots' of fire risk, where fire protection measures can be taken in advance. Further this study also demonstrate the usefulness of best-subset regression approach integrated with GIS, as an effective method to assess 'where and when' forest fires will most likely occur.

Conservation of Natural Resources↗

An analysis of in vivo hprt mutant frequency in circulating T-lymphocytes in the normal human population: a comparison of four datasets.

In this paper, we have compared mutant frequency data at the hprt locus in circulating T-lymphocytes from four large datasets obtained in the UK (Sussex), the USA (Vermont), France (Paris) and The Netherlands (Leiden). In total, data from > 500 non-exposed individuals ranging in age from newborns (cord blood samples) to > 80 years old have been included in the analysis. Based on raw data provided by the four laboratories, a model is presented for the analysis of mutant frequency estimations for population monitoring. For three of the laboratories, a considerable body of data was provided on replicate estimates of mutant frequency from single blood samples, as well as estimates from repeat blood samples obtained over a period of time from many of the individual subjects. This enabled us to analyse the sources of variation in the estimation of mutant frequency. Although some variation was apparent in the results from the four laboratories, overall the data were in general agreement. Thus, in all laboratories, cellular cloning efficiency of T-cells was generally high (> 30%), although in each laboratory considerable variation between experiments and subjects was seen. Mutant frequency per clonable T-cell was in general found to be inversely related to cloning efficiency. With the exception of a few outliers (which are to be expected), mutant frequencies at this locus were in the same range in each dataset; no effect of subject gender was found, but an overall clear age effect was apparent. When log mutant frequency was analysed vs log (age + 0.5) a consistent trend from birth to old age was seen. In contrast, the effect of the smoking habit did differ between the laboratories, there being an association of smoking with a significant increase in mutant frequency in the Sussex and Leiden datasets, but not in those from the Vermont or Paris datasets. Possible reasons for this are discussed. One of the objectives of population monitoring is an ability to detect the effect of accidental or environmental exposure to mutagens and carcinogens among exposed persons. The large body of data from non-exposed subjects we have analysed in this paper has enabled us to estimate the size of an effect that could be detected, and the number of individuals required to detect a significant effect, taking known sources of variation into account.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent↗

Parallel processing of large datasets from NanoLC-FTICR-MS measurements.

A new approach for automatic parallel processing of large mass spectral datasets in a distributed computing environment is demonstrated to significantly decrease the total processing time. The implementation of this novel approach is described and evaluated for large nanoLC-FTICR-MS datasets. The speed benefits are determined by the network speed and file transfer protocols only and allow almost real-time analysis of complex data (e.g., a 3-gigabyte raw dataset is fully processed within 5 min). Key advantages of this approach are not limited to the improved analysis speed, but also include the improved flexibility, reproducibility, and the possibility to share and reuse the pre- and postprocessing strategies. The storage of all raw data combined with the massively parallel processing approach described here allows the scientist to reprocess data with a different set of parameters (e.g., apodization, calibration, noise reduction), as is recommended by the proteomics community. This approach of parallel processing was developed in the Virtual Laboratory for e-Science (VL-e), a science portal that aims at allowing access to users outside the computer research community. As such, this strategy can be applied to all types of serially acquired large mass spectral datasets such as LC-MS, LC-MS/MS, and high-resolution imaging MS results.

Algorithms↗

Deterministic projection by growing cell structure networks for visualization of high-dimensionality datasets.

Recent advances in clinical proteomics data acquisition have led to the generation of datasets of high complexity and dimensionality. We present here a visualization method for high-dimensionality datasets that makes use of neuronal vectors of a trained growing cell structure (GCS) network for the projection of data points onto two dimensions. The use of a GCS network enables the generation of the projection matrix deterministically rather than randomly as in random projection. Three datasets were used to benchmark the performance and to demonstrate the use of this deterministic projection approach in real-life scientific applications. Comparisons are made to an existing self-organizing map projection method and random projection. The results suggest that deterministic projection outperforms existing methods and is suitable for the visualization of datasets of very high dimensionality.

Algorithms↗

Feature-guided clustering of multi-dimensional flow cytometry datasets.

BACKGROUND: Flow cytometry produces large multi-dimensional datasets of the physical and molecular characteristics of individual cells. The objective of this study was to simplify the cytometry datasets by arranging or clustering "objects" (cells) into a smaller number of relatively homogeneous groups (clusters) on the basis of interobject similarities and dissimilarities. RESULTS: The algorithm was designed to be driven by histogram features; that is, the relevant single parameter histogram features were used to guide multidimensional k-means clustering without an a priori estimate of cluster number. To test this approach, we simulated cell-derived datasets using protein-coated microspheres (artificial "cells"). The microspheres were constructed to provide 119 populations in 40 samples. The feature-guided (FG) approach accurately identified 100% of the predetermined cluster combinations. In contrast, an approach based on the partition index (PI) cluster validity measure accurately identified 83.2% of the clusters. Direct comparisons of the two methods indicated that the FG method was significantly more accurate than PI in identifying both the number of clusters and the number of objects within the clusters (p<.0001). CONCLUSION: We conclude that parameter feature analysis can be used to effectively guide k-means clustering of flow cytometry datasets.

Algorithms↗

Synthetic DNA barcodes identify singlets in scRNA-seq datasets and evaluate doublet&#xa0;algorithms.

Single-cell RNA sequencing (scRNA-seq) datasets contain true single cells, or singlets, in addition to cells that coalesce during the protocol, or doublets. Identifying singlets with high fidelity in scRNA-seq is necessary to avoid false negative and false positive discoveries. Although several methodologies have been proposed, they are typically tested on highly heterogeneous datasets and lack a priori knowledge of true singlets. Here, we leveraged datasets with synthetically introduced DNA barcodes for a hitherto unexplored application: to extract ground-truth singlets. We demonstrated the feasibility of our framework, "singletCode," to evaluate existing doublet detection methods across a range of contexts. We also leveraged our ground-truth singlets to train a proof-of-concept machine learning classifier, which outperformed other doublet detection algorithms. Our integrative framework can identify ground-truth singlets and enable robust doublet detection in non-barcoded datasets.

Algorithms↗

Visualisation of biomedical datasets by use of growing cell structure networks: a novel diagnostic classification technique.

BACKGROUND: Medical research produces large multivariable datasets that are difficult to visualise and interpret intuitively. We describe a novel growing cell structure (GCS) technique that compresses multidimensional datasets into two dimensional maps with colour overlays that can be visually interpreted. METHODS: The two-dimensional map is self-discovered from the training set by distribution of cases to different nodes according to similarity between the cases at each node. Nodes are added to the map until there is no further significant reduction in error. The Parzen window method is used to estimate the probability distribution of the training cases, and this probability is converted to posterior class probabilities by use of Bayes' theorem. Classification performance can be assessed by means of receiver operating characteristic (ROC) curves. Colour maps of the values of each input variable at each node are constructed, which illustrate the relation between each input variable and the overall distribution of cases in the network map. FINDINGS: From a dataset of 11 input variables from 692 fine-needle aspirate samples from breast lesions, a 32-node network produced an area under the ROC curve of 0.96, which was not significantly different from that for logistic regression (0.98, z=1.09, p>0.05). Colour maps of the input variables showed that some variables had discrete distributions over exclusively benign or malignant areas of the network, and were thus discriminant, whereas others, such as foamy macrophages, covered both benign and malignant regions. INTERPRETATION: This technique produces dimensional compression that allows multidimensional data to be displayed as two-dimensional colour images. This envisioning of information allows the highly developed visuospatial abilities of human observers to perceive subtle inter-relations in the dataset.

Artificial Intelligence↗

Modelling the optimal radiotherapy regime for the control of T2 laryngeal carcinoma using parameters derived from several datasets.

PURPOSE: A number of previous studies have used direct maximum-likelihood methods to derive the values of radiobiological parameters of the linear-quadratic model for head and neck tumors from large clinical datasets. Time factors for accelerated repopulation were included, along with a lag period before the start of this repopulation. This study was performed to attempt to utilise these results from clinical datasets to compare treatment regimes in common clinical use in the UK, along with other schedules used historically in a number of clinical series in North America and elsewhere, and to determine if an optimal treatment regime could be derived based on these clinical data. METHODS: The biologically-based linear-quadratic model, applied to local tumor control and late morbidity, has been used to derive theoretical optimum (maximising tumor control whilst not exceeding tolerance for late reactions) radiotherapy schedules based on daily fractions. The specific case of T2 laryngeal carcinoma was considered as this is treated primarily by radiotherapy in many centers. Parameter values for local control were taken from previous analyses of several large single-center and national datasets. A time factor and a lag period were included in the modelling. Values for the alpha/beta ratio for late morbidity were used in the range 1-4 Gy, which is compatible with the limited range of values reported in the literature for particular complications following radiotherapy for head and neck cancer. Early reactions and their consequential late morbidity were not modelled in this study, but assumed to be within tolerance. RESULTS: For treatments using daily fractions there was a broad optimum treatment time of between 3-6 weeks. The theoretical optimum depended to some extent on the value of the alpha/beta ratio for late morbidity, but in many cases was at or just beyond the end of the purported lag period of 3-4 weeks, although small values of alpha/beta between 1-2 Gy favour longer treatment times. Similar results were obtained using a range of parameter values derived from four independent clinical datasets. CONCLUSION: The mathematical modelling of this broad range of once-daily treatments for most of which differences in local control and late morbidity are essentially undetectable (< 5%) has shown how this clinically-recognised phenomenon is interpreted in terms of the combination of dose-response slopes, fractionation sensitivities and time factors for both tumor control and normal tissue morbidity. Although the conclusions are inevitably tempered by a number of caveats concerning confounding factors in different centers; for example, the use of different treatment volumes, the present analysis provides a framework with which to explore the potential value of modifications to conventional treatment schedules, such as the use of multiple fractions per day.

Carcinoma↗

Predictive QSAR modeling based on diversity sampling of experimental datasets for the training and test set selection.

One of the most important characteristics of Quantitative Structure Activity Relashionships (QSAR) models is their predictive power. The latter can be defined as the ability of a model to predict accurately the target property (e.g., biological activity) of compounds that were not used for model development. We suggest that this goal can be achieved by rational division of an experimental SAR dataset into the training and test set, which are used for model development and validation, respectively. Given that all compounds are represented by points in multidimensional descriptor space, we argue that training and test sets must satisfy the following criteria: (i) Representative points of the test set must be close to those of the training set; (ii) Representative points of the training set must be close to representative points of the test set; (iii) Training set must be diverse. For quantitative description of these criteria, we use molecular dataset diversity indices introduced recently (Golbraikh, A., J. Chem. Inf. Comput. Sci., 40 (2000) 414-425). For rational division of a dataset into the training and test sets, we use three closely related sphere-exclusion algorithms. Using several experimental datasets, we demonstrate that QSAR models built and validated with our approach have statistically better predictive power than models generated with either random or activity ranking based selection of the training and test sets. We suggest that rational approaches to the selection of training and test sets based on diversity principles should be used routinely in all QSAR modeling research.

Algorithms↗

Predictive QSAR modeling based on diversity sampling of experimental datasets for the training and test set selection.

One of the most important characteristics of Quantitative Structure Activity Relashionships (QSAR) models is their predictive power. The latter can be defined as the ability of a model to predict accurately the target property (e.g., biological activity) of compounds that were not used for model development. We suggest that this goal can be achieved by rational division of an experimental SAR dataset into the training and test set, which are used for model development and validation, respectively. Given that all compounds are represented by points in multidimensional descriptor space, we argue that training and test sets must satisfy the following criteria: (i) Representative points of the test set must be close to those of the training set; (ii) Representative points of the training set must be close to representative points of the test set; (iii) Training set must be diverse. For quantitative description of these criteria, we use molecular dataset diversity indices introduced recently (Golbraikh, A., J. Chem. Inf. Comput. Sci., 40 (2000) 414-425). For rational division of a dataset into the training and test sets, we use three closely related sphere-exclusion algorithms. Using several experimental datasets, we demonstrate that QSAR models built and validated with our approach have statistically better predictive power than models generated with either random or activity ranking based selection of the training and test sets. We suggest that rational approaches to the selection of training and test sets based on diversity principles should be used routinely in all QSAR modeling research.

Algorithms↗

Phylogeny of Trichoptera (caddisflies): characterization of signal and noise within multiple datasets.

Trichoptera are holometabolous insects with aquatic larvae that, together with the Lepidoptera, make up the Amphiesmenoptera. Despite extensive previous morphological work, little phylogenetic agreement has been reached about the relationship among the three suborders--Annulipalpia, Spicipalpia, and Integripalpia--or about the monophyly of Spicipalpia. In an effort to resolve this conflict, we sequenced fragments of the large and small subunit nuclear ribosomal RNAs (1078 nt; D1, D3, V4-5), the nuclear elongation factor 1 alpha gene (EF-1 alpha; 1098 nt), and a fragment of mitochondrial cytochrome oxidase I (COI; 411 nt). Seventy adult and larval morphological characters were reanalyzed and added to molecular data in a combined analysis. We evaluated signal and homoplasy in each of the molecular datasets and attempted to rank the particular datasets according to how appropriate they were for inferring relationships among suborders. This evaluation included testing for conflict among datasets, comparing tree lengths among alternative hypotheses, measuring the left-skew of tree-length distributions from maximally divergent sets of taxa, evaluating the recovery of expected clades, visualizing whether or not substitutions were accumulating with time, and estimating nucleotide compositional bias. Although all these measures cast doubt on the reliability of the deep-level signal coming from the nucleotides of the COI and EF-1 alpha genes, these data could still be included in combined analyses without overturning the results from the most conservative marker, the rRNA. The different datasets were found to be evolving under extremely different rates. A site-specific likelihood method for dealing with combined data with nonoverlapping parameters was proposed, and a similar weighting scheme under parsimony was evaluated. Among our phylogenetic conclusions, we found Annulipalpia to be the most basal of the three suborders, with Spicipalpia and Integripalpia forming a clade. Monophyly of Annulipalpia and Integripalpia was confirmed, but the relationships among spicipalpians remain equivocal.

Animals↗

Utilisation of proteomics datasets generated via multidimensional protein identification technology (MudPIT).

Technological developments in proteomics have had a dramatic impact on biology in recent years. One of these developments--named multidimensional protein identification technology (MudPIT)--couples two-dimensional chromatography of peptides in mass spectrometry-compatible solutions directly to tandem mass spectrometry, allowing for the identification of proteins from highly complex mixtures. Since the initial descriptions of MudPIT, this approach has been implemented in the analysis of whole proteomes, organelles and protein complexes. Key aspects of many of the analyses are the validation of MudPIT datasets with alternate strategies and the integration of MudPIT datasets with other biochemical, cell biology or molecular biology approaches. This paper presents strategies for validating MudPIT datasets and incorporating these datasets into biologically driven experimental design.

Automation↗

Pretraining improves prediction of genomic datasets across species.

MOTIVATION: Recent studies suggest that deep neural network models trained on thousands of human genomic datasets can accurately predict genomic features, including gene expression and chromatin accessibility. However, training these models is computation- and time-intensive, and datasets of comparable size do not exist for most other organisms. RESULTS: Here, we identify modifications to an existing state-of-the-art model that improve model accuracy while reducing training time and computational cost. Using this streamlined model architecture, we investigate the ability of models pretrained on human genomic datasets to transfer performance to a variety of different tasks. Models pretrained on human data but fine-tuned on genomic datasets from diverse tissues and species achieved significantly higher prediction accuracy while significantly reducing training time compared to models trained from scratch, with Pearson correlation coefficients between experimental results and predictions as high as 0.8. Further, we found that including excessive training tasks decreased model performance and that this decrease could be partially but not completely rescued by fine-tuning. Thus, simplifying model architecture, applying pretrained models, and carefully considering the number of training tasks may be effective and economical techniques for building new models across data types, tissues, and species. AVAILABILITY AND IMPLEMENTATION: Code is available on GitHub and Figshare: https://github.com/optimizedlearning/genomicsML, https://doi.org/10.6084/m9.figshare.31796116.

Genomics↗

Robust multi-scale clustering of large DNA microarray datasets with the consensus algorithm.

MOTIVATION: Hierarchical and relocation clustering (e.g. K-means and self-organizing maps) have been successful tools in the display and analysis of whole genome DNA microarray expression data. However, the results of hierarchical clustering are sensitive to outliers, and most relocation methods give results which are dependent on the initialization of the algorithm. Therefore, it is difficult to assess the significance of the results. We have developed a consensus clustering algorithm, where the final result is averaged over multiple clustering runs, giving a robust and reproducible clustering, capable of capturing small signal variations. The algorithm preserves valuable properties of hierarchical clustering, which is useful for visualization and interpretation of the results. RESULTS: We show for the first time that one can take advantage of multiple clustering runs in DNA microarray analysis by collecting re-occurring clustering patterns in a co-occurrence matrix. The results show that consensus clustering obtained from clustering multiple times with Variational Bayes Mixtures of Gaussians or K-means significantly reduces the classification error rate for a simulated dataset. The method is flexible and it is possible to find consensus clusters from different clustering algorithms. Thus, the algorithm can be used as a framework to test in a quantitative manner the homogeneity of different clustering algorithms. We compare the method with a number of state-of-the-art clustering methods. It is shown that the method is robust and gives low classification error rates for a realistic, simulated dataset. The algorithm is also demonstrated for real datasets. It is shown that more biological meaningful transcriptional patterns can be found without conservative statistical or fold-change exclusion of data. AVAILABILITY: Matlab source code for the clustering algorithm ClusterLustre, and the simulated dataset for testing are available upon request from T.G. and O.W.

Algorithms↗

PLPD: reliable protein localization prediction from imbalanced and overlapped datasets.

Subcellular localization is one of the key functional characteristics of proteins. An automatic and efficient prediction method for the protein subcellular localization is highly required owing to the need for large-scale genome analysis. From a machine learning point of view, a dataset of protein localization has several characteristics: the dataset has too many classes (there are more than 10 localizations in a cell), it is a multi-label dataset (a protein may occur in several different subcellular locations), and it is too imbalanced (the number of proteins in each localization is remarkably different). Even though many previous works have been done for the prediction of protein subcellular localization, none of them tackles effectively these characteristics at the same time. Thus, a new computational method for protein localization is eventually needed for more reliable outcomes. To address the issue, we present a protein localization predictor based on D-SVDD (PLPD) for the prediction of protein localization, which can find the likelihood of a specific localization of a protein more easily and more correctly. Moreover, we introduce three measurements for the more precise evaluation of a protein localization predictor. As the results of various datasets which are made from the experiments of Huh et al. (2003), the proposed PLPD method represents a different approach that might play a complimentary role to the existing methods, such as Nearest Neighbor method and discriminate covariant method. Finally, after finding a good boundary for each localization using the 5184 classified proteins as training data, we predicted 138 proteins whose subcellular localizations could not be clearly observed by the experiments of Huh et al. (2003).

Algorithms↗