PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Dataset”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Exploring trafficking GTPase function by mRNA expression profiling: use of the SymAtlas web-application and the Membrome datasets.

Despite complete sequencing of the human and mouse genomes, functional annotation of novel gene function still remains a major challenge in mammalian biology. Emerging strategies to help elucidate unknown gene function include the analysis of tissue-specific patterns of mRNA expression. A recent study investigated the steady-state mRNA expression profiling of the vast majority of protein-encoding human and mouse genes across a panel of 79 human and 61 mouse nonredundant tissues. The microarray data from this study constitutes the Genomics Institute of Novartis Foundation (GNF) Human and Mouse Gene Atlases and is publicly available for exploration through the SymAtlas web-application (http://symatlas.gnf.org/). We have recently reported the use of these data and hierarchical clustering algorithms to generate a global overview of the distribution of Rabs, SNAREs, and coat machinery components, as well as their respective adaptors, effectors, and regulators. This systems biology approach led us to propose Rab-centric protein activity hubs as a framework for an integrated coding system, the membrome network, which orchestrates the dynamics of specialized membrane architecture of differentiated cells. Here, we describe the use of the SymAtlas web-application and the Membrome datasets to help explore trafficking GTPase function. The human and mouse membrome datasets are available through the Membrome homepage (http://www.membrome.org/) and correspond to subsets of the SymAtlas content restricted to known membrane trafficking components. Considering the fragmentary nature of the current reductionist approaches in elucidating trafficking component functions, the membrome datasets provide a more focused systems biology perspective that not only complements our current understanding of transport in complex tissues but also provides an integrated perspective of Rab activity in controlling membrane architecture.

Animals↗

Alignment of magnetic-resonance brain datasets with the stereotactical coordinate system.

Neuroanatomical and neurofunctional studies are often referenced to high-resolution magnetic-resonance brain datasets. For the analysis of the cortical surface, mapping of functional information on to the cortex or visualization, it is necessary to remove the outer surfaces of the brain. For intersubject comparison, it is useful to align the dataset with a coordinate system and introduce a spatial normalization. We describe an image processing chain that combines all of these steps in an interaction-free procedure. We report on a period of 2 years of routine application of this procedure, with >250 successfully processed datasets from healthy subjects and patients with various forms of brain damage.

Brain Diseases↗

Molecular dataset diversity indices and their applications to comparison of chemical databases and QSAR analysis

A new mutual molecular dataset diversity index (MMDDI), individual molecular dataset diversity index (IMDDI), and volume ratio (VR) are proposed to assess molecular dataset diversity. MMDDI and IMDDI can serve as valuable instruments for selecting monomer pools for combinatorial synthesis and in decision making about acquiring new databases. MMDDI can also be used as one of the criteria to estimate the quality of quantitative structure-activity relationship (QSAR) models aimed at the prediction of biological activities. The indices can be calculated directly from molecular descriptor values. The procedures applied for MMDDI and IMDDI calculations allow one to automatically compile lists of compounds, which can simplify molecular diversity analyses and database searching. The information can also be used for forming training and test sets in QSAR analysis.

Journal Article↗

Recursive partitioning for the prediction of cytochromes P450 2D6 and 1A2 inhibition: importance of the quality of the dataset.

The purpose of this study was to explore the use of detailed biological data in combination with a statistical learning method for predicting the CYP1A2 and CYP2D6 inhibition. Data were extracted from the Aureus-Pharma highly structured databases which contain precise measures and detailed experimental protocol concerning the inhibition of the two cytochromes. The methodology used was Recursive Partitioning, an easy and quick method to implement. The building of models was preceded by the evaluation of the chemical space covered by the datasets. The descriptors used are available in the MOE software suite. The models reached at least 80% of Accuracy and often exceeded this percentage for the Sensitivity (Recall), Specificity, and Precision parameters. CYP2D6 datasets provided 11 models with Accuracy over 80%, while CYP1A2 datasets counted 5 high-accuracy models. Our models can be useful to predict the ADME properties during the drug discovery process and are indicated for high-throughput screening.

Artificial Intelligence↗

Computational methods for comparison of large genomic and proteomic datasets reveal protein markers of metastatic cancer.

Large-scale genomic and proteomic analysis has provided a wealth of information on biologically relevant systems, and the ability to analyze this information is crucial to uncovering important biological relationships. However, it has proven difficult to compare large datasets from different sources due to different gene and protein identifiers assigned by individual laboratories and database systems. Here, we describe the design of a fully automated blast program (BlastPro) that facilitates rapid comparison of large protein-protein, nucleotide--nucleotide, or nucleotide--protein datasets from numerous, independent studies. Using this system, we compared several published genomic and proteomic databases for proteins that are upregulated in highly motile, metastatic tumor cells. Analysis of five independent studies comprised of greater than 1 x 10(6) genomic sequences and greater than 1,000 proteins revealed that the cytoskeletal-associated protein alpha-actinin is increased at both the mRNA and protein level in metastatic breast, prostate, and skin cancer cells. Interestingly, spatial analysis of alpha-actinin expression revealed that it is amplified 8-fold in the leading pseudopodium compared to the cell body compartment of migrating cells. These findings indicate that amplification of alpha-actinin and its localization to the leading pseudopodium are potential biomarkers of cancer progression to a more metastatic phenotype. Together, our results demonstrate that the BlastPro system can be used to compare large genomic and proteomic datasets to reveal important biological relationships including those associated with cancer progression.

Actinin↗

Predicting membrane protein types using residue-pair models based on reduced similarity dataset.

An algorithm to predict the membrane protein types based on the multi-residue-pair effect in the Markov model is proposed. For a newly constructed dataset of 835 membrane proteins with very low sequence similarity, the overall prediction accuracy has been achieved as high as 81.1% and 71.7% in the resubstitution and jackknife test, respectively, for a prediction of type I single-pass, type II single-pass, multi-pass membrane proteins, lipid chain-anchored and GPI-anchored membrane proteins. The improvement of about 11% in the jackknife test can be achieved compared with the component-coupled algorithm merely based on the amino acid composition (AAC approach). The improvement is also confirmed on a high similarity dataset and the other extrapolating test. The result implies that designing more incisive analysis tools, one should develop algorithms based on the representative dataset with lower sequence similarity. The present algorithm is useful to expedite the determination of the types and functions of new membrane proteins and may be useful for the systematic analysis of functional genome data in a large scale. The computer program is available on request.

Algorithms↗

QSAR modeling of datasets with enantioselective compounds using chirality sensitive molecular descriptors.

Shape descriptors used in 3D QSAR studies naturally take into account chirality; however, for flexible and structurally diverse molecules such studies require extensive conformational searching and alignment. QSAR modeling studies of two datasets of fragrance compounds with complex stereochemistry using simple alignment-free chirality sensitive descriptors developed in our laboratories are presented. In the first investigation, 44 alpha-campholenic derivatives with sandalwood odor were represented as derivatives of several common structural templates with substituents numbered according to their relative spatial positions in the molecules. Both molecular and substituent descriptors were used as independent variables in MLR calculations, and the best model was characterized by the training set q2 of 0.79 and external test set r2 of 0.95. In the second study, several types of chirality descriptors were employed in combinatorial QSAR modeling of 98 ambergris fragrance compounds. Among 28 possible combinations of seven types of descriptors and four statistical modeling techniques, k nearest neighbor classification with CoMFA descriptors was initially found to generate the best models with the internal and external accuracies of 76 and 89%, respectively. The same dataset was then studied using novel atom pair chirality descriptors (cAP). The cAP are based on a modified definition of the atomic chirality, in which the seniority of the substituents is defined by their relative partial charge values: higher values correspond to higher seniorities. The resulting models were found to have higher predictive power than those developed with CoMFA descriptors; the best model was characterized by the internal and external accuracies of 82 and 94%, respectively. The success of modeling studies using simple alignment free chirality descriptors discussed in this paper suggests that they should be applied broadly to QSAR studies of many datasets when compound stereochemistry plays an important role in defining their activity.

Ambergris↗

A protocol to select high quality datasets of ecotoxicity values for pesticides.

The key to any QSAR model is the underlying dataset. In order to construct a reliable dataset to develop a QSAR model for pesticide toxicity, we have derived a protocol to critically evaluate the quality of the underlying data. In developing an appropriate protocol that would enable data to be selected in constructing a QSAR, we concentrated on one toxicity end point, the 96 h LC50 from the acute rainbow trout study. This end point is key in pesticide regulation carried out under 91/414/EEC. The dataset used for this exercise was from the US EPA-OPP database.

Animals↗

Acquiring a four-dimensional computed tomography dataset using an external respiratory signal.

Four-dimensional (4D) methods strive to achieve highly conformal radiotherapy, particularly for lung and breast tumours, in the presence of respiratory-induced motion of tumours and normal tissues. Four-dimensional radiotherapy accounts for respiratory motion during imaging, planning and radiation delivery, and requires a 4D CT image in which the internal anatomy motion as a function of the respiratory cycle can be quantified. The aims of our research were (a) to develop a method to acquire 4D CT images from a spiral CT scan using an external respiratory signal and (b) to examine the potential utility of 4D CT imaging. A commercially available respiratory motion monitoring system provided an 'external' tracking signal of the patient's breathing. Simultaneous recording of a TTL 'X-Ray ON' signal from the CT scanner indicated the start time of CT image acquisition, thus facilitating time stamping of all subsequent images. An over-sampled spiral CT scan was acquired using a pitch of 0.5 and scanner rotation time of 1.5 s. Each image from such a scan was sorted into an image bin that corresponded with the phase of the respiratory cycle in which the image was acquired. The complete set of such image bins accumulated over a respiratory cycle constitutes a 4D CT dataset. Four-dimensional CT datasets of a mechanical oscillator phantom and a patient undergoing lung radiotherapy were acquired. Motion artefacts were significantly reduced in the images in the 4D CT dataset compared to the three-dimensional (3D) images, for which respiratory motion was not accounted. Accounting for respiratory motion using 4D CT imaging is feasible and yields images with less distortion than 3D images. 4D images also contain respiratory motion information not available in a 3D CT image.

Algorithms↗

Estimating dataset size requirements for classifying DNA microarray data.

A statistical methodology for estimating dataset size requirements for classifying microarray data using learning curves is introduced. The goal is to use existing classification results to estimate dataset size requirements for future classification experiments and to evaluate the gain in accuracy and significance of classifiers built with additional data. The method is based on fitting inverse power-law models to construct empirical learning curves. It also includes a permutation test procedure to assess the statistical significance of classification performance for a given dataset size. This procedure is applied to several molecular classification problems representing a broad spectrum of levels of complexity.

Algorithms↗

Estimating total Barthel scores from just three items: the European Stroke Database 'minimum dataset' for assessing functional status at discharge from hospital.

BACKGROUND: The European Stroke Database (ESDB) Project aims to develop a 'common clinical language' for stroke care by agreeing on terminology, definitions and clinical assessments. Each area of stroke assessment has a 'minimum dataset', which can be collected routinely and more detailed information can be added for particular studies. Measurement of patients' functional status at discharge is essential for assessing the impact of hospital care, but even simple activities of daily living scales like the Barthel index may not be easy to use routinely on busy acute units. We thus aimed to further simplify the 20-point Barthel index by reducing it to a few key items. METHODS: We initially analysed data on 169 consecutive stroke patients discharged from one British hospital and found that a simple formula involving the combined subscores for urinary continence (Blad), bed-chair transfers C (Trans) and indoor mobility (Mob)-(Blad + Trans + Mob) x 2.39 + 0.14-predicted the total BI score to within 1 point in 79% and to within 2 points in 95% of cases. We then tested this three-item Barthel index (BI3) in four different stroke datasets (total n = 824). RESULTS: The predictions were accurate to +/-1 point in 72-81% and to +/-2 points in 88-97% of cases, and BI3 accounted for 95% of the variance in total Barthel score. It was more accurate in patients without obvious mental impairment. CONCLUSIONS: For studies involving large groups of stroke patients, it is sufficient to know about each patient's continence, transfers and indoor mobility at discharge, in order to estimate the total Barthel score. These measures have therefore been incorporated, together with a simple observational measure of cognitive status, into the database minimum dataset for short-term functional outcome and are now being validated in international studies.

Activities of Daily Living↗

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256 Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms↗

SpecAlign--processing and alignment of mass spectra datasets.

SUMMARY: Pre-processing of chromatographic profile or mass spectral data is an important aspect of many types of proteomics and biomarker discovery experiments. Here we present a graphical computational tool, SpecAlign, that enables simultaneous visualization and manipulation of multiple datasets. SpecAlign not only provides all common processing functions, but also uniquely implements an algorithm that enables the complete alignment of each mass spectrum within a loaded dataset. We demonstrate its utility by aligning two datasets each containing six spectra; one set was acquired prior to instrument calibration and the other following calibration. AVAILABILITY: The software is free of charge and available for download from http://ptcl.chem.ox.ac.uk/~jwong/specalign. Supports Windows operating systems including Windows 9X/NT/2000/XP.

Algorithms↗

Inferring gene regulatory networks from multiple microarray datasets.

MOTIVATION: Microarray gene expression data has increasingly become the common data source that can provide insights into biological processes at a system-wide level. One of the major problems with microarrays is that a dataset consists of relatively few time points with respect to a large number of genes, which makes the problem of inferring gene regulatory network an ill-posed one. On the other hand, gene expression data generated by different groups worldwide are increasingly accumulated on many species and can be accessed from public databases or individual websites, although each experiment has only a limited number of time-points. RESULTS: This paper proposes a novel method to combine multiple time-course microarray datasets from different conditions for inferring gene regulatory networks. The proposed method is called GNR (Gene Network Reconstruction tool) which is based on linear programming and a decomposition procedure. The method theoretically ensures the derivation of the most consistent network structure with respect to all of the datasets, thereby not only significantly alleviating the problem of data scarcity but also remarkably improving the prediction reliability. We tested GNR using both simulated data and experimental data in yeast and Arabidopsis. The result demonstrates the effectiveness of GNR in terms of predicting new gene regulatory relationship in yeast and Arabidopsis. AVAILABILITY: The software is available from http://zhangorup.aporc.org/bioinfo/grninfer/, http://digbio.missouri.edu/grninfer/ and http://intelligent.eic.osaka-sandai.ac.jp or upon request from the authors.

Algorithms↗

Innovations in user-defined analysis: dynamic grouping and customized user datasets in VistaPHw.

Flexible, ready access to community health assessment data is a feature of innovative Web-based data query systems. An example is VistaPHw, which provides access to Washington state data and statistics used in community health assessment. Because of its flexible analysis options, VistaPHw customizes local, population-based results to be relevant to public health decision-making. The advantages of two innovations, dynamic grouping and the Custom Data Module, are described. Dynamic grouping permits the creation of user-defined aggregations of geographic areas, age groups, race categories, and years. Standard VistaPHw measures such as rates, confidence intervals, and other statistics may then be calculated for the new groups. Dynamic grouping has provided data for major, successful grant proposals, building partnerships with local governments and organizations, and informing program planning for community organizations. The Custom Data Module allows users to prepare virtually any dataset so it may be analyzed in VistaPHw. Uses for this module may include datasets too sensitive to be placed on a Web server or datasets that are not standardized across the state. Limitations and other system needs are also discussed.

Community Health Planning↗

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals↗

Combining datasets to predict the effects of regulation of environmental lead exposure in housing stock.

A model for children's blood lead concentrations as a function of environmental lead exposures was developed by combining two nationally representative sources of data that characterize the marginal distributions of blood lead and environmental lead with a third regional dataset that contains joint measures of blood lead and environmental lead. The complicating factor addressed in this article was the fact that methods for assessing environmental lead were different in the national and regional datasets. Relying on an assumption of transportability (that although the marginal distributions of blood lead and environmental lead may be different between the regional dataset and the nation as a whole, the joint relationship between blood lead and environmental lead is the same), the model makes use of a latent variable approach to estimate the joint distribution of blood lead and environmental lead nationwide.

Biometry↗

A chemical dataset for evaluation of alternative approaches to skin-sensitization testing.

Allergic contact dermatitis resulting from skin sensitization is a common occupational and environmental health problem. In recent years, the local lymph node assay (LLNA) has emerged as a practical option for assessing the skin-sensitization potential of chemicals. In addition to accurate identification of skin sensitizers, the LLNA can also provide a reliable measure of relative sensitization potency, information that is pivotal in successful management of human health risks. However, even with the significant animal welfare benefits provided by the LLNA, there is interest still in the development of non-animal test methods for skin sensitization. Here, we provide a dataset of chemicals that have been tested in the LLNA and the activity of which correspond with what is known of their potential to cause skin sensitization in humans. It is anticipated that this will be of value to other investigators in the evaluation and calibration of novel approaches to skin-sensitization testing. The materials that comprise this dataset encompass both the chemical and biological diversity of known chemical allergens and provide also examples of negative controls. It is hoped that this dataset will accelerate the development, evaluation and eventual validation of new approaches to skin-sensitization testing.

Allergens↗