PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “biological data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

BIOZON: a hub of heterogeneous biological data.

Biological entities are strongly related and mutually dependent on each other. Therefore, there is a growing need to corroborate and integrate data from different resources and aspects of biological systems in order to analyze them effectively. Biozon is a unified biological database that integrates heterogeneous data types such as proteins, structures, domain families, protein-protein interactions and cellular pathways, and establishes the relationships between them. All data are integrated on to a single graph schema centered around the non-redundant set of biological objects that are shared by each source. This integration results in a highly connected graph structure that provides a more complete picture of the known context of a given object that cannot be determined from any one source. Currently, Biozon integrates roughly 2 million protein sequences, 42 million DNA or RNA sequences, 32,000 protein structures, 150,000 interactions and more from sources such as GenBank, UniProt, Protein Data Bank (PDB) and BIND. Biozon augments source data with locally derived data such as 5 billion pairwise protein alignments and 8 million structural alignments. The user may form complex cross-type queries on the graph structure, add similarity relations to form fuzzy queries and rank the results based on analysis of the edge structure similar to Google PageRank, online at Biozon.org.

Computer Graphics↗

All systems go: launching cell simulation fueled by integrated experimental biology data.

Biological simulation serves to unify the basic elements of systems biology, namely, model selection, experimentation and model refinement. To select biochemical models for simulation, metabolome analysis can be performed using capillary electrophoresis or liquid chromatography coupled with mass spectrometry. In this manner, selected models can be elaborated with temporal/spatial gene and protein expression data obtained from model organisms such as Escherichia coli. The E. coli single gene deletion mutant library (KO collection) and His-tag/GFP-fusion single open reading frame clone expression library (ASKA) are powerful resources for this task. The integration of parallel experimental datasets into dynamic simulation tools forms the remaining challenge for the systematic analysis and elucidation of biological networks and holds promise for biotechnological applications.

Cell Physiological Phenomena↗

Connection between chromatographic data and biological data.

There are no previous references to the direct use of GLC data in the correlation of biological processes, but we show that GLC retention data can be used in the correlation of several such processes involving gaseous solutes. There are a number of reports of RP-HPLC and MEKC data being used in the correlation of biological processes, but they are mostly restricted as to the number and type of solute studied. We show that if chromatographic data are used to obtain solvation descriptors for solutes, and if these descriptors are then used in the correlation of biological processes, that this indirect connection is a much more powerful and generally applicable method than is the direct connection between chromatographic data and biological data.

Chromatography, Gel↗

Design of a description language for generating wrapper to collect biological data.

The biological data are scattered in various areas with various formats and they are changing continuously. Therefore, data integration becomes an important issue to provide researcher a dynamic access of data. In the data integration process, the method of extracting heterogeneous data dynamically from the data source is an essential part. Data extraction method using wrapper can provide flexibility and extensibility to an integration system.

Computational Biology↗

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem↗

MRS: a fast and compact retrieval system for biological data.

The biological data explosion of the 'omics' era requires fast access to many data types in rapidly growing data banks. The MRS server allows for very rapid queries in a large number of flat-file data banks, such as EMBL, UniProt, OMIM, dbEST, PDB, KEGG, etc. This server combines a fast and reliable backend with a very user-friendly implementation of all the commonly used information retrieval facilities. The MRS server is freely accessible at http://mrs.cmbi.ru.nl/. Moreover, the MRS software is freely available at http://mrs.cmbi.ru.nl/download/ for those interested in making their own data banks available via a web-based server.

Databases, Genetic↗

Computer video acquisition and analysis system for biological data.

The BIAS (Biological Image Analysis System) was developed to: (i) permit accurate entry and image processing of biological data; (ii) minimize the need for specialized hardware; and (iii) aid in the human genome mapping and other projects. The first mouse/cursor key-driven module was designed to be user interactive and readily accessible to many laboratories. It contains the DRSNDS programs which automate the entering of data in a systematic format. The types of data that can be entered utilizing this module are DNA-RNA gels from either a positive or negative Polaroid image, autoradiograms or biotinylated images from Southern, Northern and dot or slot blot hybridization analyses. The image is acquired using a video camera and then digitized for subsequent analysis. During the analysis graphical representations of the intermediate results are provided to assure user confidence. At any point within the program the user may obtain on-line help with the current task. The output displays the mol. wt of each individual component in the appropriate context. The present version of the program produces results comparable with a human interpreter for some data. Band shifting and optical density calculations are in a prototype form to permit evaluation of various techniques. Future work is directed at expanding the system's capabilities to interpret data from other biological analyses including DNA sequencing gels.

Algorithms↗

Secured distributed service to manage biological data on EGEE grid.

Biological data are most times published and then become public ones. They, then, do not need to be isolated or encrypted. But, in some cases, these data stemed from patients or are analyzed with, for instance, pharmaceutical or agronomics goals. Also in simple ways , these data, before to become public, have to be kept confidential while researchers haven't been able to publish their work or to register them. So they are a lot of cases where the integrity and the confidentiality of biological data have to be protected against unauthorized accesses. But, as these private data are also large datasets, they need high-throughput computing and huge data storage to processed, such as ones produced by complete genome projects. These requirements are enhanced in the context of a Grid such EGEE, where the computing and storage resources are distributed across a large-scale platform. We have developed a secured distributed service to manage biological data on grid: the EncFile encrypted files management system. We have deployed it on the production platform of the EGEE grid project. Thus we provided grid users with a user-friendly component that doesn't require any user privileges. And we have integrated into a bioinformatics grid portal associated to encrypted representative biological resources: world-famous databases and programs.

Computational Biology↗

Identification of global data and partitioning scheme for modeling biological data within the electronic medical record.

Using "Black Box" theory we analyzed human physiology. The major physiological means of communication are the vascular and nervous systems. The fundamental partitions of physiology are the vascular capillary fields and efferent and afferent fields of the nervous system. These fields are generally associated with organs and organ systems. Such analysis leads to the conclusion that the global biological data are information carried within the vascular and nervous systems. Data elements and processes within organs are important to other organs only through their effects on these global elements. Incorporation of these concepts into medical databases would allow the partitioning of the software around physiological systems. As a result of partitioning the utility of the electronic medical record, software could be greatly expanded.

Humans↗

Genomic pathways database and biological data management.

In this paper, we discuss the properties of biological data and challenges it poses for data management, and argue that, in order to meet the data management requirements for 'digital biology', careful integration of the existing technologies and the development of new data management techniques for biological data are needed. Based on this premise, we present PathCase: Case Pathways Database System. PathCase is an integrated set of software tools for modelling, storing, analysing, visualizing and querying biological pathways data at different levels of genetic, molecular, biochemical and organismal detail. The novel features of the system include: (i) genomic information integrated with other biological data and presented starting from pathways; (ii) design for biologists who are possibly unfamiliar with genomics, but whose research is essential for annotating gene and genome sequences with biological functions; (iii) database design, implementation and graphical tools which enable users to visualize pathways data in multiple abstraction levels and to pose exploratory queries; (iv) a wide range of different types of queries including, 'path' and 'neighbourhood queries' and graphical visualization of query outputs; and (v) an implementation that allows for web (XML)-based dissemination of query outputs (i.e. pathways data in BIOPAX format) to researchers in the community, giving them control on the use of pathways data.

Computational Biology↗

Biological data warehousing system for identifying transcriptional regulatory sites from gene expressions of microarray data.

Identification of transcriptional regulatory sites plays an important role in the investigation of gene regulation. For this propose, we designed and implemented a data warehouse to integrate multiple heterogeneous biological data sources with data types such as text-file, XML, image, MySQL database model, and Oracle database model. The utility of the biological data warehouse in predicting transcriptional regulatory sites of coregulated genes was explored using a synexpression group derived from a microarray study. Both of the binding sites of known transcription factors and predicted over-represented (OR) oligonucleotides were demonstrated for the gene group. The potential biological roles of both known nucleotides and one OR nucleotide were demonstrated using bioassays. Therefore, the results from the wet-lab experiments reinforce the power and utility of the data warehouse as an approach to the genome-wide search for important transcription regulatory elements that are the key to many complex biological systems.

Algorithms↗

A program for fitting of hypnograms and other biological data by orthogonal polynomials.

Many biological phenomenon are nonlinear and poorly approximated by linear regression. The program POLFIT calculates curvilinear regressions by fitting of orthogonal polynomials. This program, as well as two supporting programs (CORREC: disk storage, and CALCML: calculation of cumulated values of series of observations), was primarily designed for the study of the temporal organization of sleep components. They can be used as well for any other kind of biological data.

Computers↗

EnsMart: a generic system for fast and flexible access to biological data.

The EnsMart system (www.ensembl.org/EnsMart) provides a generic data warehousing solution for fast and flexible querying of large biological data sets and integration with third-party data and tools. The system consists of a query-optimized database and interactive, user-friendly interfaces. EnsMart has been applied to Ensembl, where it extends its genomic browser capabilities, facilitating rapid retrieval of customized data sets. A wide variety of complex queries, on various types of annotations, for numerous species are supported. These can be applied to many research problems, ranging from SNP selection for candidate gene screening, through cross-species evolutionary comparisons, to microarray annotation. Users can group and refine biological data according to many criteria, including cross-species analyses, disease links, sequence variations, and expression patterns. Both tabulated list data and biological sequence output can be generated dynamically, in HTML, text, Microsoft Excel, and compressed formats. A wide range of sequence types, such as cDNA, peptides, coding regions, UTRs, and exons, with additional upstream and downstream regions, can be retrieved. The EnsMart database can be accessed via a public Web site, or through a Java application suite. Both implementations and the database are freely available for local installation, and can be extended or adapted to 'non-Ensembl' data sets.

Animals↗

Self-organizing and self-correcting classifications of biological data.

MOTIVATION: Rapid, automated means of organizing biological data are required if we hope to keep abreast of the flood of data emanating from sequencing, microarray and similar high-throughput analyses. Faced with the need to validate the annotation of thousands of sequences and to generate biologically meaningful classifications based on the sequence data, we turned to statistical methods in order to automate these processes. RESULTS: An algorithm for automated classification based on evolutionary distance data was written in S. The algorithm was tested on a dataset of 1436 small subunit ribosomal RNA sequences and was able to classify the sequences according to an extant scheme, use statistical measurements of group membership to detect sequences that were misclassified within this scheme and produce a new classification. In this study, the use of the algorithm to address problems in prokaryotic taxonomy is discussed. AVAILABILITY: S-Plus is available from Insightful, Inc. An S-Plus implementation of the algorithm and the associated data are available at http://taxoweb.mmg.msu.edu/datasets

Algorithms↗

Multivariate analysis of clinical and biological data in cirrhotic patients: application to prognosis.

One hundred and thirty-one patients underwent clinical and biological investigation with the following determinations performed on the same day; presence or absence of ascites, icterus and/or encephalopathy, coagulation study, biochemical determinations including albumin, transferrin and immunoglobulins immunoassays. The principal component analysis of biological data showed two sets of highly representative and inversely correlated data; one included coagulation tests, albumin and transferrin, and the other included immunoglobulin A/transferrin ratio, immunoglobulin A and total bilirubin. Clinical and biological data were computed using discriminant analysis between dead and survivors. Six parameters were then selected (total bilirubin, encephalopathy, factor V, AST, antithrombin III and transferrin) giving a correct prognosis in 81.6% (31/38) of cases in a test sample. Neither ascites nor immunoglobulins were useful for the estimation of the prognosis.

Adult↗

HDBStat!: a platform-independent software suite for statistical analysis of high dimensional biology data.

BACKGROUND: Many efforts in microarray data analysis are focused on providing tools and methods for the qualitative analysis of microarray data. HDBStat! (High-Dimensional Biology-Statistics) is a software package designed for analysis of high dimensional biology data such as microarray data. It was initially developed for the analysis of microarray gene expression data, but it can also be used for some applications in proteomics and other aspects of genomics. HDBStat! provides statisticians and biologists a flexible and easy-to-use interface to analyze complex microarray data using a variety of methods for data preprocessing, quality control analysis and hypothesis testing. RESULTS: Results generated from data preprocessing methods, quality control analysis and hypothesis testing methods are output in the form of Excel CSV tables, graphs and an Html report summarizing data analysis. CONCLUSION: HDBStat! is a platform-independent software that is freely available to academic institutions and non-profit organizations. It can be downloaded from our website http://www.soph.uab.edu/ssg_content.asp?id=1164.

Algorithms↗

Physicochemical and biological data for the development of predictive organophosphorus pesticide QSARs and PBPK/PD models for human risk assessment.

A search of the scientific literature was carried out for physiochemical and biological data [i.e., IC50, LD50, Kp (cm/h) for percutaneous absorption, skin/water and tissue/blood partition coefficients, inhibition ki values, and metabolic parameters such as Vmax and Km] on 31 organophosphorus pesticides (OPs) to support the development of predictive quantitative structure-activity relationship (QSAR) and physiologically based pharmacokinetic and pharmacodynamic (PBPK/PD) models for human risk assessment. Except for work on parathion, chlorpyrifos, and isofenphos, very few modeling data were found on the 31 OPs of interest. The available percutaneous absorption, partition coefficients and metabolic parameters were insufficient in number to develop predictive QSAR models. Metabolic kinetic parameters (Vmax, Km) varied according to enzyme source and the manner in which the enzymes were characterized. The metabolic activity of microsomes should be based on the kinetic activity of purified or cDNA-expressed cytochrome P450s (CYPs) and the specific content of each active CYP in tissue microsomes. Similar requirements are needed to assess the activity of tissue A- and B-esterases metabolizing OPs. A limited amount of acetylcholinesterase (AChE), butyrylcholinesterase (BChE), and carboxylesterase (CaE) inhibition and recovery data were found in the literature on the 31 OPs. A program is needed to require the development of physicochemical and biological data to support risk assessment methodologies involving QSAR and PBPK/PD models.

Animals↗