PubMed Health⌕ Search

PubMed · 11038333

Object-oriented parsing of biological databases with Python.

Abstract

MOTIVATION: While database activities in the biological area are increasing rapidly, rather little is done in the area of parsing them in a simple and object-oriented way. RESULTS: We present here an elegant, simple yet powerful way of parsing biological flat-file databases. We have taken EMBL, SWISSPROT and GENBANK as examples. EMBL and SWISS-PROT do not differ much in the format structure. GENBANK has a very different format structure than EMBL and SWISS-PROT. Extracting the desired fields in an entry (for example a sub-sequence with an associated feature) for later analysis is a constant need in the biological sequence-analysis community: this is illustrated with tools to make new splice-site databases. The interface to the parser is abstract in the sense that the access to all the databases is independent from their different formats, since parsing instructions are hidden.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

C Ramu, C Gemünd, T J Gibson. 2000. Object-oriented parsing of biological databases with Python.. https://doi.org/10.1093/bioinformatics%2F16.7.628

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Using a geographical information system to plan a malaria control programme in South Africa.

INTRODUCTION: Sustainable control of malaria in sub-Saharan Africa is jeopardized by dwindling public health resources resulting from competing health priorities that include an overwhelming acquired immunodeficiency syndrome (AIDS) epidemic. In Mpumalanga province, South Africa, rational planning has historically been hampered by a case surveillance system for malaria that only provided estimates of risk at the magisterial district level (a subdivision of a province). METHODS: To better map control programme activities to their geographical location, the malaria notification system was overhauled and a geographical information system implemented. The introduction of a simplified notification form used only for malaria and a carefully monitored notification system provided the good quality data necessary to support an effective geographical information system. RESULTS: The geographical information system displays data on malaria cases at a village or town level and has proved valuable in stratifying malaria risk within those magisterial districts at highest risk, Barberton and Nkomazi. The conspicuous west-to-east gradient, in which the risk rises sharply towards the Mozambican border (relative risk = 4.12, 95% confidence interval = 3.88-4.46 when the malaria risk within 5 km of the border was compared with the remaining areas in these two districts), allowed development of a targeted approach to control. DISCUSSION: The geographical information system for malaria was enormously valuable in enabling malaria risk at town and village level to be shown. Matching malaria control measures to specific strata of endemic malaria has provided the opportunity for more efficient malaria control in Mpumalanga province.

Databases, Factual↗

Development of a virtual screening method for identification of "frequent hitters" in compound libraries.

A computer-based method was developed for rapid and automatic identification of potential "frequent hitters". These compounds show up as hits in many different biological assays covering a wide range of targets. A scoring scheme was elaborated from substructure analysis, multivariate linear and nonlinear statistical methods applied to several sets of one and two-dimensional molecular descriptors. The final model is based on a three-layered neural network, yielding a predictive Matthews correlation coefficient of 0.81. This system was able to correctly classify 90% of the test set molecules in a 10-times cross-validation study. The method was applied to database filtering, yielding between 8% (compilation of trade drugs) and 35% (Available Chemicals Directory) potential frequent hitters. This filter will be a valuable tool for the prioritization of compounds from large databases, for compound purchase and biological testing, and for building new virtual libraries.

Databases, Factual↗