PubMed Health⌕ Search

Biomedical subjects

Barry Robson

Publications and source records attributed to Barry Robson.

9 recordsLinked to original sources

Data mining and clinical data repositories: Insights from a 667,000 patient data set.

Clinical repositories containing large amounts of biological, clinical, and administrative data are increasingly becoming available as health care systems integrate patient information for research and utilization objectives. To investigate the potential value of searching these databases for novel insights, we applied a new data mining approach, HealthMiner, to a large cohort of 667,000 inpatient and outpatient digital records from an academic medical system. HealthMiner approaches knowledge discovery using three unsupervised methods: CliniMiner, Predictive Analysis, and Pattern Discovery. The initial results from this study suggest that these approaches have the potential to expand research capabilities through identification of potentially novel clinical disease associations.

Clinical Chemistry Tests↗

The future of highly personalized health care.

We can surely lean well towards the optimistic in envisioning health care. In the world 10-25 years ahead of us. This optimism is based on rapid developments in genomics, the essential basis of molecular medicine, and on advances in computer power. At the time of writing this paper, the Human Genome Project was planned to have a working draft by 2000 and indeed completion was announced on June 26th from Washington. This paper describes the situation and vision at that time. Though there has been much subsequent more thought about the influence of genomics on healthcare, the aspirations and visions have not fundamentally changed from those of 2000, except for the greater attention to practical details that comes from increased confidence in the practicality of the vision.

Biomedical Technology↗

Studies in the assessment of folding quality for protein modeling and structure prediction.

A diagnostic for assessing the quality of a fold has been developed to which further criteria can be progressively added. The goal is to create a measure that can follow the status of a protein structure in a simulation or modeling process, when the answer (the experimental structure) is not known in advance, rather than simply reject deliberate misfolds. This places greater emphasis on the need to study, and calibrate against, marginal cases, i.e., unusual native structures, incomplete structures, partially erroneous X-ray structures, good models, poor models, and the effect of cofactors. The first three terms introduced in the diagnostic are appropriate core-forming properties or noncore properties of residues in relation to tertiary structure, appropriate neighboring structure density for each residue in relation to tertiary structure, and secondary structure consistency. While the method emerges as a useful simulation analysis tool, we find a need for further fine-tuning to diminish sensitivity to minor conformational changes that retain essential features of the fold, balanced against the need to obtain a more sensitive response when a conformational change involves less physically meaningful interatomic interactions. This dual utility is difficult to obtain: the investigation highlights some of the issues. Initial attempts to obtain it have led to terms in the diagnostic that are admittedly complex: simplifications must also be explored.

Animals↗

Clinical and pharmacogenomic data mining: 1. Generalized theory of expected information and application to the development of tools.

New scientific problems, arising from the human genome project, are challenging the classical means of using statistics. Yet quantified knowledge in the form of rules and rule strengths based on real relationships in data, as opposed to expert opinion, is urgently required for researcher and physician decision support. The problem is that with many parameters, the space to be analyzed is highly dimensional. That is, the combinations of data to examine are subject to a combinatorial explosion as the number of possible events (entries, items, sub-records) (a),(b),(c),... per record (a,b,c,..) increases, and hence much of the space is sparsely populated. These combinatorial considerations are particularly problematic for identifying those associations called "Unicorn Events" which occur significantly less than expected to the extent that they are never seen to be counted. To cope with the combinatorial explosion, a novel numerical "book keeping" approach is taken to generate information terms relating to the combinatorial subsets of events (a,b,c,..), and, most importantly, the zeta (Zeta) function is employed. The incomplete Zeta function zeta(s,n) with s = 1, in which frequencies of occurrence such as n = n(a,b,c,...) determine the range of summation n, is argued to be the natural choice of information function. It emerges from Bayesian integration, taken over the distribution of possible values of information measures for sparse and ample data alike. Expected mutual information l(a;b;c) in nats (i.e., natural units analogous to bits but based on the natural logarithm), such as is available to the observer, is measured as e.g., the difference zeta(s,o(a,b,c..)) - zeta(s,e(a,b,c..)) where o(a,b,c,..) and e(a,b,c,..) are, or relate to, the observed and expected frequencies of occurrence, respectively. For real values of s > 1 the qualitative impact of strongly (positively or negatively) ranked data is preserved despite several numerical approximations. As real s increases, and the output of the information functions converge into three values +1, 0, and -1 nats representing a trinary logic system. For quantitative data, a useful ad hoc method, to report sigma-normalized covariations in an analogous manner to mutual information for significance comparison purposes, is demonstrated. Finally, the potential ability to make use of mutual information in a complex biomedical study, and to include Bayesian prior information derived from statistical, tabular, anecdotal, and expert opinion is briefly illustrated.

Clinical Trials as Topic↗

Clinical and pharmacogenomic data mining: 2. A simple method for the combination of information from associations and multivariances to facilitate analysis, decision, and design in clinical research and practice.

The physician and researcher must ultimately be able to combine qualitative and quantitative features from a variety of combinations of observations on data of many component items (i.e., many dimensions), and hence reach simple conclusions about interpretation, rational courses of action, and design. In the first paper of this series, it was noted that such needs are challenging the classical means of using statistics. Hence, the paper proposed the use of a Generalized Theory of Expected Information or "Zeta Theory". The conjoint event [a,b,c,..] is seen as a rule of association for a,b,c,.. associated with a rule strength I(a;b;c;...) = xi(s,o[a,b,c,..]) - xi (s,e[a,b,c,...]), where xi is the incomplete Zeta Function. Here, o[a,b,c,...] is the observed, and e[a,b,c,..] the expected, frequency of occurrence of conjoint event [a,b,c,...]. The present paper explores how output from this approach might be assembled in a form better suited for decision support. Related to this is the difficulty that the treatment of covariance and multivariance was previously rendered as a "fuzzy association" so that the output would fall into a similar form as the true associations, but this was a somewhat ad hoc approach in which only the final I( ) had any meaning. Users at clinical research sites had subsequently requested an alternative approach in which "effective frequencies" o[ ] and e[ ] calculated from the above variances and used to evaluate I( ) give some intuitive feeling analogous to the association treatment, and this is explored here. Though the present paper is theoretical, real examples are used to illustrate application. One clinical-genomic example illustrates experimental design by identifying data which is, or is not, statistically germane to the study. We also report on some impressions based on applying these techniques in studies of real, extensive patient record data which are now emerging, as well as on molecular design data originally studied in part to test the ability to deduce the effects of simple natural patient sequence variations ("SNPs") on patient protein activity. On the basis of these study experiences, methods of rationalizing and condensing the rules implied by associations and variances between data, as well as discussion of the difficulty of what is meant by "condensed", are presented in the Appendix.

Biomedical Research↗

Genomic messaging system and DNA mark-up language for information-based personalized medicine with clinical and proteome research applications.

The convergence of clinical medicine and the Life Sciences, commencing with opportunities in clinical trials and clinically linked medical research, presents many novel challenges. The Genomic Messaging System (GMS) described here was originally developed as a tool for assembling clinical genomic records of individual and collective patients, and was then generalized to become a flexible workflow component that will link clinical records to a variety of computational biology research tools, for research and ultimately for a more personalized, focused, and preventative healthcare system. Prominent among the applications linked are protein science applications, including the rapid automated modeling of patient proteins with their individual structural polymorphisms. In an initial study, GMS formed the basis of a fully automated system for modeling patient proteins with structural polymorphisms as a basis for drug selection and ultimately design on an individual patient basis.

Clinical Medicine↗

Clinical and pharmacogenomic data mining: 3. Zeta theory as a general tactic for clinical bioinformatics.

A new approach, a Zeta Theory of observations, data, and data mining, is being forged from a theory of expected information into an even more cohesive and comprehensive form by the challenge of general genomic, pharmacogenomic, and proteomic data. In this paper, the focus is not on studies using the specific tool FANO (CliniMiner) but on extensions to a new broader theoretical approach, aspects of which can easily be implemented into, or otherwise support, excellent existing methods, such as forms of multivariate analysis and IBM's product Intelligent Miner. The theory should perhaps be distinguished from an existing purely number-theoretic area sometimes also known as Zeta Theory, which focuses on the Riemann Zeta Function and the ways in which it governs the distribution of prime numbers. However, Zeta Theory as used here overlaps heavily with it and actually makes use of these same matters. The distinction is that it enters from a Bayesian information theory and data representation perspective. It could thus be considered an application of the 'mathematician's version'. The application is by no means confined to areas of modern biomedicine, and indeed its generality, even merging into quantum mechanics, is a key feature. Other areas with some similar challenges as modern biology, and which have inspired data mining methods such as IBM's Intelligent Miner, include commerce. But for several reasons discussed, modern molecular biology and medicine seem particularly challenging, and this relates to the often irreducible high dimensionality of the data. This thus remains our main target.

Computational Biology↗

Genomic messaging system language including command extensions for clinical data categories.

This paper, in the area of clinical bioinformatics, highlights relatively efficient means of storing, exchanging, protecting, and searching human and other genomic data, so as to make the data securely accessible to researchers while respecting patient privacy. One important idea is that the GMSL language can be considered as an extension of the way DNA and protein sequences are written so as to carry with them the wishes of the patient in regard to fine-grained consent (as well as retaining the medical experts' cautions, instructions for use, and annotation), and this is carried, whatever environment (e.g., XML) that the data is from or whatever it is going to. At the deepest level, a stream of data expressed in GMSL resembles highly compressed stream of self-checking machine code. For the reader less familiar with the computational aspects, some simple examples illustrate how the raw language looks and works as a raw stream of (interpreted) bytes. The bioinformatics applications are not confined to the clinical domain. This paper completes the initial specification of the language as previously presented and reports on some important extensions including clinical data categories.

Base Sequence↗

The dragon on the gold: myths and realities for data mining in biomedicine and biotechnology using digital and molecular libraries.

To develop bioscience and personalized medicine in the post-genomic era, the biggest problem may be how to extract knowledge from the rich libraries of biomedical data. A particular dragon protects the gold therein: the dragon is the "curse of dimensionality" and its formidable fire weapon, which is burning researchers, is the "combinatorial explosion". This arises because many genomic, proteomic, clinical, and lifestyle factors may interact that cannot necessarily be considered on a simple pairwise or additive basis. A suggested theoretical solution--or at least "road map" that ameliorates management of these problems--borrows from several disciplines. It is undertaken also in the hope might also lead to research with broader impact on several unresolved issues in biotechnology: conversely, mathematical understanding of processes involving molecular libraries, such as cDNA libraries and DNA in the living cell itself, may open the opportunities to use biotechnology to construct nanotechnological storage and query systems.

Biomedical Research↗