PubMed Health⌕ Search

Biomedical subjects

Perry L Miller

Publications and source records attributed to Perry L Miller.

11 recordsLinked to original sources

Exploring the portability of informatics capabilities from a clinical application to a bioscience application.

This report describes XDesc (eXperiment Description), a pilot project that serves as a case study exploring the degree to which an informatics capability developed in a clinical application can be ported for use in the biosciences. In particular, XDesc uses the Entity-Attribute-Value database implementation (including a great deal of metadata-based functionality) developed in TrialDB, a clinical research database, for use in describing the samples used in microarray experiments stored in the Yale Microarray Database (YMD). XDesc was linked successfully to both TrialDB and YMD, and was used to describe the data in three different microarray research projects involving Drosophila. In the process, a number of new desirable capabilities were identified in the bioscience domain. These were implemented on a pilot basis in XDesc, and subsequently "folded back" into TrialDB itself, enhancing its capabilities for dealing with clinical data. This case study provides a concrete example of how informatics research and development in clinical and bioscience domains has the potential for synergy and for cross-fertilization.

Clinical Medicine↗

Training the next generation of informaticians: the impact of "BISTI" and bioinformatics--a report from the American College of Medical Informatics.

In 2002-2003, the American College of Medical Informatics (ACMI) undertook a study of the future of informatics training. This project capitalized on the rapidly expanding interest in the role of computation in basic biological research, well characterized in the National Institutes of Health (NIH) Biomedical Information Science and Technology Initiative (BISTI) report. The defining activity of the project was the three-day 2002 Annual Symposium of the College. A committee, comprised of the authors of this report, subsequently carried out activities, including interviews with a broader informatics and biological sciences constituency, collation and categorization of observations, and generation of recommendations. The committee viewed biomedical informatics as an interdisciplinary field, combining basic informational and computational sciences with application domains, including health care, biological research, and education. Consequently, effective training in informatics, viewed from a national perspective, should encompass four key elements: (1). curricula that integrate experiences in the computational sciences and application domains rather than just concatenating them; (2). diversity among trainees, with individualized, interdisciplinary cross-training allowing each trainee to develop key competencies that he or she does not initially possess; (3). direct immersion in research and development activities; and (4). exposure across the wide range of basic informational and computational sciences. Informatics training programs that implement these features, irrespective of their funding sources, will meet and exceed the challenges raised by the BISTI report, and optimally prepare their trainees for careers in a field that continues to evolve.

Computational Biology↗

Achieving evolvable Web-database bioscience applications using the EAV/CR framework: recent advances.

The EAV/CR framework, designed for database support of rapidly evolving scientific domains, utilizes metadata to facilitate schema maintenance and automatic generation of Web-enabled browsing interfaces to the data. EAV/CR is used in SenseLab, a neuroscience database that is part of the national Human Brain Project. This report describes various enhancements to the framework. These include (1) the ability to create "portals" that present different subsets of the schema to users with a particular research focus, (2) a generic XML-based protocol to assist data extraction and population of the database by external agents, (3) a limited form of ad hoc data query, and (4) semantic descriptors for interclass relationships and links to controlled vocabularies such as the UMLS.

Database Management Systems↗

Neuroscience data and tool sharing: a legal and policy framework for neuroinformatics.

The requirements for neuroinformatics to make a significant impact on neuroscience are not simply technical--the hardware, software, and protocols for collaborative research--they also include the legal and policy frameworks within which projects operate. This is not least because the creation of large collaborative scientific databases amplifies the complicated interactions between proprietary, for-profit R&D and public "open science." In this paper, we draw on experiences from the field of genomics to examine some of the likely consequences of these interactions in neuroscience. Facilitating the widespread sharing of data and tools for neuroscientific research will accelerate the development of neuroinformatics. We propose approaches to overcome the cultural and legal barriers that have slowed these developments to date. We also draw on legal strategies employed by the Free Software community, in suggesting frameworks neuroinformatics might adopt to reinforce the role of public-science databases, and propose a mechanism for identifying and allowing "open science" uses for data whilst still permitting flexible licensing for secondary commercial research.

Computational Biology↗

The integration of similar clinical research data collection instruments.

We devised an algorithm for integrating similar clinical research data collection instruments to create a common measurement instrument. We tested this algorithm using questions from several similar surveys. We encountered differing levels of granularity among questions and responses across surveys resulting in either the loss of granularity or data. This algorithm may make survey integration more systematic and efficient.

Algorithms↗

Creating knowledgebases to text-mine PUBMED articles using clustering techniques.

Knowledgebase-mediated text-mining approaches work best when processing the natural language of domain-specific text. To enhance the utility of our successfully tested program-NeuroText, and to extend its methodologies to other domains, we have designed clustering algorithms, which is the principal step in automatically creating a knowledgebase. Our algorithms are designed to improve the quality of clustering by parsing the test corpus to include semantic and syntactic parsing

Algorithms↗

Metadata-driven creation of data marts from an EAV-modeled clinical research database.

Generic clinical study data management systems can record data on an arbitrary number of parameters in an arbitrary number of clinical studies without requiring modification of the database schema. They achieve this by using an Entity-Attribute-Value (EAV) model for clinical data. While very flexible for creating transaction-oriented systems for data entry and browsing of individual forms, EAV-modeled data is unsuitable for direct analytical processing, which is the focus of data marts. For this purpose, such data must be extracted and restructured appropriately. This paper describes how such a process, which is non-trivial and highly error prone if performed using non-systematic approaches, can be automated by judicious use of the study metadata-the descriptions of measured parameters and their higher-level grouping. The metadata, in addition to driving the process, is exported along with the data, in order to facilitate its human interpretation.

Breast Neoplasms↗

ALFRED: An allele frequency database for anthropology.

The deluge of data from the human genome project (HGP) presents new opportunities for molecular anthropologists to study human variation through the promise of vast numbers of new polymorphisms (e.g., single nucleotide polymorphisms or SNPs). Collecting the resulting data into a single, easily accessible resource will be important to facilitate this research. We created a prototype Web-accessible database named ALFRED (ALelle FREquency Database, http://alfred.med.yale.edu/alfred/) to store and make publicly available allele frequency data on diverse polymorphic sites for many populations. In constructing this database, we considered many different concerns relating to the types of information needed for anthropology, population genetics, molecular genetics, and statistics, as well as issues of data integrity and ease of access to data. We also developed links to other Web-based databases as well as procedures for others to make links to the data in ALFRED. Here we present an overview of the issues considered and provisional solutions, as well as an example of data already available. It is our hope that this database will be useful for research and teaching in a wide range of fields, and that colleagues from various fields will contribute to making ALFRED an important resource for many studies as yet unforeseen.

Anthropology, Physical↗

Neuroinformatics: the integration of shared databases and tools towards integrative neuroscience.

There is significant interest amongst neuroscientists in sharing neuroscience data and analytical tools. The exchange of neuroscience data and tools between groups affords the opportunity to differently re-analyze previously collected data, encourage new neuroscience interpretations and foster otherwise uninitiated collaborations, and provide a framework for the further development of theoretically based models of brain function. Data sharing will ultimately reduce experimental and analytical error. Many small Internet accessible database initiatives have been developed and specialized analytical software and modeling tools are distributed within different fields of neuroscience. However, in addition large-scale international collaborations are required which involve new mechanisms of coordination and funding. Provided sufficient government support is given to such international initiatives, sharing of neuroscience data and tools can play a pivotal role in human brain research and lead to innovations in neuroscience, informatics and treatment of brain disorders. These innovations will enable application of theoretical modeling techniques to enhance our understanding of the integrative aspects of neuroscience. This article, authored by a multinational working group on neuroinformatics established by the Organization for Economic Co-operation and Development (OECD), articulates some of the challenges and lessons learned to date in efforts to achieve international collaborative neuroscience.

Computational Biology↗

Exploring issues of quality of service in a Next Generation Internet testbed: a case study using PathMaster.

This case study describes a project that explores issues of quality of service (QoS) relevant to the next-generation Internet (NGI), using the PathMaster application in a testbed environment. PathMaster is a prototype computer system that analyzes digitized cell images from cytology specimens and compares those images against an image database, returning a ranked set of "similar" cell images from the database. To perform NGI testbed evaluations, we used a cluster of nine parallel computation workstations configured as three subclusters using Cisco routers. This architecture provides a local "simulated Internet" in which we explored the following QoS strategies: (1) first-in-first-out queuing, (2) priority queuing, (3) weighted fair queuing, (4) weighted random early detection, and (5) traffic shaping. The study describes the results of using these strategies with a distributed version of the PathMaster system in the presence of different amounts of competing network traffic and discusses certain of the issues that arise. The goal of the study is to help introduce NGI QoS issues to the Medical Informatics community and to use the PathMaster NGI testbed to illustrate concretely certain of the QoS issues that arise.

Cell Biology↗