PubMed Health⌕ Search

Biomedical subjects

Jules J Berman

Publications and source records attributed to Jules J Berman.

10 recordsLinked to original sources

Classifying the precancers: a metadata approach.

BACKGROUND: During carcinogenesis, precancers are the morphologically identifiable lesions that precede invasive cancers. In theory, the successful treatment of precancers would result in the eradication of most human cancers. Despite the importance of these lesions, there has been no effort to list and classify all of the precancers. The purpose of this study is to describe the first comprehensive taxonomy and classification of the precancers. As a novel approach to disease classification, terms and classes were annotated with metadata (data that describes the data) so that the classification could be used to link precancer terms to data elements in other biological databases. METHODS: Terms in the UMLS (Unified Medical Language System) related to precancers were extracted. Extracted terms were reviewed and additional terms added. Each precancer was assigned one of six general classes. The entire classification was assembled as an XML (eXtensible Mark-up Language) file. A Perl script converted the XML file into a browser-viewable HTML (HyperText Mark-up Language) file. RESULTS: The classification contained 4700 precancer terms, 568 distinct precancer concepts and six precancer classes: 1) Acquired microscopic precancers; 2) acquired large lesions with microscopic atypia; 3) Precursor lesions occurring with inherited hyperplastic syndromes that progress to cancer; 4) Acquired diffuse hyperplasias and diffuse metaplasias; 5) Currently unclassified entities; and 6) Superclass and modifiers. CONCLUSION: This work represents the first attempt to create a comprehensive listing of the precancers, the first attempt to classify precancers by their biological properties and the first attempt to create a pathologic classification of precancers using standard metadata (XML). The classification is placed in the public domain, and comment is invited by the authors, who are prepared to curate and modify the classification.

Decision Support Systems, Clinical↗

A tool for sharing annotated research data: the "Category 0" UMLS (Unified Medical Language System) vocabularies.

BACKGROUND: Large biomedical data sets have become increasingly important resources for medical researchers. Modern biomedical data sets are annotated with standard terms to describe the data and to support data linking between databases. The largest curated listing of biomedical terms is the the National Library of Medicine's Unified Medical Language System (UMLS). The UMLS contains more than 2 million biomedical terms collected from nearly 100 medical vocabularies. Many of the vocabularies contained in the UMLS carry restrictions on their use, making it impossible to share or distribute UMLS-annotated research data. However, a subset of the UMLS vocabularies, designated Category 0 by UMLS, can be used to annotate and share data sets without violating the UMLS License Agreement. METHODS: The UMLS Category 0 vocabularies can be extracted from the parent UMLS metathesaurus using a Perl script supplied with this article. There are 43 Category 0 vocabularies that can be used freely for research purposes without violating the UMLS License Agreement. Among the Category 0 vocabularies are: MESH (Medical Subject Headings), NCBI (National Center for Bioinformatics) Taxonomy and ICD-9-CM (International Classification of Diseases-9-Clinical Modifiers). RESULTS: The extraction file containing all Category 0 terms and concepts is 72,581,138 bytes in length and contains 1,029,161 terms. The UMLS Metathesaurus MRCON file (January, 2003) is 151,048,493 bytes in length and contains 2,146,899 terms. Therefore the Category 0 vocabularies, in aggregate, are about half the size of the UMLS metathesaurus.A large publicly available listing of 567,921 different medical phrases were automatically coded using the full UMLS metatathesaurus and the Category 0 vocabularies. There were 545,321 phrases with one or more matches against UMLS terms while 468,785 phrases had one or more matches against the Category 0 terms. This indicates that when the two vocabularies are evaluated by their fitness to find at least one term for a medical phrase, the Category 0 vocabularies performed 86% as well as the complete UMLS metathesaurus. CONCLUSION: The Category 0 vocabularies of UMLS constitute a large nomenclature that can be used by biomedical researchers to annotate biomedical data. These annotated data sets can be distributed for research purposes without violating the UMLS License Agreement. These vocabularies may be of particular importance for sharing heterogeneous data from diverse biomedical data sets. The software tools to extract the Category 0 vocabularies are freely available Perl scripts entered into the public domain and distributed with this article.

Algorithms↗

The tissue microarray data exchange specification: a community-based, open source tool for sharing tissue microarray data.

BACKGROUND: Tissue Microarrays (TMAs) allow researchers to examine hundreds of small tissue samples on a single glass slide. The information held in a single TMA slide may easily involve Gigabytes of data. To benefit from TMA technology, the scientific community needs an open source TMA data exchange specification that will convey all of the data in a TMA experiment in a format that is understandable to both humans and computers. A data exchange specification for TMAs allows researchers to submit their data to journals and to public data repositories and to share or merge data from different laboratories. In May 2001, the Association of Pathology Informatics (API) hosted the first in a series of four workshops, co-sponsored by the National Cancer Institute, to develop an open, community-supported TMA data exchange specification. METHODS: A draft tissue microarray data exchange specification was developed through workshop meetings. The first workshop confirmed community support for the effort and urged the creation of an open XML-based specification. This was to evolve in steps with approval for each step coming from the stakeholders in the user community during open workshops. By the fourth workshop, held October, 2002, a set of Common Data Elements (CDEs) was established as well as a basic strategy for organizing TMA data in self-describing XML documents. RESULTS: The TMA data exchange specification is a well-formed XML document with four required sections: 1) Header, containing the specification Dublin Core identifiers, 2) Block, describing the paraffin-embedded array of tissues, 3)Slide, describing the glass slides produced from the Block, and 4) Core, containing all data related to the individual tissue samples contained in the array. Eighty CDEs, conforming to the ISO-11179 specification for data elements constitute XML tags used in the TMA data exchange specification. A set of six simple semantic rules describe the complete data exchange specification. Anyone using the data exchange specification can validate their TMA files using a software implementation written in Perl and distributed as a supplemental file with this publication. CONCLUSION: The TMA data exchange specification is now available in a draft form with community-approved Common Data Elements and a community-approved general file format and data structure. The specification can be freely used by the scientific community. Efforts sponsored by the Association for Pathology Informatics to refine the draft TMA data exchange specification are expected to continue for at least two more years. The interested public is invited to participate in these open efforts. Information on future workshops will be posted at http://www.pathologyinformatics.org (API we site).

Community Health Services↗

Concept-match medical data scrubbing. How pathology text can be used in research.

CONTEXT: In the normal course of activity, pathologists create and archive immense data sets of scientifically valuable information. Researchers need pathology-based data sets, annotated with clinical information and linked to archived tissues, to discover and validate new diagnostic tests and therapies. Pathology records can be used for research purposes (without obtaining informed patient consent for each use of each record), provided the data are rendered harmless. Large data sets can be made harmless through 3 computational steps: (1) deidentification, the removal or modification of data fields that can be used to identify a patient (name, social security number, etc); (2) rendering the data ambiguous, ensuring that every data record in a public data set has a nonunique set of characterizing data; and (3) data scrubbing, the removal or transformation of words in free text that can be used to identify persons or that contain information that is incriminating or otherwise private. This article addresses the problem of data scrubbing. OBJECTIVE: To design and implement a general algorithm that scrubs pathology free text, removing all identifying or private information. METHODS: The Concept-Match algorithm steps through confidential text. When a medical term matching a standard nomenclature term is encountered, the term is replaced by a nomenclature code and a synonym for the original term. When a high-frequency "stop" word, such as a, an, the, or for, is encountered, it is left in place. When any other word is encountered, it is blocked and replaced by asterisks. This produces a scrubbed text. An open-source implementation of the algorithm is freely available. RESULTS: The Concept-Match scrub method transformed pathology free text into scrubbed output that preserved the sense of the original sentences, while it blocked terms that did not match terms found in the Unified Medical Language System (UMLS). The scrubbed product is safe, in the restricted sense that the output retains only standard medical terms. The software implementation scrubbed more than half a million surgical pathology report phrases in less than an hour. CONCLUSIONS: Computerized scrubbing can render the textual portion of a pathology report harmless for research purposes. Scrubbing and deidentification methods allow pathologists to create and use large pathology databases to conduct medical research.

Computing Methodologies↗

Threshold protocol for the exchange of confidential medical data.

BACKGROUND: Medical researchers often need to share clinical data without violating patient confidentiality. Threshold cryptographic protocols divide messages into multiple pieces, no single piece containing information that can reconstruct the original message. The author describes and implements a novel threshold protocol that can be used to search, annotate or transform confidential data without breaching patient confidentiality. METHODS: The basic threshold protocol is: 1) Text is divided into short phrases; 2) Each phrase is converted by a one-way hash algorithm into a seemingly-random set of characters; 3) Threshold Piece 1 is composed of the list of all phrases, with each phrase followed by its one-way hash; 4) Threshold Piece 2 is composed of the text with all phrases replaced by their one-way hash values, and with high-frequency words preserved. Neither Piece 1 nor Piece 2 contains information linking patients to their records. The original text can be re-constructed from Piece 1 and Piece 2. RESULTS: The threshold algorithm produces two files (threshold pieces). In typical usage, Piece 2 is held by the data owner, and Piece 1 is freely distributed. Piece 1 can be annotated and returned to the owner of the original data to enhance the complete data set. Collections of Piece 1 files can be merged and distributed without identifying patient records. Variations of the threshold protocol are described. The author's Perl implementation is freely available. CONCLUSIONS: Threshold files are safe in the sense that they are de-identified and can be used for research purposes. The threshold protocol is particularly useful when the receiver of the threshold file needs to obtain certain concepts or data-types found in the original data, but does not need to fully understand the original data set.

Algorithms↗

Bethesda proposals for classification of nonlymphoid hematopoietic neoplasms in mice.

The hematopathology subcommittee of the Mouse Models of Human Cancers Consortium recognized the need for a classification of murine hematopoietic neoplasms that would allow investigators to diagnose lesions as well-defined entities according to accepted criteria. Pathologists and investigators worked cooperatively to develop proposals for the classification of lymphoid and nonlymphoid hematopoietic neoplasms. It is proposed here that nonlymphoid hematopoietic neoplasms of mice be classified in 4 broad categories: nonlymphoid leukemias, nonlymphoid hematopoietic sarcomas, myeloid dysplasias, and myeloid proliferations (nonreactive). Criteria for diagnosis and subclassification of these lesions include peripheral blood findings, cytologic features of hematopoietic tissues, histopathology, immunophenotyping, genetic features, and clinical course. Differences between murine and human lesions are reflected in the terminology and methods used for classification. This classification will be of particular value to investigators seeking to develop, use, and communicate about mouse models of human hematopoietic neoplasms.

Animals↗

Diagnosis of gastrointestinal stromal tumors: A consensus approach.

As a result of major recent advances in understanding the biology of gastrointestinal stromal tumors (GISTs), specifically recognition of the central role of activating KIT mutations and associated KIT protein expression in these lesions, and the development of novel and effective therapy for GISTs using the receptor tyrosine kinase inhibitor STI-571, these tumors have become the focus of considerable attention by pathologists, clinicians, and patients. Stromal/mesenchymal tumors of the gastrointestinal tract have long been a source of confusion and controversy with regard to classification, line(s) of differentiation, and prognostication. Characterization of the KIT pathway and its phenotypic implications has helped to resolve some but not all of these issues. Given the now critical role of accurate and reproducible pathologic diagnosis in ensuring appropriate treatment for patients with GIST, the National Institutes of Health convened a GIST workshop in April 2001 with the goal of developing a consensus approach to diagnosis and morphologic prognostication. Key elements of the consensus, as described herein, are the defining role of KIT immunopositivity in diagnosis and a proposed scheme for estimating metastatic risk in these lesions, based on tumor size and mitotic count, recognizing that it is probably unwise to use the definitive term "benign" for any GIST, at least at the present time.

Antineoplastic Agents↗

Diagnosis of gastrointestinal stromal tumors: a consensus approach.

As a result of major recent advances in understanding the biology of gastrointestinal stromal tumors (GIST), specifically recognition of the central role of activating KIT mutations and associated KIT protein expression in these lesions, and the development of novel and effective therapy for GISTs using the receptor tyrosine kinase inhibitor STI-571, these tumors have become the focus of considerable attention among pathologists, clinicians, and patients. Stromal/mesenchymal tumors of the gastrointestinal tract have long been a source of confusion and controversy with regard to classification, line(s) of differentiation, and prognostication. Characterization of the KIT pathway and its phenotypic implications has helped to resolve some but not all of these issues. Given the now critical role of accurate and reproducible pathologic diagnosis in ensuring appropriate treatment for patients with GIST, the National Institutes of Health (NIH) convened a GIST workshop in April 2001 with the goal of developing a consensus approach to diagnosis and morphologic prognostication. Key elements of the consensus, as described herein, are the defining role of KIT immunopositivity in diagnosis and a proposed scheme for estimating metastatic risk in these lesions, based on tumor size and mitotic count, recognizing that it is probably unwise to use the definitive term benign for any GIST, at least at the present time.

Antineoplastic Agents↗

Confidentiality issues for medical data miners.

The first task in any medical data mining effort is ensuring patient confidentiality. In the past, most data mining efforts ensured confidentiality by the dubious policy of withholding their raw data from colleagues and the public. A cursory review of medical informatics literature in the past decade reveals that much of what we have "learned" consists of assertions derived from confidential datasets unavailable for anyone's review. Without access to the original data, it is impossible to validate or improve upon a researcher's conclusions. Without access to research data, we are asked to accept findings as an act of faith, rather than as a scientific conclusion. This special issue of Artificial Intelligence in Medicine is devoted to medical data mining. The medical data miner has an obligation to conduct valid research in a way that protects human subjects. Today, data miners have the technical tools to merge large data collections and to distribute queries over disparate databases. In order to include patient-related data in shared databases, data miners will need methods to anonymize and deidentify data. This article reviews the human subject risks associated with medical data mining. This article also describes some of the innovative computational remedies that will permit researchers to conduct research AND share their data without risk to patient or institution.

Computer Security↗