PubMed Health⌕ Search

PubMed · 15165399

A classification paradigm for distributed vertically partitioned data.

Abstract

In general, pattern classification algorithms assume that all the features are available during the construction of a classifier and its subsequent use. In many practical situations, data are recorded in different servers that are geographically apart, and each server observes features of local interest. The underlying infrastructure and other logistics (such as access control) in many cases do not permit continual synchronization. Each server thus has a partial view of the data in the sense that feature subsets (not necessarily disjoint) are available at each server. In this article, we present a classification algorithm for this distributed vertically partitioned data. We assume that local classifiers can be constructed based on the local partial views of the data available at each server. These local classifiers can be any one of the many standard classifiers (e.g., neural networks, decision tree, k nearest neighbor). Often these local classifiers are constructed to support decision making at each location, and our focus is not on these individual local classifiers. Rather, our focus is constructing a classifier that can use these local classifiers to achieve an error rate that is as close as possible to that of a classifier having access to the entire feature set. We empirically demonstrate the efficacy of the proposed algorithm and also provide theoretical results quantifying the loss that results as compared to the situation where the entire feature set is available to any single classifier.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jayanta Basak, Ravi Kothari. 2004. A classification paradigm for distributed vertically partitioned data.. https://doi.org/10.1162/089976604323057470

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Ten-year experience with an emergency medicine resident research project requirement.

BACKGROUND: Controversy exists regarding the value and quality of required emergency medicine (EM) resident scholarly projects. OBJECTIVES: To describe the research designs and presentation rate at national scientific meetings and the publication rate of EM resident scholarly projects at a university-based residency program. METHODS: The authors reviewed the initial ten years (1993-2002) of resident scholarly projects from an EM residency program. Since the inception of the program, a formal research study has been required of all residents for residency graduation. Scholarly projects were reviewed and categorized by study design. Abstracts from the American Academy of Emergency Medicine (AAEM), American College of Emergency Physicians (ACEP), and Society for Academic Emergency Medicine (SAEM) annual meetings were searched to identify projects presented at any of these national meetings. A PubMed search for resident and faculty investigators was performed, and faculty and graduated residents were queried to identify all resident scholarly projects published in peer-reviewed journals. RESULTS: Eighty-seven residents produced 90 scholarly projects. Study designs were prospective data collection, 42 (47%); retrospective chart review, 38 (42%); survey, 5 (6%); animal, 4 (4%); and computer program development, 1 (1%). Of the 80 projects collecting patient data, 72 were conducted at a single center; 6, at two centers; and 2, at five centers each. Of the 42 prospective clinical studies, 27 (64%) were observational and 15 (36%) were interventional. Forty-six (51%) abstracts were presented at national meetings (SAEM, 20; ACEP, 19; AAEM, 3; and other, 4). Thirty-six (40%) of the projects have been published in peer-reviewed journals. Abstract presentation at national meetings (range, 13%-64% of projects per yr) and manuscript publication rates (range, 0-67% of projects per yr) were variable from year to year. CONCLUSIONS: Resident scholarly projects at one institution were equally likely to use a prospective or retrospective design, and most were conducted at a single center. More than half of the projects were presented at national research meetings, and more than a third were subsequently developed into manuscripts and published in peer-reviewed journals. When an original research study is required for satisfying the scholarly requirement for EM residency graduation, resident projects can contribute to the EM literature.

Abstracting and Indexing↗

Reviewer agreement trends from four years of electronic submissions of conference abstract.

BACKGROUND: The purpose of this study was to determine the inter-rater agreement between reviewers on the quality of abstract submissions to an annual national scientific meeting (Canadian Association of Emergency Physicians; CAEP) to identify factors associated with low agreement. METHODS: All abstracts were submitted using an on-line system and assessed by three volunteer CAEP reviewers blinded to the abstracts' source. Reviewers used an on-line form specific for each type of study design to score abstracts based on nine criteria, each contributing from two to six points toward the total (maximum 24). The final score was determined to be the mean of the three reviewers' scores using Intraclass Correlation Coefficient (ICC). RESULTS: 495 Abstracts were received electronically during the four-year period, 2001-2004, increasing from 94 abstracts in 2001 to 165 in 2004. The mean score for submitted abstracts over the four years was 14.4 (95% CI: 14.1-14.6). While there was no significant difference between mean total scores over the four years (p = 0.23), the ICC increased from fair (0.36; 95% CI: 0.24-0.49) to moderate (0.59; 95% CI: 0.50-0.68). Reviewers agreed less on individual criteria than on the total score in general, and less on subjective than objective criteria. CONCLUSION: The correlation between reviewers' total scores suggests general recognition of "high quality" and "low quality" abstracts. Criteria based on the presence/absence of objective methodological parameters (i.e., blinding in a controlled clinical trial) resulted in higher inter-rater agreement than the more subjective and opinion-based criteria. In future abstract competitions, defining criteria more objectively so that reviewers can base their responses on empirical evidence may lead to increased consistency of scoring and, presumably, increased fairness to submitters.

Abstracting and Indexing↗

Exploring supervised and unsupervised methods to detect topics in biomedical text.

BACKGROUND: Topic detection is a task that automatically identifies topics (e.g., "biochemistry" and "protein structure") in scientific articles based on information content. Topic detection will benefit many other natural language processing tasks including information retrieval, text summarization and question answering; and is a necessary step towards the building of an information system that provides an efficient way for biologists to seek information from an ocean of literature. RESULTS: We have explored the methods of Topic Spotting, a task of text categorization that applies the supervised machine-learning technique naïve Bayes to assign automatically a document into one or more predefined topics; and Topic Clustering, which apply unsupervised hierarchical clustering algorithms to aggregate documents into clusters such that each cluster represents a topic. We have applied our methods to detect topics of more than fifteen thousand of articles that represent over sixteen thousand entries in the Online Mendelian Inheritance in Man (OMIM) database. We have explored bag of words as the features. Additionally, we have explored semantic features; namely, the Medical Subject Headings (MeSH) that are assigned to the MEDLINE records, and the Unified Medical Language System (UMLS) semantic types that correspond to the MeSH terms, in addition to bag of words, to facilitate the tasks of topic detection. Our results indicate that incorporating the MeSH terms and the UMLS semantic types as additional features enhances the performance of topic detection and the naïve Bayes has the highest accuracy, 66.4%, for predicting the topic of an OMIM article as one of the total twenty-five topics. CONCLUSION: Our results indicate that the supervised topic spotting methods outperformed the unsupervised topic clustering; on the other hand, the unsupervised topic clustering methods have the advantages of being robust and applicable in real world settings.

Abstracting and Indexing↗